AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

Cache Writes Are Counted Twice For Priority And Flex In Langchain-Openai

In langchain-openai 1.6.2, Responses API usage metadata for 'priority' and 'flex' model providers double-counts cache_write_tokens: the non-cached input bucket only subtracts cached_tokens, not cache_write_tokens, causing inflated cost estimates.

mediumConfidence 95%Langchain-OpenaiAffected V1.6.2

Origin Analysis

The `_create_usage_metadata_responses` function (and the chat completions converter) computes the `priority`/`flex` input bucket as `input_tokens - cached_tokens`, omitting `cache_write_tokens`. Since `cache_write_tokens` are included in `input_tokens` and also recorded separately in `priority_cache_creation`/`flex_cache_creation`, they are counted twice. OpenAI's cost formula requires subtracting both cached reads and writes from the base input tokens.
Run the provided Python snippet: ```python from langchain_openai.chat_models.base import _create_usage_metadata_responses usage = { 'input_tokens': 2304, 'input_tokens_details': {'cached_tokens': 256, 'cache_write_tokens': 1024}, 'output_tokens': 50, 'total_tokens': 2354, } details = _create_usage_metadata_responses(usage, 'priority')['input_token_details'] print(details) ``` Actual output: {'priority_cache_read': 256, 'priority_cache_creation': 1024, 'priority': 2048} Expected: {'priority_cache_read': 256, 'priority_cache_creation': 1024, 'priority': 1024}

Fixing Code Block

from typing import Any, Union from langchain_core.usage_metadata import UsageMetadata def _create_usage_metadata_responses(usage: dict, model_provider: str) -> UsageMetadata: input_token_details: dict = {} output_token_details: dict = {} if model_provider in ('openai', 'azure-openai', 'azure_openai'): input_token_details = { 'prompt_tokens': usage['input_tokens'], 'total_tokens': usage['total_tokens'], 'cached_tokens': usage.get('input_tokens_details', {}).get('cached_tokens', 0), 'cache_read_input_tokens': usage.get('input_tokens_details', {}).get('cached_tokens', 0), 'cache_creation_input_tokens': usage.get('input_tokens_details', {}).get('cache_write_tokens', 0), } elif model_provider in ('priority', 'flex'): cached_tokens = usage.get('input_tokens_details', {}).get('cached_tokens', 0) cache_write_tokens = usage.get('input_tokens_details', {}).get('cache_write_tokens', 0) input_tokens = usage['input_tokens'] cache_field = f'{model_provider}_cache_read' creation_field = f'{model_provider}_cache_creation' input_field = model_provider input_token_details[cache_field] = cached_tokens input_token_details[creation_field] = cache_write_tokens input_token_details[input_field] = input_tokens - cached_tokens - cache_write_tokens else: input_token_details = { 'prompt_tokens': usage['input_tokens'], 'total_tokens': usage['total_tokens'], } output_token_details = { 'completion_tokens': usage['output_tokens'], 'total_tokens': usage['total_tokens'], } return UsageMetadata( input_tokens=usage['input_tokens'], output_tokens=usage['output_tokens'], total_tokens=usage['total_tokens'], input_token_details=input_token_details, output_token_details=output_token_details, ) def _create_usage_metadata(usage: Union[dict, Any], model_provider: str) -> UsageMetadata: input_token_details: dict = {} output_token_details: dict = {} if isinstance(usage, dict): input_tokens = usage.get('prompt_tokens', 0) output_tokens = usage.get('completion_tokens', 0) total_tokens = usage.get('total_tokens', 0) prompt_tokens_details = usage.get('prompt_tokens_details', {}) or {} cached_tokens = prompt_tokens_details.get('cached_tokens', 0) cache_write_tokens = prompt_tokens_details.get('cache_write_tokens', 0) else: input_tokens = getattr(usage, 'prompt_tokens', 0) output_tokens = getattr(usage, 'completion_tokens', 0) total_tokens = getattr(usage, 'total_tokens', 0) prompt_tokens_details = getattr(usage, 'prompt_tokens_details', None) or {} cached_tokens = getattr(prompt_tokens_details, 'cached_tokens', 0) cache_write_tokens = getattr(prompt_tokens_details, 'cache_write_tokens', 0) if model_provider in ('priority', 'flex'): cache_field = f'{model_provider}_cache_read' creation_field = f'{model_provider}_cache_creation' input_field = model_provider input_token_details[cache_field] = cached_tokens input_token_details[creation_field] = cache_write_tokens input_token_details[input_field] = input_tokens - cached_tokens - cache_write_tokens else: input_token_details = { 'prompt_tokens': input_tokens, 'total_tokens': total_tokens, 'cached_tokens': cached_tokens, 'cache_read_input_tokens': cached_tokens, 'cache_creation_input_tokens': cache_write_tokens, } output_token_details = { 'completion_tokens': output_tokens, 'total_tokens': total_tokens, } return UsageMetadata( input_tokens=input_tokens, output_tokens=output_tokens, total_tokens=total_tokens, input_token_details=input_token_details, output_token_details=output_token_details, )
The fix subtracts `cache_write_tokens` from the `input_tokens` when building the `priority` or `flex` input bucket. This aligns with OpenAI's documented formula: non-cached input tokens = total input tokens - cached reads - cache writes. Both the Responses and Chat Completions converters are updated to avoid inconsistent metadata. Existing tests for priority and flex should be updated to reflect the corrected values.

Edge Case Audit

This change alters the `input_token_details` mapping for priority/flex. Users relying on the previous `priority`/`flex` value for cost estimation will see lower numbers, which is correct. Ensure `cached_tokens` and `cache_write_tokens` are present (default to 0) and never exceed `input_tokens` to avoid negative values. The Chat Completions converter may have a different function name in older versions; verify and apply the same subtraction. Rollback: revert the subtraction line to `input_tokens - cached_tokens` if unexpected cost calculation issues arise. Additionally, update unit tests to assert the corrected values, as existing tests may fail.

Ecosystem Topology