Cache Writes Are Counted Twice For Priority And Flex In Langchain-Openai
In langchain-openai 1.6.2, Responses API usage metadata for 'priority' and 'flex' model providers double-counts cache_write_tokens: the non-cached input bucket only subtracts cached_tokens, not cache_write_tokens, causing inflated cost estimates.
The `_create_usage_metadata_responses` function (and the chat completions converter) computes the `priority`/`flex` input bucket as `input_tokens - cached_tokens`, omitting `cache_write_tokens`. Since `cache_write_tokens` are included in `input_tokens` and also recorded separately in `priority_cache_creation`/`flex_cache_creation`, they are counted twice. OpenAI's cost formula requires subtracting both cached reads and writes from the base input tokens.
The fix subtracts `cache_write_tokens` from the `input_tokens` when building the `priority` or `flex` input bucket. This aligns with OpenAI's documented formula: non-cached input tokens = total input tokens - cached reads - cache writes. Both the Responses and Chat Completions converters are updated to avoid inconsistent metadata. Existing tests for priority and flex should be updated to reflect the corrected values.
Edge Case Audit
This change alters the `input_token_details` mapping for priority/flex. Users relying on the previous `priority`/`flex` value for cost estimation will see lower numbers, which is correct. Ensure `cached_tokens` and `cache_write_tokens` are present (default to 0) and never exceed `input_tokens` to avoid negative values. The Chat Completions converter may have a different function name in older versions; verify and apply the same subtraction. Rollback: revert the subtraction line to `input_tokens - cached_tokens` if unexpected cost calculation issues arise. Additionally, update unit tests to assert the corrected values, as existing tests may fail.