AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

Inflated Token Usage Reporting For Fireworks Qwen3p6-Plus With Reasoning_effort='None'

The langchain-fireworks integration reports corrupted token counts (e.g., completion_tokens=1229292929) when using accounts/fireworks/models/qwen3p6-plus, breaking downstream cost tracking and database operations.

highConfidence 75%LangchainAffected Vlangchain-Fireworks>=1.3.1Affected Vlangchain>=1.2.15Affected Vlangchain-Core>=0.3.0Affected Vlangchain-Openai>=1.2.1

Origin Analysis

The Fireworks API returns invalid usage metadata for the qwen3p6-plus reasoning model, and the LangChain Fireworks adapter does not validate or sanitize these values before exposing them in AIMessage.usage_metadata and response_metadata.
1. Set FIREWORKS_API_KEY. 2. Initialize model with init_chat_model(model='accounts/fireworks/models/qwen3p6-plus', model_provider='fireworks', model_kwargs={'reasoning_effort':'none'}). 3. Invoke with a short query. 4. Inspect response.usage_metadata; observe inflated token counts.

Fixing Code Block

Edge Case Audit

This is a client-side workaround; it does not fix the upstream Fireworks API or LangChain integration. Token counts may differ from the provider's exact accounting if the model uses a different tokenizer. If the API later returns correct values, remove this wrapper to avoid inaccurate local counts. For streaming or multi-modal inputs, additional handling is required. Rollback by reverting to direct llm.invoke.

Ecosystem Topology