AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

Count_tokens_approximately Underestimates Image Tokens With Fixed Constant

LangChain core's count_tokens_approximately assigns a fixed 85 tokens to every image regardless of resolution, causing severe underestimation for high-resolution images and leading to context window overflow failures.

highConfidence 95%Langchain-CoreAffected V1.4.0Affected V1.6.1

Origin Analysis

The function uses a constant tokens_per_image (85) for all image blocks, ignoring actual image dimensions. The optional usage_metadata scaling is clamped to at most 25% increase, which cannot correct large underestimations.
1. Create a PNG image of resolution 1024x1024 or higher. 2. Encode it as a data URL. 3. Construct a HumanMessage with content containing a text block and an image_url block. 4. Construct an AIMessage with known usage_metadata indicating the real input tokens (e.g., 1398). 5. Call count_tokens_approximately([human, ai], use_usage_metadata_scaling=True). 6. Observe that the returned count is around 103/129, far below the actual 1401, confirming the undercount.

Fixing Code Block

Edge Case Audit

This patch may not perfectly match all providers' token formulas (e.g., Anthropic's w*h/750). Remote URLs or file_id images still fall back to constant 85, which may undercount. The JPEG/WebP parsing could fail for some valid files, again falling back. If the new estimate causes overcounting in some edge cases, consider rolling back to the constant by reverting the image block change. Do not loosen the usage_metadata scaling clamp without extensive testing, as it could amplify errors for other message types.

Ecosystem Topology