AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

Count_tokens_approximately Over-Counts Audio/Video/File Blocks By Char-Counting Base64 Payload

The `count_tokens_approximately` function in langchain-core incorrectly counts base64 payloads of non-image data blocks (audio, video, file, text_plain) as characters, leading to massive token overestimation and potentially breaking trim_messages and token budgeting.

highConfidence 95%Langchain

Origin Analysis

The function handles only 'image' and 'image_url' block types with a fixed per-image token penalty; all other block types fall through to `repr(block)` and are char-counted, treating base64 data as raw text. This design fail to generalize a fixed penalty to all non-text data blocks.
1. Create a HumanMessage with a content list containing a non-text block (e.g., audio, video, file) with a large base64 payload. 2. Call `count_tokens_approximately([msg])`. 3. Observe that the returned token count is proportional to the length of the base64 string (e.g., 10000 chars -> ~2500 tokens) instead of a small fixed penalty.

Fixing Code Block

Edge Case Audit

Using a fixed token count (1100) is an approximation and may not match actual model token usage for all data types. If users relied on previous overestimation for conservative budgeting, the new lower estimates could lead to underestimation, potentially causing context overflow. Recommend making the constant configurable and documenting the behavior change. Rollback to previous version would restore old overcounting. No concurrency or thread-safety issues.

Ecosystem Topology