AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

SummarizationMiddleware Fails To Identify Claude On Bedrock, Causing Late Summarization And Potential Context Overflow

The SummarizationMiddleware uses a hardcoded check on model._llm_type to apply Claude-specific token calibration. For ChatBedrockConverse (and ChatBedrock), _llm_type is 'amazon_bedrock_converse_chat', which does not match the expected 'anthropic-chat' prefix. As a result, it falls back to the generic chars_per_token=4.0 instead of Claude's calibrated 3.3, underestimating token counts by ~20%. This leads to summarization triggering too late and can cause model calls to fail with 'Input is too long for requested model'.

highConfidence 95%LangChainAffected V<=1.4.1

Origin Analysis

In langchain.agents.middleware.SummarizationMiddleware._get_approximate_token_counter, the logic checks if self.model._llm_type.startswith("anthropic-chat") to apply chars_per_token=3.3. Claude models served via Amazon Bedrock use _llm_type='amazon_bedrock_converse_chat' (or similar), failing this check and defaulting to 4.0. The design flaw is coupling token calibration to a specific provider-specific _llm_type instead of detecting the underlying model family (Claude).
Run the following Python snippet with langchain, langchain-anthropic, and langchain-aws installed:\n```python\nfrom langchain.agents.middleware import SummarizationMiddleware\nfrom langchain_anthropic import ChatAnthropic\nfrom langchain_aws import ChatBedrockConverse\n\nmodels = [\n ChatAnthropic(model=\"claude-sonnet-4-5\", api_key=\"unused\"),\n ChatBedrockConverse(\n model=\"us.anthropic.claude-sonnet-4-5-20250929-v1:0\", region_name=\"us-east-1\"\n ),\n]\nfor model in models:\n mw = SummarizationMiddleware(model=model, trigger=(\"tokens\", 170_000))\n print(f\"{type(model).__name__:22} _llm_type={model._llm_type!r:34} chars_per_token={mw.token_counter.keywords.get('chars_per_token', 4.0)}\")\n```\nExpected output shows ChatAnthropic gets 3.3 but ChatBedrockConverse gets 4.0.

Fixing Code Block

Edge Case Audit

The proposed fix relies on the model's _get_ls_params() method returning a dictionary with 'ls_model_name'. While this is generally stable in current LangChain versions, future changes to the internal LS params structure could break detection. Additionally, a model name containing 'claude' is assumed to be Anthropic Claude; this is safe today but could theoretically cause false positives if another vendor uses that string. No concurrency or multithreading issues are introduced as the method is stateless. To roll back, simply restore the original _llm_type-only check. Recommend adding unit tests for ChatAnthropic, ChatBedrock, ChatBedrockConverse, and a generic model to ensure correct calibration.

Ecosystem Topology