high95%
SummarizationMiddleware fails to identify Claude on Bedrock, causing late summarization and potential context overflow
The SummarizationMiddleware uses a hardcoded check on model._llm_type to apply Claude-specific token calibration. For ChatBedrockConverse (and ChatBedrock), _llm_type is 'amazon_bedrock_converse_chat', which does not match the expected 'anthropic-chat' prefix. As a result, it falls back to the generic chars_per_token=4.0 instead of Claude's calibrated 3.3, underestimating token counts by ~20%. This leads to summarization triggering too late and can cause model calls to fail with 'Input is too long for requested model'.
