AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

Output Parsers Drop Tool Calls When A Chat Model Falls Back To Invoke Inside Stream, So Structured Output Streams Nothing

BaseCumulativeTransformOutputParser converts incoming AIMessage to BaseMessageChunk using chunk.model_dump(), but BaseMessageChunk lacks tool_calls, causing tool calls to be dropped. This results in empty stream/astream output for non-streaming chat models (e.g., disable_streaming="tool_calling") even though invoke works correctly.

highConfidence 95%LangchainAffected V1.6.5

Origin Analysis

In langchain_core/output_parsers/transform.py, BaseCumulativeTransformOutputParser._transform and _atransform attempt to wrap a BaseMessage as BaseMessageChunk(**chunk.model_dump()). However, BaseMessageChunk does not have a tool_calls field, so any tool_calls present on an AIMessage are silently discarded before parsing. The correct approach is to convert to the specific chunk subclass (AIMessageChunk) that preserves tool_calls.
Run the provided Python code with langchain-core 1.6.5 and pydantic 2.13.5. The NonStreamingModel's _generate returns an AIMessage with tool_calls. Invoking the chain yields Person(name='Ada', age=36), but stream() and astream() yield empty lists.

Fixing Code Block

Edge Case Audit

This change only affects standard LangChain message types that have defined chunk counterparts. If custom message subclasses with additional fields are used, convert_to_chunk falls back to BaseMessageChunk, which may still drop those fields. Test thoroughly with custom message types and all relevant streaming configurations. The methods are stateless per iteration, so concurrency is safe as long as convert_to_chunk does not introduce shared state (it does not). To rollback, revert to the original BaseMessageChunk(**chunk.model_dump()) lines.

Ecosystem Topology