AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

Langchain V3 Stream_events Final Response Tokens Buffered Instead Of Streamed

When using astream_events version='v3' with create_agent, the final response from the top-level agent is not streamed token-by-token; all tokens arrive at once with identical timestamps, while intermediate outputs stream correctly. This impacts real-time streaming applications.

highConfidence 78%LangchainAffected V1.3.11

Origin Analysis

In LangGraph v3 streaming protocol, the final model node after tool execution appears to buffer its output tokens until the node completes, instead of emitting them incrementally through the messages stream. This is likely a bug in the v3 event emission logic when an agent transitions from tool execution to final answer generation. Additionally, the synchronous `weather_agent.invoke` inside the tool may block the event loop, contributing to delayed flushing.
Run the provided code with DeepSeek model and tool delegation. Observe that thinking, tool calls, and sub-agent messages stream token-by-token, but the final top-level agent response appears all at once with identical timestamps after a delay.

Fixing Code Block

Edge Case Audit

This fallback loses structured v3 features like subagent streaming and may require manual handling of tool events. The async tool change requires that all internal operations be truly async; otherwise blocking persists. On Windows with Python 3.14, asyncio event loop policies may interact differently. If the DeepSeek endpoint does not support token-level streaming, neither v2 nor v3 will produce real-time tokens. Test thoroughly and retain the original v3 code for when the upstream bug is fixed. Rollback: revert to original call_weather (sync) and run() (v3) if the fallback introduces regressions.

Ecosystem Topology