AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

Agent Factory `_Execute_model_async` Uses Non-Streaming `Ainvoke`, Preventing `On_llm_new_token` Callbacks

The agent factory's async model execution path uses `await model_.ainvoke(messages)`, which returns a complete response without emitting token-level callbacks. LangGraph's `use_astream=True` only controls graph-level streaming, so model token streaming via `on_llm_new_token` is never triggered. This breaks real-time token streaming for agents.

highConfidence 75%LangChainAffected V1.3.4

Origin Analysis

In `langchain/agents/factory.py`, the `_execute_model_async` method calls `await model_.ainvoke(messages)` instead of using streaming `astream`. Non-streaming invocation aggregates the full model response internally and never invokes the `on_llm_new_token` callback. LangGraph's `use_astream=True` streams graph events but does not change how the model node calls the LLM, so the model still returns a complete `AIMessage`.
1. Use langchain==1.3.4 and langgraph==1.2.4. 2. Create an agent with a callback handler that implements `on_llm_new_token`. 3. Run the agent with `astream_events` or graph streaming (e.g., `use_astream=True`). 4. Observe that the `on_llm_new_token` callback is never invoked, even though graph-level events are streamed.

Fixing Code Block

Edge Case Audit

Streaming aggregation may introduce overhead and memory usage proportional to response length. If the model does not support streaming or returns incomplete tool call chunks, the merge logic must handle `tool_call_chunks` to avoid losing tool calls. This change may alter timing and concurrency behavior; ensure `config` is not None when callbacks are required. Rollback: revert to the original `ainvoke` call if streaming causes regressions or tool call parsing issues.

Ecosystem Topology