AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

Reasoning_content Is Not Streamed When Response_format=Json_object (Aggregated Into A Single Chunk)

ChatDeepSeek with response_format=json_object streams reasoning_content as one aggregated chunk because langchain-openai uses the beta streaming helper when response_format is present, which buffers the full response.

mediumConfidence 90%LangchainAffected Vlangchain-Openai==1.4.0Affected Vlangchain-Deepseek==1.1.0

Origin Analysis

langchain_openai.ChatOpenAI._stream/_astream detects response_format in the payload and switches to self.client.beta.chat.completions.stream, which accumulates all chunks before yielding a final chunk; therefore reasoning_content is not streamed incrementally. The underlying design flaw is that the beta helper does not provide true token-level streaming for reasoning_content when response_format is set.
Instantiate ChatDeepSeek with model='deepseek-v4-flash', thinking enabled, and bind response_format={'type': 'json_object'}; call astream with a system and human message; count chunks containing reasoning_content in additional_kwargs. Observed: 1 aggregated chunk instead of many incremental chunks.

Fixing Code Block

diff --git a/libs/partners/openai/langchain_openai/chat_models/base.py b/libs/partners/openai/langchain_openai/chat_models/base.py --- a/libs/partners/openai/langchain_openai/chat_models/base.py +++ b/libs/partners/openai/langchain_openai/chat_models/base.py @@ -XXXX,XX +XXXX,XX @@ class ChatOpenAI(BaseChatModel): if "response_format" in payload: - response = self.client.beta.chat.completions.stream(**payload) + response = self.client.chat.completions.create(**payload, stream=True) else: response = self.client.chat.completions.create(**payload, stream=True) @@ -XXXX,XX +XXXX,XX @@ class ChatOpenAI(BaseChatModel): if "response_format" in payload: - response = await self.async_client.beta.chat.completions.stream(**payload) + response = await self.async_client.chat.completions.create(**payload, stream=True) else: response = await self.async_client.chat.completions.create(**payload, stream=True)
Replace the beta streaming branch with the standard streaming endpoint, which the DeepSeek API supports for response_format. This ensures token-by-token reasoning_content emission as verified by the native SDK. For async, use the async client's standard create method.

Edge Case Audit

Applying this patch to langchain-openai affects all OpenAI-compatible models using response_format with streaming. Some providers or older API versions may still require the beta endpoint for structured outputs; if such models break, rollback by reverting to the beta branch or introducing a provider-specific flag. Test with OpenAI and DeepSeek explicitly, and validate that response_format still enforces JSON schema. Also consider concurrency: the standard endpoint may have different rate limits than the beta endpoint.

Ecosystem Topology