Reasoning_content Is Not Streamed When Response_format=Json_object (Aggregated Into A Single Chunk)
ChatDeepSeek with response_format=json_object streams reasoning_content as one aggregated chunk because langchain-openai uses the beta streaming helper when response_format is present, which buffers the full response.
langchain_openai.ChatOpenAI._stream/_astream detects response_format in the payload and switches to self.client.beta.chat.completions.stream, which accumulates all chunks before yielding a final chunk; therefore reasoning_content is not streamed incrementally. The underlying design flaw is that the beta helper does not provide true token-level streaming for reasoning_content when response_format is set.
Instantiate ChatDeepSeek with model='deepseek-v4-flash', thinking enabled, and bind response_format={'type': 'json_object'}; call astream with a system and human message; count chunks containing reasoning_content in additional_kwargs. Observed: 1 aggregated chunk instead of many incremental chunks.
Fixing Code Block
diff --git a/libs/partners/openai/langchain_openai/chat_models/base.py b/libs/partners/openai/langchain_openai/chat_models/base.py
--- a/libs/partners/openai/langchain_openai/chat_models/base.py
+++ b/libs/partners/openai/langchain_openai/chat_models/base.py
@@ -XXXX,XX +XXXX,XX @@ class ChatOpenAI(BaseChatModel):
if "response_format" in payload:
- response = self.client.beta.chat.completions.stream(**payload)
+ response = self.client.chat.completions.create(**payload, stream=True)
else:
response = self.client.chat.completions.create(**payload, stream=True)
@@ -XXXX,XX +XXXX,XX @@ class ChatOpenAI(BaseChatModel):
if "response_format" in payload:
- response = await self.async_client.beta.chat.completions.stream(**payload)
+ response = await self.async_client.chat.completions.create(**payload, stream=True)
else:
response = await self.async_client.chat.completions.create(**payload, stream=True)
Replace the beta streaming branch with the standard streaming endpoint, which the DeepSeek API supports for response_format. This ensures token-by-token reasoning_content emission as verified by the native SDK. For async, use the async client's standard create method.
Edge Case Audit
Applying this patch to langchain-openai affects all OpenAI-compatible models using response_format with streaming. Some providers or older API versions may still require the beta endpoint for structured outputs; if such models break, rollback by reverting to the beta branch or introducing a provider-specific flag. Test with OpenAI and DeepSeek explicitly, and validate that response_format still enforces JSON schema. Also consider concurrency: the standard endpoint may have different rate limits than the beta endpoint.