AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

Surface LangSmith Gateway Metadata On Successful Responses In ChatFireworks

ChatFireworks does not parse the X-LangSmith-Gateway-Metadata header on successful responses when routed through the LangSmith LLM Gateway, unlike ChatOpenAI and ChatAnthropic. This causes traces to miss gateway info such as the actually resolved model and provider, leading to incorrect cost tracking and inconsistent behavior across integrations.

mediumConfidence 92%LangChain

Origin Analysis

ChatFireworks calls `chat.completions.create` directly, which returns only the parsed body without HTTP headers. The LangSmith gateway sends metadata in the `X-LangSmith-Gateway-Metadata` header on successful responses, but this header is discarded because the raw response is not accessed.
1. Set `LANGSMITH_GATEWAY=true` and configure a gateway with Fireworks model routing (including fallbacks). 2. Use `ChatFireworks` to make a successful call. 3. Check the LangSmith trace: observe that `ls_gateway_info` is missing, and `ls_model_name` / `ls_provider` reflect the requested model rather than the actual model resolved by the gateway. 4. Compare with `ChatOpenAI` or `ChatAnthropic` under the same setup, where the metadata is present.

Fixing Code Block

from typing import Any, Mapping, Optional from langchain_core.utils._gateway import _parse_gateway_metadata from langchain_core.outputs import ChatGeneration, ChatGenerationChunk # Add these methods and modifications to ChatFireworks class @property def _uses_gateway(self) -> bool: return self.fireworks_api_key is not None and self.fireworks_api_key.startswith("lsv2_") def _add_gateway_metadata(self, generation_info: dict, headers: Mapping[str, str]) -> None: metadata = _parse_gateway_metadata(headers) if metadata: generation_info[GATEWAY_METADATA_RESPONSE_KEY] = metadata def _completion_with_retry(self, **kwargs: Any) -> tuple[Any, Optional[Mapping[str, str]]]: if self._uses_gateway: raw_response = self.client.with_raw_response.create(**kwargs) response = raw_response.parse() headers = getattr(raw_response, "headers", {}) return response, headers else: return self.client.create(**kwargs), None async def _acompletion_with_retry(self, **kwargs: Any) -> tuple[Any, Optional[Mapping[str, str]]]: if self._uses_gateway: raw_response = await self.async_client.with_raw_response.create(**kwargs) response = raw_response.parse() headers = getattr(raw_response, "headers", {}) return response, headers else: return await self.async_client.create(**kwargs), None def _generate(self, messages, stop=None, run_manager=None, **kwargs): params = self._get_invocation_params(stop=stop, **kwargs) response, headers = self._completion_with_retry(**params) chat_result = self._create_chat_result(response) if headers: for generation in chat_result.generations: self._add_gateway_metadata(generation.generation_info or {}, headers) return chat_result async def _agenerate(self, messages, stop=None, run_manager=None, **kwargs): params = self._get_invocation_params(stop=stop, **kwargs) response, headers = await self._acompletion_with_retry(**params) chat_result = self._create_chat_result(response) if headers: for generation in chat_result.generations: self._add_gateway_metadata(generation.generation_info or {}, headers) return chat_result def _stream(self, messages, stop=None, run_manager=None, **kwargs): params = self._get_invocation_params(stop=stop, **kwargs) response, headers = self._completion_with_retry(stream=True, **params) for chunk in response: generation_info = {} if headers: self._add_gateway_metadata(generation_info, headers) yield ChatGenerationChunk(message=AIMessageChunk(content=chunk.choices[0].delta.content if chunk.choices else ""), generation_info=generation_info) async def _astream(self, messages, stop=None, run_manager=None, **kwargs): params = self._get_invocation_params(stop=stop, **kwargs) response, headers = await self._acompletion_with_retry(stream=True, **params) async for chunk in response: generation_info = {} if headers: self._add_gateway_metadata(generation_info, headers) yield ChatGenerationChunk(message=AIMessageChunk(content=chunk.choices[0].delta.content if chunk.choices else ""), generation_info=generation_info) # Ensure to import GATEWAY_METADATA_RESPONSE_KEY from langchain_core.utils._gateway if not already. # Note: The above code assumes existing imports and context; adjust as necessary.
The fix introduces a `_uses_gateway` property to detect if the API key is a LangSmith gateway key (prefix `lsv2_`). When gateway is used, the completion methods switch to `with_raw_response.create`, which returns a raw response object containing headers. The headers are captured via `getattr` and the response body is parsed with `.parse()` to maintain the same return type as before. The gateway metadata is extracted using the shared `_parse_gateway_metadata` helper and attached to each generation's `generation_info` for non-streaming, and to the first chunk for streaming, mirroring the behavior in ChatOpenAI and ChatAnthropic.

Edge Case Audit

1. **Custom clients**: If a user passes a custom client but also uses an API key starting with `lsv2_`, the code will attempt `with_raw_response` which may not be supported by that custom client, causing errors. Mitigation: check for `hasattr(client, 'with_raw_response')` before using it, or allow an explicit opt-out. 2. **SDK version compatibility**: The feature relies on `with_raw_response` and `parse` methods, which are available in fireworks-ai>=1.2.0a71. Older SDK versions may not have these and will break. Pin a minimum version. 3. **Streaming metadata**: Metadata is only attached to the first chunk. If the first chunk has no choices (e.g., an empty delta), metadata might still be attached but could be missed by downstream consumers. Consider attaching to all chunks or using a wrapper stream. 4. **Async client**: Ensure `async_client` is correctly implemented and used; otherwise async gateway calls may still lose metadata. 5. **Rollback**: If issues arise, rollback by reverting to the original direct `create` call for all cases (simply remove the gateway-specific branch).

Ecosystem Topology