AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

Add First-Class Retry Configuration To ChatOllama

ChatOllama lacks a constructor-level max_retries field. Although Pydantic accepts the extra keyword, it is not retained as a model field and is not used to configure retry behavior in the underlying Ollama HTTP client. This prevents provider-agnostic retry configuration from being applied to Ollama requests.

mediumConfidence 65%LangChain

Origin Analysis

The ChatOllama model does not define max_retries as a field. Client construction in _client and _async_client creates Ollama Client/AsyncClient instances from client_kwargs without any retry transport configuration. The user-provided max_retries argument is silently swallowed as extra data and never influences HTTP retries.
1. Install langchain-ollama and instantiate ChatOllama(model='llama3', max_retries=3). 2. Check ChatOllama.model_fields: 'max_retries' is absent. 3. Check getattr(llm, 'max_retries', None): returns None. 4. Inspect the underlying HTTP client transport for the Ollama client; no retry policy is configured.

Fixing Code Block

from typing import Optional import httpx from ollama import AsyncClient, Client class ChatOllama(BaseChatModel): # Add this field to the existing model fields max_retries: Optional[int] = None @property def _client(self) -> Client: if self._sync_client is None: kwargs = self.client_kwargs or {} sync_kwargs = self.sync_client_kwargs or {} if self.max_retries is not None: transport = httpx.HTTPTransport(retries=self.max_retries) client = httpx.Client(transport=transport, **sync_kwargs) self._sync_client = Client(host=self.base_url, client=client, **kwargs) else: self._sync_client = Client(host=self.base_url, **kwargs) return self._sync_client @property def _async_client(self) -> AsyncClient: if self._async_client is None: kwargs = self.client_kwargs or {} async_kwargs = self.async_client_kwargs or {} if self.max_retries is not None: transport = httpx.AsyncHTTPTransport(retries=self.max_retries) client = httpx.AsyncClient(transport=transport, **async_kwargs) self._async_client = AsyncClient(host=self.base_url, client=client, **kwargs) else: self._async_client = AsyncClient(host=self.base_url, **kwargs) return self._async_client
Add max_retries as an optional integer field. When set, construct httpx.Client and httpx.AsyncClient with HTTPTransport/AsyncHTTPTransport configured with the requested retry count and pass those clients to the Ollama Client and AsyncClient constructors. This applies retry behavior at the SDK HTTP layer while preserving existing client_kwargs and host configuration.

Edge Case Audit

Retrying non-idempotent Ollama API calls (e.g., chat generations) can cause duplicate completions if the request succeeds on the server but the response is lost. Users must ensure that application logic handles potential duplicate generations. Additionally, if sync_client_kwargs or async_client_kwargs already contains a custom transport, this fix may conflict by overriding it. Concurrency is safe as long as each ChatOllama instance has its own clients; sharing instances across threads may lead to retry counters affecting all threads. Rollback: set max_retries=None to restore previous behavior.

Ecosystem Topology