AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

`ProviderToolSearchMiddleware` Defers Tools As Flat Top-Level Functions Instead Of Wrapping Them In A `Namespace` — No Token Savings

The OpenAI Responses API excludes deferred tool schemas from billed input only when they are inside a `namespace` block. `ProviderToolSearchMiddleware` emits flat `function` tools with `defer_loading: true` plus `tool_search`, causing OpenAI to bill the full schemas and adding overhead, so token usage increases instead of decreasing.

highConfidence 85%LangchainAffected V1.3.11

Origin Analysis

The middleware appends `defer_loading: true` to each tool definition but does not wrap them in a `{"type": "namespace", ...}` block. OpenAI's Responses API interprets flat `defer_loading` functions as regular billed tools, while the added `tool_search` tool further inflates input token count.
1. Create a list of tool specifications. 2. Use `ProviderToolSearchMiddleware` with `defer_loading=True, tool_search=True`. 3. Capture the outbound `tools` payload: it will contain flat `function` objects with `defer_loading: true` and a separate `tool_search` object. 4. Send the same request with tools manually wrapped in a `namespace` block (and `tool_search`). 5. Compare `response.usage.input_tokens`: flat form shows +24% vs inline, namespace form shows -47%.

Fixing Code Block

Edge Case Audit

This is a global monkey patch; it affects all instances of `ProviderToolSearchMiddleware`. The hardcoded namespace name `deferred_tools` may collide with user-defined namespaces if the request already contains one. OpenAI models/endpoints that do not support the `namespace` tool type may reject the request. Rollback: remove the monkey patch and restart the process. For thread-safety, the patch is global, but it checks instance attributes (`defer_loading`, `tool_search`) before applying, so instances with different settings fall back to the original behavior.

Ecosystem Topology