AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

ContextEditingMiddleware Blocks The Event Loop With Model-Based Token Counting

When ContextEditingMiddleware uses token_count_method="model" through the async method awrap_model_call, it directly invokes the synchronous model.get_num_tokens_from_messages on the event-loop thread. For models like ChatAnthropic that perform a synchronous HTTP request for token counting, this blocks the entire event loop for the duration of each count, causing severe latency spikes and starving other async tasks.

highConfidence 85%LangchainAffected V1.4.0

Origin Analysis

The async middleware path calls the synchronous token-counting method directly instead of offloading it to a worker thread, so blocking I/O or CPU work in get_num_tokens_from_messages freezes the event loop.
1. Install langchain 1.4.0, langchain-core 1.6.2, langgraph 1.2.11, and other dependencies as in the issue.\n2. Create a custom FakeChatModel whose get_num_tokens_from_messages sleeps for 0.15 seconds to simulate a remote token-counting API.\n3. Build an async main() that creates several AIMessage/ToolMessage pairs and sets up a heartbeat asyncio task that expects 10 ms intervals.\n4. Instantiate ContextEditingMiddleware with edits=[ClearToolUsesEdit(trigger=0, clear_at_least=10_000, keep=3)] and token_count_method="model".\n5. Call await middleware.awrap_model_call(request, handler) and observe the maximum heartbeat gap. The reproduction shows a maximum gap of approximately 0.6 seconds instead of the expected 10 ms.

Fixing Code Block

Edge Case Audit

Using asyncio.to_thread relies on the default ThreadPoolExecutor, whose size is limited (typically min(32, os.cpu_count()+4)). Under very high concurrency this can cause thread pool exhaustion or queueing, which may reduce throughput or introduce latency. If the token-counting method is CPU-bound Python code, the GIL may still serialize execution globally even if the event loop is not blocked. This change may also alter the timing of token-counting calls relative to other async tasks. Rollback suggestion: if thread pool saturation or unexpected performance regressions occur, revert this change and either set token_count_method="length" or use a dedicated executor with a bounded queue. Additionally, test on Python versions 3.9+ where asyncio.to_thread is available; no other version-specific issues are expected.

Ecosystem Topology