AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

CommaSeparatedListOutputParser Corrupts Quoted CSV Fields Across Streaming Chunks

Streaming CommaSeparatedListOutputParser produces incorrect fields when a quoted CSV field containing a comma is split across chunk boundaries. Non-streaming parsing is correct, but transform/atransform silently corrupts data.

highConfidence 85%Langchain-CoreAffected V1.6.2

Origin Analysis

ListOutputParser's generic streaming fallback parses the partial buffer, emits all but the last field, and stores the decoded last field as the next buffer. This loses the raw opening quote and quote state, so subsequent chunks are parsed with incorrect CSV field boundaries.
```python from langchain_core.output_parsers import CommaSeparatedListOutputParser parser = CommaSeparatedListOutputParser() chunks = ['alpha, "beta,', ' gamma", delta'] print(list(parser.transform(iter(chunks)))) ``` Expected: [['alpha'], ['beta, gamma'], ['delta']] Actual: [['alpha'], ['beta'], ['gamma"'], ['delta']]

Fixing Code Block

Edge Case Audit

This subclass-based hotfix must be used in place of the original CommaSeparatedListOutputParser. It assumes the parser uses csv.reader semantics with skipinitialspace=True and doublequote=True. If those settings change, _last_delimiter may need adjustment. The fix is not process-wide unless you monkeypatch langchain_core.output_parsers.CommaSeparatedListOutputParser. Rollback is straightforward: remove/replace the subclass usage. Concurrency is safe because state is local to the generator/async generator.

Ecosystem Topology