AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

XMLOutputParser Streaming Drops Valid Unicode Root Tags

Streaming with XMLOutputParser silently loses data when the root XML element uses a valid Unicode name (e.g., Chinese tags). The parser's start-tag detection regex only matches ASCII letters, colon, and underscore, causing it to skip Unicode roots and either return empty output or omit the root hierarchy.

highConfidence 95%Langchain-CoreAffected V1.6.6Affected Vmaster

Origin Analysis

The _StreamingParser in langchain_core/output_parsers/xml.py uses the regex r"<[a-zA-Z:_]" to locate the first XML element. This regex does not account for Unicode NameStartChar characters defined in XML 1.0, so roots like <答案> are ignored. If the root is Unicode but a child is ASCII, the parser starts at the child, losing the root; if both root and child are Unicode, no start tag is ever matched and streaming yields nothing.
Run the provided Python script with langchain-core==1.6.6. It creates a FakeStreamingListLLM that returns fixed XML strings and pipes them through XMLOutputParser. Compare invoke vs stream/astream for three cases: <答案><选项>蓝色</选项></答案>, <答案><choice>blue</choice></答案>, and <answer><选项>蓝色</选项></answer>. The first returns empty for stream/astream, the second returns only the child without root, the third works correctly.

Fixing Code Block

Edge Case Audit

This fix broadens the start-tag detection to include all Unicode word characters. While unlikely, some Unicode symbols or characters that are valid as word characters but not valid XML NameStartChar could be incorrectly matched, leading to attempts to parse non-XML content. Rollback suggestion: revert to the original regex r"<[a-zA-Z:_]" if the new behavior causes regressions, but this will reintroduce the Unicode root loss. Additionally, this fix only addresses start detection; other streaming logic may still assume ASCII tag names, so thorough testing with mixed Unicode/ASCII tags is advised.

Ecosystem Topology