AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

XMLOutputParser Incorrectly Strips Valid Element Tags Containing Encoding Attribute

XMLOutputParser.parse() raises OutputParserException for valid XML where a non-declaration element has an encoding attribute and the opening tag is followed by a newline. The regex used to remove the XML declaration is too broad and matches such element tags, causing the parser to remove the opening tag and produce malformed XML.

mediumConfidence 95%Langchain-CoreAffected V1.6.1

Origin Analysis

The regex in libs/core/langchain_core/output_parsers/xml.py used to strip the XML declaration matches any tag containing an 'encoding' attribute (e.g., <root encoding="utf-8">), not just the actual XML declaration (<?xml ...?>). When such a tag appears at the start and is followed by a newline, the regex removes it entirely, leaving the parser with malformed XML like '<item>x</item>\n</root>'.
Run the following Python code: ```python from langchain_core.output_parsers.xml import XMLOutputParser from langchain_core.exceptions import OutputParserException xml = '<root encoding="utf-8">\n<item>x</item>\n</root>' for backend in ("xml", "defusedxml"): parser = XMLOutputParser(parser=backend) print("Backend:", backend) print("Streaming:", list(parser.transform(iter([xml])))) try: print("Parse:", parser.parse(xml)) except OutputParserException as exc: print("Parse failed:", type(exc).__name__) print("Input passed to parser:", repr(exc.llm_output)) ``` Observe that parse() fails while transform() succeeds.

Fixing Code Block

def _remove_xml_declaration(xml: str) -> str: """Remove only the XML declaration, not element tags with encoding attributes.""" import re return re.sub(r"<\?xml.*?\?>", "", xml, flags=re.DOTALL) # In the parse() method (or wherever the XML is preprocessed), replace the existing regex with the above function call: # xml = _remove_xml_declaration(xml)
The new regex anchors on the literal '<?xml' and uses non-greedy matching with DOTALL to capture the entire declaration up to the first '?>'. This ensures only the actual XML declaration is removed, leaving element tags like <root encoding="utf-8"> intact.

Edge Case Audit

This fix is safe for most cases, but if the XML contains '<?xml' inside a CDATA section or comment, the regex might incorrectly remove that content. However, such occurrences are extremely rare and the XML standard prohibits '<?xml' outside the declaration in prolog. Rollback advice: if unexpected parsing issues arise after applying this fix, revert to the original regex and consider a more robust XML parsing approach (e.g., using an XML parser's prolog handling).

Ecosystem Topology