XMLOutputParser Incorrectly Strips Valid Element Tags Containing Encoding Attribute
XMLOutputParser.parse() raises OutputParserException for valid XML where a non-declaration element has an encoding attribute and the opening tag is followed by a newline. The regex used to remove the XML declaration is too broad and matches such element tags, causing the parser to remove the opening tag and produce malformed XML.
The regex in libs/core/langchain_core/output_parsers/xml.py used to strip the XML declaration matches any tag containing an 'encoding' attribute (e.g., <root encoding="utf-8">), not just the actual XML declaration (<?xml ...?>). When such a tag appears at the start and is followed by a newline, the regex removes it entirely, leaving the parser with malformed XML like '<item>x</item>\n</root>'.
Run the following Python code:
```python
from langchain_core.output_parsers.xml import XMLOutputParser
from langchain_core.exceptions import OutputParserException
xml = '<root encoding="utf-8">\n<item>x</item>\n</root>'
for backend in ("xml", "defusedxml"):
parser = XMLOutputParser(parser=backend)
print("Backend:", backend)
print("Streaming:", list(parser.transform(iter([xml]))))
try:
print("Parse:", parser.parse(xml))
except OutputParserException as exc:
print("Parse failed:", type(exc).__name__)
print("Input passed to parser:", repr(exc.llm_output))
```
Observe that parse() fails while transform() succeeds.
Fixing Code Block
def _remove_xml_declaration(xml: str) -> str:
"""Remove only the XML declaration, not element tags with encoding attributes."""
import re
return re.sub(r"<\?xml.*?\?>", "", xml, flags=re.DOTALL)
# In the parse() method (or wherever the XML is preprocessed), replace the existing regex with the above function call:
# xml = _remove_xml_declaration(xml)
The new regex anchors on the literal '<?xml' and uses non-greedy matching with DOTALL to capture the entire declaration up to the first '?>'. This ensures only the actual XML declaration is removed, leaving element tags like <root encoding="utf-8"> intact.
Edge Case Audit
This fix is safe for most cases, but if the XML contains '<?xml' inside a CDATA section or comment, the regex might incorrectly remove that content. However, such occurrences are extremely rare and the XML standard prohibits '<?xml' outside the declaration in prolog. Rollback advice: if unexpected parsing issues arise after applying this fix, revert to the original regex and consider a more robust XML parsing approach (e.g., using an XML parser's prolog handling).