AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

MarkdownListOutputParser Silently Drops '+' Bullet Items

The MarkdownListOutputParser in langchain-core uses a regex that only matches '-' and '*' as bullet markers, omitting '+'. As a result, parsing a Markdown list with '+' bullets returns an empty list instead of the expected items, causing silent data loss.

highConfidence 97%LangchainAffected V<=1.6.1

Origin Analysis

The regular expression r"^\s*[-*]\s([^\n]+)$" in MarkdownListOutputParser's pattern does not include '+' in its character class, despite '+' being a valid CommonMark bullet list marker.
1. Install langchain-core (e.g., version 1.6.1). 2. Execute the following Python code: from langchain_core.output_parsers import MarkdownListOutputParser parser = MarkdownListOutputParser() print(parser.parse("+ a\n+ b")) 3. Observe that the output is [] instead of ['a', 'b'].

Fixing Code Block

Edge Case Audit

This change is backward compatible; it only expands the set of recognized bullet markers. No existing parsing behavior is broken. However, if the parser is used in environments where '+' was intentionally discouraged, this may now accept previously ignored lines. Rollback is trivial: revert the pattern to r"^\s*[-*]\s([^\n]+)$".

Ecosystem Topology