AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

Text-Splitters: `MarkdownHeaderTextSplitter` Keeps The Closing `##` Sequence In Header Metadata

MarkdownHeaderTextSplitter incorrectly includes CommonMark optional closing hash sequences in header metadata values, e.g. '# Title ##' yields metadata 'Title ##' instead of 'Title'. This breaks filtering and grouping by header.

mediumConfidence 92%Langchain-Text-SplittersAffected V1.1.2

Origin Analysis

The splitter extracts header text by taking everything after the opening marker and applying `.strip()`, without removing the optional closing sequence of `#` characters. According to CommonMark, a run of trailing `#` that is preceded by whitespace or constitutes the whole header is not part of the heading text.
```python from langchain_text_splitters import MarkdownHeaderTextSplitter splitter = MarkdownHeaderTextSplitter([("#", "Header 1"), ("##", "Header 2")]) for doc in splitter.split_text("# Title ##\nIntro.\n## Section ###\nBody."): print(doc.metadata) # {'Header 1': 'Title ##'} # {'Header 1': 'Title ##', 'Header 2': 'Section ###'} ```

Fixing Code Block

import re def _strip_closing_hashes(header: str) -> str: """Remove CommonMark optional closing hashes from a header.""" header = header.strip() if header and set(header) == {"#"}: return "" return re.sub(r"(?<=\s)#+\s*$", "", header).strip() # In MarkdownHeaderTextSplitter.aggregate_lines_to_chunks, replace: # header = line[len(sep):].strip() # with: header = _strip_closing_hashes(line[len(sep):])
A helper function first strips surrounding whitespace, then removes a trailing run of '#' only if it is preceded by whitespace (using a regex lookbehind) or if the entire header consists solely of '#' characters. This matches CommonMark ATX heading rules: '# Title ##' becomes 'Title', '# C#' stays 'C#', and '# ##' becomes an empty header.

Edge Case Audit

The regex uses a single-char lookbehind for whitespace, which is safe in Python 3 but may not cover all Unicode edge cases if the language version is very old. The fix should be released with tests for '# Title ##', '# Title ### ', '# C#', '# Title#', and '# ##'. If unexpected removal occurs (e.g., headers with trailing hashes that are actually part of the content), the extraction line should be reverted to `header = line[len(sep):].strip()` and the splitter version pinned until a better fix is available. No concurrency or rollback concerns beyond metadata changes.

Ecosystem Topology