Text-Splitters: `MarkdownHeaderTextSplitter` Keeps The Closing `##` Sequence In Header Metadata
MarkdownHeaderTextSplitter incorrectly includes CommonMark optional closing hash sequences in header metadata values, e.g. '# Title ##' yields metadata 'Title ##' instead of 'Title'. This breaks filtering and grouping by header.
The splitter extracts header text by taking everything after the opening marker and applying `.strip()`, without removing the optional closing sequence of `#` characters. According to CommonMark, a run of trailing `#` that is preceded by whitespace or constitutes the whole header is not part of the heading text.
```python
from langchain_text_splitters import MarkdownHeaderTextSplitter
splitter = MarkdownHeaderTextSplitter([("#", "Header 1"), ("##", "Header 2")])
for doc in splitter.split_text("# Title ##\nIntro.\n## Section ###\nBody."):
print(doc.metadata)
# {'Header 1': 'Title ##'}
# {'Header 1': 'Title ##', 'Header 2': 'Section ###'}
```
Fixing Code Block
import re
def _strip_closing_hashes(header: str) -> str:
"""Remove CommonMark optional closing hashes from a header."""
header = header.strip()
if header and set(header) == {"#"}:
return ""
return re.sub(r"(?<=\s)#+\s*$", "", header).strip()
# In MarkdownHeaderTextSplitter.aggregate_lines_to_chunks, replace:
# header = line[len(sep):].strip()
# with:
header = _strip_closing_hashes(line[len(sep):])
A helper function first strips surrounding whitespace, then removes a trailing run of '#' only if it is preceded by whitespace (using a regex lookbehind) or if the entire header consists solely of '#' characters. This matches CommonMark ATX heading rules: '# Title ##' becomes 'Title', '# C#' stays 'C#', and '# ##' becomes an empty header.
Edge Case Audit
The regex uses a single-char lookbehind for whitespace, which is safe in Python 3 but may not cover all Unicode edge cases if the language version is very old. The fix should be released with tests for '# Title ##', '# Title ### ', '# C#', '# Title#', and '# ##'. If unexpected removal occurs (e.g., headers with trailing hashes that are actually part of the content), the extraction line should be reverted to `header = line[len(sep):].strip()` and the splitter version pinned until a better fix is available. No concurrency or rollback concerns beyond metadata changes.