AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

MarkdownHeaderTextSplitter Fails To Aggregate Custom Header Patterns When Strip_headers=False

When using custom header patterns (e.g., **, ***) and strip_headers=False, nested custom headers are split into separate chunks instead of being merged into a single chunk like standard Markdown headers. This creates a header-only chunk for the parent header, which can degrade downstream processing.

mediumConfidence 87%Langchain

Origin Analysis

In aggregate_lines_to_chunks(), the preserve-headers merge branch checks whether the previous chunk's last line starts with '#' before merging. This check only recognizes standard Markdown ATX headers and ignores custom header patterns configured via headers_to_split_on or custom_header_patterns. As a result, when a deeper custom header follows a parent custom header, the code does not merge them and instead creates a new chunk for the child, leaving the parent as a header-only chunk.
1. Instantiate MarkdownHeaderTextSplitter with headers_to_split_on=[("**", "Header 1"), ("***", "Header 2")], custom_header_patterns={"**": 1, "***": 2}, and strip_headers=False. 2. Call split_text("**H1**\n***H2***\nbody\n"). 3. Observe output: two chunks: '**H1**' (metadata Header 1) and '***H2***\nbody' (metadata Header 1 and Header 2). 4. Compare with standard headers: headers_to_split_on=[("#", "Header 1"), ("##", "Header 2")], strip_headers=False, input "# H1\n## H2\nbody\n" yields one chunk '# H1 \n## H2\nbody' with both metadata.

Fixing Code Block

def _is_header_line(self, line: str) -> bool: """Return True if line starts with a configured header pattern.""" stripped = line.lstrip() for pattern, _ in self.headers_to_split_on: if stripped.startswith(pattern): return True for pattern in self.custom_header_patterns: if stripped.startswith(pattern): return True return False def aggregate_lines_to_chunks(self, lines: List[LineType]) -> List[Document]: """Combine lines into chunks.""" chunks: List[Document] = [] for line in lines: if len(chunks) > 0 and chunks[-1].metadata.get("_split") == "header": if line["metadata"]: last_line = chunks[-1].page_content.split("\n")[-1] if self._is_header_line(last_line): chunks[-1].page_content += " \n" + line["content"] chunks[-1].metadata.update(line["metadata"]) else: chunks[-1].page_content += " \n" chunks.append(Document(page_content="", metadata=line["metadata"])) else: chunks[-1].page_content += " \n" + line["content"] chunks[-1].metadata.update(line["metadata"]) else: chunks.append(Document(page_content=line["content"], metadata=line["metadata"])) return chunks
The fix replaces the hardcoded startswith('#') check with a helper function _is_header_line() that checks the last line against all configured header patterns from both headers_to_split_on and custom_header_patterns. This ensures that any recognized header pattern, not just Markdown ATX headers, triggers the merging logic in the preserve-headers branch. The rest of the aggregation logic remains unchanged.

Edge Case Audit

This change may cause merging for lines that happen to start with a configured custom pattern even if they are not intended as headers in a particular context (e.g., normal text beginning with '**'). Users should ensure their custom patterns are specific enough. If unintended merging occurs, rollback by reverting the condition to last_line.lstrip().startswith('#'). No thread-safety concerns as this method uses only local state. Test with existing standard header cases to confirm no regression.

Ecosystem Topology