AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

MarkdownHeaderTextSplitter Silently Drops Tabs And NBSP (Including Inside Fenced Code)

MarkdownHeaderTextSplitter corrupts chunk content by applying line.strip() and str.isprintable() filtering, which removes tabs, NBSP, and leading whitespace inside fenced code blocks, causing silent data loss during text splitting.

highConfidence 85%Langchain

Origin Analysis

The splitter treats every line with line.strip() without tracking fenced code blocks, then filters characters using str.isprintable(), which returns False for tab (\t) and NBSP (\u00a0). This strips legitimately significant whitespace, including indentation inside code fences, while preserving only a blanket non-printable removal.
Run the following Python code: from langchain_text_splitters import MarkdownHeaderTextSplitter s = MarkdownHeaderTextSplitter(headers_to_split_on=[("#", "h1")]) docs = s.split_text("# Title\n\n\ndef foo():\n\treturn 1\n\n") print(repr(docs[0].page_content)) # tab indent destroyed docs2 = s.split_text("# Title\n\nhello\tworld") print(repr(docs2[0].page_content)) # 'helloworld' docs3 = s.split_text("# Title\n\nhello\u00a0world") print(repr(docs3[0].page_content)) # 'helloworld'

Fixing Code Block

Edge Case Audit

This change may break downstream code that expected tabs and NBSP to be removed (unlikely but possible). The naive fence detection toggles on any line whose stripped form starts with ```, which can be confused by inline triple backticks or multiple fences in a block. The fix does not address header detection inside code fences; lines containing '#' inside a fenced block may still be split as headers. Rollback is safe by reverting to the previous split_text implementation, but tests should be run to ensure no regression in existing metadata construction or return_each_line behavior.

Ecosystem Topology