AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

MarkdownHeaderTextSplitter Suffers O(N²) Slowdown When Merging Many Paragraphs Under The Same Header

aggregate_lines_to_chunks() repeatedly appends to a growing string, causing quadratic time complexity when many adjacent lines share the same metadata. The issue is observable with large inputs (e.g., 16,000 paragraphs take ~5.8 seconds) and becomes progressively worse.

highConfidence 95%LangchainAffected V1.1.2

Origin Analysis

In aggregate_lines_to_chunks(), the code uses `aggregated_chunks[-1]["content"] += " \n" + line["content"]` inside a loop. Each concatenation creates a new string and copies the entire accumulated content, leading to O(n²) time and memory traffic for n lines with identical metadata.
Install langchain-text-splitters==1.1.2. Run the provided benchmark: create text with N paragraphs separated by blank lines and split using MarkdownHeaderTextSplitter(headers_to_split_on=[("#", "Header 1")], strip_headers=False). Time splitting for N = 2000, 4000, 8000, 16000. Observed times: 0.02s, 0.27s, 1.27s, 5.7s respectively, demonstrating super-linear growth.

Fixing Code Block

Edge Case Audit

The fix introduces a temporary `content_parts` key in the chunk dictionaries. Ensure no downstream code expects only `content` and `metadata` keys in those intermediate dicts. If the method is ever overridden or externally referenced, behavior remains compatible. For rollback, replace this method with the original implementation, but re-run performance tests to confirm old behavior. No concurrency or multi-threading concerns are introduced. Test with empty lines, single lines, and mixed metadata to verify identical output.

Ecosystem Topology