AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

HTMLSemanticPreservingSplitter Reorders Inline Text Due To Incorrect Text Extraction Order

Inline tags like <strong> cause text reordering in split chunks, e.g., 'should' is moved before preceding text, breaking semantic coherence.

highConfidence 88%LangchainAffected V1.1.2

Origin Analysis

A recent change (PR #34587) to text extraction in HTMLSemanticPreservingSplitter concatenates child inline element text before the parent element's preceding text, violating document order. The function likely builds text as child.text + parent.text + child.tail instead of parent.text + child.text + child.tail (or equivalently using itertext()).
Use the provided Python snippet: parse HTML with <h1> and <p> containing <strong>, split with max_chunk_size=50, overlap=5; observe the first chunk is 'should This is some long text that be split into' instead of 'This is some long text that should be split into'.

Fixing Code Block

Edge Case Audit

This hotfix relies on lxml being installed (it is a required dependency for langchain-text-splitters). The monkey-patch fallback may inadvertently patch unrelated functions if the internal naming differs; test with nested inline tags (<a>, <em>, <span>) and ensure no regressions in header metadata extraction. Rollback by removing the patch and restoring the original module state. For production, prefer a proper upstream fix and version pinning.

Ecosystem Topology