AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

HTMLSectionSplitter Leaks #TITLE# Metadata And Raises KeyError In Split_documents

HTMLSectionSplitter leaks its internal #TITLE# sentinel into Document metadata for content before the first configured header, and split_documents() raises KeyError: 'Title' when the parent Document metadata lacks a Title key.

highConfidence 85%LangchainAffected V1.1.2

Origin Analysis

The split_html_by_headers() method uses '#TITLE#' as a sentinel for pre-header content. split_text() then directly places this sentinel into the Document metadata. Later, split_documents() (or an overridden create_documents) attempts to replace this sentinel with metadata['Title'] unconditionally, causing a KeyError when the parent metadata does not contain 'Title'.
1. Create an HTMLSectionSplitter with headers_to_split_on=[('h1', 'Header 1')]. 2. Call split_text() on an HTML string containing a paragraph before the first <h1> tag. 3. Observe that the returned Document metadata includes {'Header 1': '#TITLE#'}. 4. Call split_documents() on a Document whose metadata is {'source': 'example'} (no 'Title' key). 5. Observe the raised KeyError: 'Title'.

Fixing Code Block

Edge Case Audit

This hotfix changes the metadata structure for pre-header content: previously it included {'Header 1': '#TITLE#'}, now it omits the key altogether (or sets it to empty string if replacement occurs). Users who relied on the presence of the key with any value may need to adjust. The merge order gives child header metadata precedence over parent metadata; if parent metadata contains a key matching a header name, it will be overridden. This matches typical expectations but could break if someone expected parent metadata to win. Rollback advice: revert to the original implementation if this causes unexpected metadata merging issues, but note that the original code has known bugs. A safer alternative is to wait for the official LangChain fix.

Ecosystem Topology