HTMLSemanticPreservingSplitter Preserves Empty And Javascript: Href Links As Markdown
The HTMLSemanticPreservingSplitter with preserve_links=True does not filter out empty or unsafe (e.g., javascript:) href attributes, resulting in invalid markdown links and potential XSS vectors.
In the preserve_links processing, the splitter directly uses the href attribute without validating it, leading to generation of markdown links with empty targets or unsafe protocols.
1. Instantiate HTMLSemanticPreservingSplitter with headers_to_split_on and preserve_links=True.
2. Pass HTML containing <a href="">, <a href="javascript:void(0)">, and <a href="https://example.com">.
3. Observe that output includes [Empty Link]() and [Bad Link](javascript:void(0)).
Fixing Code Block
from langchain_text_splitters.html import HTMLSemanticPreservingSplitter
from bs4 import BeautifulSoup
import re
def _safe_link_replacement(a_tag):
href = a_tag.get('href')
text = a_tag.get_text(strip=True)
if not href or not href.strip():
return text
href = href.strip()
unsafe_protocols = re.compile(r'^\s*(javascript|data|vbscript):', re.IGNORECASE)
if unsafe_protocols.match(href):
return text
return f'[{text}]({href})'
original_split_text = HTMLSemanticPreservingSplitter.split_text
def patched_split_text(self, text: str):
if self.preserve_links:
soup = BeautifulSoup(text, 'html.parser')
for a_tag in soup.find_all('a'):
replacement = _safe_link_replacement(a_tag)
a_tag.replace_with(replacement)
text = str(soup)
return original_split_text(self, text)
HTMLSemanticPreservingSplitter.split_text = patched_split_text
The patch pre-processes the HTML before the original split_text method runs. It parses with BeautifulSoup, iterates over <a> tags, and replaces each with either plain text (for empty or unsafe hrefs) or a valid markdown link. Then the original split_text is called with the cleaned HTML, preventing unsafe links from being preserved.
Edge Case Audit
This is a monkey patch that relies on the internal structure of HTMLSemanticPreservingSplitter and may break in future versions if the class implementation changes. It filters javascript:, data:, and vbscript: protocols; other potentially dangerous schemes (e.g., file:) are not covered and should be added if needed. The patch temporarily modifies the class method and is not thread-safe if multiple threads use the splitter while patching; apply it once at startup before any concurrent usage. Rollback: restore the original method via HTMLSemanticPreservingSplitter.split_text = original_split_text, or restart the process. For production, prefer an upstream fix.