ExperimentalMarkdownSyntaxTextSplitter Closes Code Blocks On Shorter Or Mismatched Fences
The splitter terminates fenced code blocks whenever it sees any line of three or more backticks or tildes, ignoring CommonMark rules that require a closing fence to match the opening character and have at least the opening length. This creates extra chunks and wrong `Code` metadata when the content includes an inner shorter fence or a mismatched fence.
`_resolve_code_chunk` reuses the opening-fence matcher for every candidate closing line, so any ` ``` ` or `~~~` line of length at least 3 is accepted as a terminator without checking that the fence character matches the opener, is at least as long, and is not followed by trailing text.
Run the provided Python snippet: create a Markdown string with an outer quadruple-backtick fence containing an inner triple-backtick Python fence, then call `ExperimentalMarkdownSyntaxTextSplitter().split_text(text)`. Observed: four chunks are returned instead of two; the outer block ends immediately after the inner opening fence and its `Code` metadata is '`markdown' instead of 'markdown'. Expected: two chunks, first containing the entire outer code block with `Code='markdown'`, second containing the trailing 'After'.
Fixing Code Block
@staticmethod
def _is_closing_fence(line: str, fence_char: str, min_len: int) -> bool:
stripped = line.strip()
if len(stripped) < min_len:
return False
if stripped[0] != fence_char:
return False
return all(ch == fence_char for ch in stripped)
def _resolve_code_chunk(self, lines: list[str], start: int) -> tuple[int, str, str]:
opening_line = lines[start].strip()
match = re.match(r'^(?P<fence>`{3,}|~{3,})(?P<info>.*)$', opening_line)
if not match:
raise ValueError(f'Invalid opening fence at line {start}: {opening_line!r}')
fence_char = match.group('fence')[0]
fence_len = len(match.group('fence'))
language = match.group('info').strip()
content_lines = []
for i in range(start + 1, len(lines)):
line = lines[i]
if self._is_closing_fence(line, fence_char, fence_len):
content = '\n'.join(content_lines)
return i + 1, language, content
content_lines.append(line)
content = '\n'.join(content_lines)
return len(lines), language, content
The new `_is_closing_fence` helper only returns true when the stripped line consists solely of the opening fence character and has at least the opening fence length. `_resolve_code_chunk` now captures the full opening fence with a greedy match and uses this helper instead of the generic opening-fence regex, preventing shorter inner fences, mismatched fence characters, and fence-like lines with trailing text from closing the block.
Edge Case Audit
This changes behavior for documents that previously relied on mismatched or shorter fences as terminators; such inputs may now merge separate code blocks or swallow text after an unclosed outer fence. Rollback suggestion: keep a copy of the original `_resolve_code_chunk` and `_is_closing_fence` methods, pin `langchain-text-splitters==1.1.2` during deployment, and re-run all affected RAG/indexing pipelines. Verify inputs with unclosed code fences, as the fallback now consumes the remaining document into one large code chunk, which may be worse than the old behavior for some use cases.