AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

RecursiveCharacterTextSplitter Skips Strip_whitespace For Unsplittable Pieces

When a piece cannot be split further because no separators remain, it is appended directly to the output without applying strip_whitespace, causing inconsistent leading/trailing whitespace compared to merged chunks.

mediumConfidence 95%LangchainAffected V1.1.2

Origin Analysis

In RecursiveCharacterTextSplitter._split_text, the branch where `new_separators` is empty appends the raw segment `s` to final_chunks, bypassing the strip_whitespace handling that _merge_splits applies to assembled chunks. This only occurs when the separator list does not end with an empty string and a segment is longer than chunk_size.
```python from langchain_text_splitters import RecursiveCharacterTextSplitter splitter = RecursiveCharacterTextSplitter(separators=["\n", " "], chunk_size=5, chunk_overlap=0) print(splitter.split_text("hi supercalifragilistic\nok")) # Output: ['hi', ' supercalifragilistic', 'ok'] # leading space splitter = RecursiveCharacterTextSplitter(separators=[" "], chunk_size=3, chunk_overlap=0) print(splitter.split_text("ab ab ab")) # Output: ['ab', ' ab', ' ab'] ```

Fixing Code Block

def _split_text(self, text: str, separators: list[str]) -> list[str]: """Split incoming text and return chunks.""" final_chunks = [] separator = separators[-1] new_separators = [] for i, _s in enumerate(separators): _separator = _s if self._keep_separator else "" if _s == "": separator = _s new_separators = separators[i + 1 :] break if re.search(_s, text): separator = _s new_separators = separators[i + 1 :] break _separator = separator if self._keep_separator else "" splits = _split_text_with_regex(text, _separator) _good_splits = [] _separator = "" if self._keep_separator else separator for s in splits: if self._length_function(s) < self._chunk_size: _good_splits.append(s) else: if _good_splits: merged_text = self._merge_splits(_good_splits, _separator) final_chunks.extend(merged_text) _good_splits = [] if not new_separators: # Apply strip_whitespace handling to unsplittable piece processed = s.strip() if self.strip_whitespace else s if processed: final_chunks.append(processed) else: other_info = self._split_text(s, new_separators) final_chunks.extend(other_info) if _good_splits: merged_text = self._merge_splits(_good_splits, _separator) final_chunks.extend(merged_text) return final_chunks
The fix replaces the direct `final_chunks.append(s)` with conditional processing: if `self.strip_whitespace` is True, strip the segment and only append if non-empty; otherwise append unchanged. This mirrors the behavior of `_merge_splits`, ensuring consistent whitespace handling for all output chunks.

Edge Case Audit

This change only affects unsplittable pieces that are entirely whitespace or have leading/trailing whitespace. If downstream code relies on the exact raw string from unsplittable pieces (including leading separator spaces when keep_separator=True), those chunks will now be stripped and possibly dropped. For existing installations, re-indexing may be needed because chunk boundaries and counts can shift. Rollback: if unexpected results occur, set `strip_whitespace=False` or revert to the previous version; otherwise validate with a representative corpus.

Ecosystem Topology