RecursiveCharacterTextSplitter Skips Strip_whitespace For Unsplittable Pieces
When a piece cannot be split further because no separators remain, it is appended directly to the output without applying strip_whitespace, causing inconsistent leading/trailing whitespace compared to merged chunks.
In RecursiveCharacterTextSplitter._split_text, the branch where `new_separators` is empty appends the raw segment `s` to final_chunks, bypassing the strip_whitespace handling that _merge_splits applies to assembled chunks. This only occurs when the separator list does not end with an empty string and a segment is longer than chunk_size.
```python
from langchain_text_splitters import RecursiveCharacterTextSplitter
splitter = RecursiveCharacterTextSplitter(separators=["\n", " "], chunk_size=5, chunk_overlap=0)
print(splitter.split_text("hi supercalifragilistic\nok"))
# Output: ['hi', ' supercalifragilistic', 'ok'] # leading space
splitter = RecursiveCharacterTextSplitter(separators=[" "], chunk_size=3, chunk_overlap=0)
print(splitter.split_text("ab ab ab"))
# Output: ['ab', ' ab', ' ab']
```
Fixing Code Block
def _split_text(self, text: str, separators: list[str]) -> list[str]:
"""Split incoming text and return chunks."""
final_chunks = []
separator = separators[-1]
new_separators = []
for i, _s in enumerate(separators):
_separator = _s if self._keep_separator else ""
if _s == "":
separator = _s
new_separators = separators[i + 1 :]
break
if re.search(_s, text):
separator = _s
new_separators = separators[i + 1 :]
break
_separator = separator if self._keep_separator else ""
splits = _split_text_with_regex(text, _separator)
_good_splits = []
_separator = "" if self._keep_separator else separator
for s in splits:
if self._length_function(s) < self._chunk_size:
_good_splits.append(s)
else:
if _good_splits:
merged_text = self._merge_splits(_good_splits, _separator)
final_chunks.extend(merged_text)
_good_splits = []
if not new_separators:
# Apply strip_whitespace handling to unsplittable piece
processed = s.strip() if self.strip_whitespace else s
if processed:
final_chunks.append(processed)
else:
other_info = self._split_text(s, new_separators)
final_chunks.extend(other_info)
if _good_splits:
merged_text = self._merge_splits(_good_splits, _separator)
final_chunks.extend(merged_text)
return final_chunks
The fix replaces the direct `final_chunks.append(s)` with conditional processing: if `self.strip_whitespace` is True, strip the segment and only append if non-empty; otherwise append unchanged. This mirrors the behavior of `_merge_splits`, ensuring consistent whitespace handling for all output chunks.
Edge Case Audit
This change only affects unsplittable pieces that are entirely whitespace or have leading/trailing whitespace. If downstream code relies on the exact raw string from unsplittable pieces (including leading separator spaces when keep_separator=True), those chunks will now be stripped and possibly dropped. For existing installations, re-indexing may be needed because chunk boundaries and counts can shift. Rollback: if unexpected results occur, set `strip_whitespace=False` or revert to the previous version; otherwise validate with a representative corpus.