AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

VectorStore.Add_texts Silently Truncates Texts When Len(Ids) != Len(Texts)

The `add_texts` method in `langchain-core` does not raise a `ValueError` when the number of provided IDs does not match the number of texts, despite the docstring stating otherwise. This leads to silent data loss as extra texts are discarded via `zip(strict=False)`. A regression from commit `6dd9f053e3` removed explicit validation.

highConfidence 95%Langchain-CoreAffected V1.4.7

Origin Analysis

The internal document construction uses `zip(texts, metadatas_, ids_, strict=False)`, which silently truncates to the shortest iterable. The explicit ID-length validation that existed prior to commit `6dd9f053e3` was removed, leaving the documented `ValueError` contract unenforced.
```python from langchain_core.embeddings import DeterministicFakeEmbedding from langchain_core.vectorstores import InMemoryVectorStore store = InMemoryVectorStore(embedding=DeterministicFakeEmbedding(size=3)) store.add_texts(["a", "b"], ids=["only-one"]) docs = store.get_by_ids(["only-one"]) print(len(docs)) # prints 1, but should raise ValueError before adding ```

Fixing Code Block

Edge Case Audit

Any caller that currently passes mismatched IDs and relies on the silent truncation (unlikely but possible) will now receive a `ValueError`. This is a breaking change in behavior but aligns with the documented contract. It is recommended to update tests to expect the exception. Rollback to the previous version would re-introduce the silent data loss. Additionally, the fix does not validate metadata length, so mismatched metadata lengths may still cause truncation, though this is outside the scope of this issue.

Ecosystem Topology