AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

Pyproject.Toml Is Read With Locale Encoding In Unit Tests, Causing UnicodeDecodeError On Non-UTF-8 Windows Locales

Multiple unit test call sites read pyproject.toml without specifying encoding='utf-8'. On Windows with a non-UTF-8 locale (e.g., cp936), Python defaults to the ANSI code page, leading to UnicodeDecodeError because the file is valid UTF-8 and contains non-ASCII characters. Issue #40588 fixed two call sites, but two remain: one active in langchain_v1 summarization tests and one latent in langchain-classic dependencies tests.

mediumConfidence 96%Langchain

Origin Analysis

Path.open() and Path.read_text() are called without an explicit encoding argument. When the locale is not UTF-8 (common on Windows), Python uses the locale’s preferred encoding (e.g., cp936) instead of UTF-8. The pyproject.toml file contains UTF-8 non-ASCII bytes (an en-dash in allowed-confusables), which cannot be decoded by cp936, raising UnicodeDecodeError.
On Windows with a non-UTF-8 system locale (cp936), from a fresh clone: cd libs/langchain_v1 uv sync --group test PYTHONUTF8=0 uv run --group test pytest tests/unit_tests/test_version.py tests/unit_tests/test_dependencies.py "tests/unit_tests/agents/middleware/implementations/test_summarization.py::test_trigger_conditions_legacy_tuple_view_remove_in_2_0" -q The tests fail with UnicodeDecodeError: 'gbk' codec can't decode byte 0x93 in position 4345: illegal multibyte sequence. The same failure can be reproduced with a one-liner: PYTHONUTF8=0 python -c "open('libs/langchain_v1/pyproject.toml').read()"

Fixing Code Block

--- a/libs/langchain_v1/tests/unit_tests/agents/middleware/implementations/test_summarization.py +++ b/libs/langchain_v1/tests/unit_tests/agents/middleware/implementations/test_summarization.py @@ -54,7 +54,7 @@ def test_trigger_conditions_legacy_tuple_view_remove_in_2_0(): - data = pyproject_path.read_text() + data = pyproject_path.read_text(encoding="utf-8") --- a/libs/langchain/tests/unit_tests/test_dependencies.py +++ b/libs/langchain/tests/unit_tests/test_dependencies.py @@ -16,7 +16,7 @@ def test_required_dependencies(): - with open(pyproject_path) as f: + with open(pyproject_path, encoding="utf-8") as f: ...
Adding encoding='utf-8' explicitly tells Python to decode the file as UTF-8 regardless of the platform default. This aligns with the fix already applied in #40588 for the other two call sites and is the standard best practice for reading text files with a known encoding. The change is minimal and directly eliminates the UnicodeDecodeError.

Edge Case Audit

Explicit UTF-8 is safe and backward compatible; it does not change behavior on systems already using UTF-8. Rollback is trivial: remove the encoding parameter to restore the previous locale-dependent behavior. The main residual risk is that other future file reads in tests may still omit encoding; to prevent regressions, enforce a lint rule (e.g., use tokenize.open for all source file reads) or set PYTHONUTF8=1 in CI. No concurrency or multithreading concerns because the file is read once during test setup.

Ecosystem Topology