Pyproject.Toml Is Read With Locale Encoding In Unit Tests, Causing UnicodeDecodeError On Non-UTF-8 Windows Locales
Multiple unit test call sites read pyproject.toml without specifying encoding='utf-8'. On Windows with a non-UTF-8 locale (e.g., cp936), Python defaults to the ANSI code page, leading to UnicodeDecodeError because the file is valid UTF-8 and contains non-ASCII characters. Issue #40588 fixed two call sites, but two remain: one active in langchain_v1 summarization tests and one latent in langchain-classic dependencies tests.
Path.open() and Path.read_text() are called without an explicit encoding argument. When the locale is not UTF-8 (common on Windows), Python uses the locale’s preferred encoding (e.g., cp936) instead of UTF-8. The pyproject.toml file contains UTF-8 non-ASCII bytes (an en-dash in allowed-confusables), which cannot be decoded by cp936, raising UnicodeDecodeError.
On Windows with a non-UTF-8 system locale (cp936), from a fresh clone:
cd libs/langchain_v1
uv sync --group test
PYTHONUTF8=0 uv run --group test pytest tests/unit_tests/test_version.py tests/unit_tests/test_dependencies.py "tests/unit_tests/agents/middleware/implementations/test_summarization.py::test_trigger_conditions_legacy_tuple_view_remove_in_2_0" -q
The tests fail with UnicodeDecodeError: 'gbk' codec can't decode byte 0x93 in position 4345: illegal multibyte sequence. The same failure can be reproduced with a one-liner: PYTHONUTF8=0 python -c "open('libs/langchain_v1/pyproject.toml').read()"
Fixing Code Block
--- a/libs/langchain_v1/tests/unit_tests/agents/middleware/implementations/test_summarization.py
+++ b/libs/langchain_v1/tests/unit_tests/agents/middleware/implementations/test_summarization.py
@@ -54,7 +54,7 @@ def test_trigger_conditions_legacy_tuple_view_remove_in_2_0():
- data = pyproject_path.read_text()
+ data = pyproject_path.read_text(encoding="utf-8")
--- a/libs/langchain/tests/unit_tests/test_dependencies.py
+++ b/libs/langchain/tests/unit_tests/test_dependencies.py
@@ -16,7 +16,7 @@ def test_required_dependencies():
- with open(pyproject_path) as f:
+ with open(pyproject_path, encoding="utf-8") as f:
...
Adding encoding='utf-8' explicitly tells Python to decode the file as UTF-8 regardless of the platform default. This aligns with the fix already applied in #40588 for the other two call sites and is the standard best practice for reading text files with a known encoding. The change is minimal and directly eliminates the UnicodeDecodeError.
Edge Case Audit
Explicit UTF-8 is safe and backward compatible; it does not change behavior on systems already using UTF-8. Rollback is trivial: remove the encoding parameter to restore the previous locale-dependent behavior. The main residual risk is that other future file reads in tests may still omit encoding; to prevent regressions, enforce a lint rule (e.g., use tokenize.open for all source file reads) or set PYTHONUTF8=1 in CI. No concurrency or multithreading concerns because the file is read once during test setup.