AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

Langchain-Profiles Refresh Corrupts Strings Containing JSON Literal Substrings

`langchain-model-profiles` CLI `refresh` performs global string replacements (`true`->`True`, `false`->`False`, `null`->`None`) on a `json.dumps` result, corrupting quoted JSON keys and values that contain those substrings. The generated `_profiles.py` still imports successfully but with silently altered profile data.

mediumConfidence 95%LangchainAffected V0.0.6

Origin Analysis

The serializer uses `json_str.replace("true", "True").replace("false", "False").replace("null", "None")`, which cannot distinguish JSON literal tokens from substrings inside JSON string values or string keys. This is a classic design flaw of regex/string-replace based JSON-to-Python literal conversion.
1. Create a temporary directory. 2. Patch `httpx.get` to return a profile payload where a model id contains `null` and a profile name contains `true false null`. 3. Run `langchain_model_profiles.cli.refresh("anthropic", tmp_dir)`. 4. Import generated `_profiles.py` and inspect the profile dict: `null-model` becomes `None-model`, and `true false null` becomes `True False None`.

Fixing Code Block

import io import json import tokenize _LITERAL_REPLACEMENTS = {"true": "True", "false": "False", "null": "None"} def _json_to_python_literals(json_str: str) -> str: tokens = [] for token in tokenize.generate_tokens(io.StringIO(json_str).readline): if token.type == tokenize.NAME and token.string in _LITERAL_REPLACEMENTS: token = (token.type, _LITERAL_REPLACEMENTS[token.string], token.start, token.end, token.line) tokens.append(token) return tokenize.untokenize(tokens) # Replace the existing json_str.replace chain in refresh() with: json_str = json.dumps(merged_profile) json_str = _json_to_python_literals(json_str)
The fix tokenizes the JSON output using Python's standard `tokenize` module. In this token stream, true/false/null appear only as standalone NAME tokens when unquoted; quoted strings remain STRING tokens and are never replaced. The token tuple's string field is changed for the three mapping keys and the stream is reconstructed with `untokenize`, preserving original formatting because replacement strings have identical lengths.

Edge Case Audit

The fix relies on Python's tokenizer maintaining compatibility with JSON token boundaries; this holds for standard JSON emitted by `json.dumps`. No known version-specific issues, but if the tokenizer behavior changes in a future Python release, regression tests should cover quoted substrings. Rollback: reapply the original `json_str.replace(...)` chain if the tokenize import is unavailable (Python < 3.8? tokenize exists; no issue). Concurrency/multi-thread: CLI runs in a single process, no shared mutable state.

Ecosystem Topology