AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

Parse_json_markdown Fails On Uppercase JSON Fence Tags

The parse_json_markdown function in langchain_core.utils.json uses a case-sensitive regex to identify fenced code blocks, causing parsing failures when models output fences with uppercase 'JSON' or 'Json'. This affects users of JsonOutputParser and other utilities relying on this function.

mediumConfidence 95%Langchain-CoreAffected V1.6.2

Origin Analysis

The regex _json_markdown_re = re.compile(r"```(json)?(.*)", re.DOTALL) matches the optional language tag 'json' case-sensitively. When the fence tag is uppercase (e.g., 'JSON'), the tag is not captured as the optional group and remains part of the JSON payload, leading to invalid JSON parse errors.
import json from langchain_core.utils.json import parse_json_markdown # Lowercase works print(parse_json_markdown('\n```json\n{"foo": "bar"}\n```\n')) # {'foo': 'bar'} # Uppercase fails print(parse_json_markdown('\n```JSON\n{"foo": "bar"}\n```\n')) # raises OutputParserException or JSONDecodeError

Fixing Code Block

import re # Replace the existing regex definition with the following line: _json_markdown_re = re.compile(r"```(json)?(.*)", re.DOTALL | re.IGNORECASE)
Adding re.IGNORECASE flag to the regex makes the optional 'json' language tag match case-insensitively. This ensures that fences tagged 'JSON', 'Json', or any other casing are correctly recognized and stripped before the JSON content is parsed.

Edge Case Audit

The change is minimal and backward compatible for existing lowercase and untagged fences. However, it may cause previously failing uppercase-tagged inputs to now succeed, which is the intended behavior. No destructive operations are performed. If any unforeseen issues arise, revert the regex flag change. Ensure comprehensive tests are added for all case variants.

Ecosystem Topology