AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

PIIMiddleware Fails To Detect Email, IP, And MAC Addresses Adjacent To Non-ASCII Text Due To Unicode-Aware \B

The PII detectors for email, IP, and MAC addresses use \b word boundaries which are Unicode-aware. Non-ASCII letters (CJK, Cyrillic, Arabic, accented Latin) are treated as word characters, so the boundary does not trigger between them and an ASCII PII token. This allows PII to pass through unredacted, unblocked, or unmasked.

highConfidence 95%LangchainAffected V1.3.18

Origin Analysis

Python's regex \b is Unicode-aware by default. It matches between a word character and a non-word character, where word characters include all Unicode letters. Patterns anchored with \b therefore fail to start or end when PII is directly adjacent to non-ASCII text. The email, IP, and MAC address detectors all use \b at both ends; the URL detector does not, which is why URL detection works across scripts.
Run the following Python code with langchain 1.3.18:\n```python\nfrom langchain.agents.middleware._redaction import detect_email, detect_ip, detect_mac_address\nprint(detect_email('联系alice@example.com')) # []\nprint(detect_email('Contact alice@example.com')) # ['alice@example.com']\nprint(detect_ip('почта192.168.1.100')) # []\nprint(detect_mac_address('café00:1A:2B:3C:4D:5E')) # []\n```\nThe ASCII-prefixed cases return matches, while non-ASCII-prefixed cases return empty lists.

Fixing Code Block

Edge Case Audit

This change expands detection scope: previously hidden PII next to non-ASCII characters will now be redacted/blocked, which is the intended security improvement. Review any application logic that relies on the old Unicode boundary semantics, especially for mixed-script strings. The credit_card detector still uses \b and remains vulnerable; do not assume full PII coverage after this fix. Rollback: revert the boundary constants to \b in the affected functions. Test across Python versions, as lookbehind assertions with fixed-width alternatives are safe in Python 3.6+. Differential fuzzing reported by the issue author shows no loss of ASCII detection, but run your own regression tests.

Ecosystem Topology