AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

Core: _Convert_openai_format_to_data_block Hard-Codes Mime_type On Base64 File Blocks

In langchain-core's OpenAI message translation, base64-encoded file blocks always get mime_type='application/pdf', ignoring the MIME type specified in the data URI. This causes non-PDF files (e.g., CSV, text) to be silently mislabeled, which can break downstream processing that relies on content_blocks.

highConfidence 95%Langchain-CoreAffected V1.3.0

Origin Analysis

In _convert_openai_format_to_data_block, the file branch hard-codes 'application/pdf' instead of extracting the MIME type from the parsed data URI via parsed['mime_type'].
Run the following Python code: ```python from langchain_core.messages import HumanMessage msg = HumanMessage(content=[ { "type": "file", "file": { "filename": "sheet.csv", "file_data": "data:text/csv;base64,aGVsbG8=", }, }, ]) for block in msg.content_blocks: print(block) ``` Observe that the printed block shows mime_type='application/pdf' instead of 'text/csv'.

Fixing Code Block

Edge Case Audit

The fallback for missing MIME type (parsed is None) still defaults to 'application/pdf', which could still mislabel files with data URIs lacking a MIME type. Consider changing that fallback to a more generic type like 'application/octet-stream' in a separate improvement. Downstream code that implicitly assumed all file blocks were PDF may need adjustment. Rollback is straightforward: revert to the previous hard-coded version if unexpected behavior arises.

Ecosystem Topology