Grep_search Python Fallback Skips Valid UTF-8 Files On Windows Due To Missing Encoding Parameter In Read_text()
When ripgrep is unavailable, the Python fallback in FilesystemFileSearchMiddleware.grep_search reads files using Path.read_text() without specifying an encoding. On Windows, the default encoding is often cp1252, causing UnicodeDecodeError for UTF-8 files. These errors are silently swallowed, leading to incomplete or empty search results even when matching content exists.
The _python_search() method calls file_path.read_text() without an explicit encoding argument. Python's read_text() uses the platform default encoding (cp1252 on many Windows systems) when encoding is not provided. UTF-8 encoded files then raise UnicodeDecodeError, which is caught by the existing except (UnicodeDecodeError, PermissionError) block and skipped silently. This is a design flaw in the fallback path that assumes the default encoding is UTF-8.
1. On a Windows machine with default locale set to cp1252 (common for en-US), ensure ripgrep is not installed or set the environment to force the Python fallback.
2. Create a directory with a UTF-8 encoded markdown file containing a searchable string, e.g., 'needle'.
3. Use FilesystemFileSearchMiddleware.grep_search with the search root set to that directory and query 'needle'.
4. Observe that no results are returned and no error is raised, because the UnicodeDecodeError during file reading is silently skipped.
5. Same search works on Linux or when ripgrep is available.
Explicitly pass encoding="utf-8" to read_text(). This ensures the Python fallback reads text files as UTF-8 on all platforms, independent of the OS default. Files that are binary or in an incompatible encoding will still raise UnicodeDecodeError and be safely skipped by the existing exception handler, preserving the current behavior for truly unreadable files.
Edge Case Audit
This change can cause files that are not encoded in UTF-8 (e.g., legacy cp1252, Latin-1) to be skipped even if they were previously readable under the Windows default encoding. This may lead to missing results in those non-UTF-8 files. To mitigate, consider a fallback strategy: first try utf-8, then try locale.getpreferredencoding(False), and only skip if both fail. No new concurrency or threading risks are introduced. For rollback, revert to the original read_text() call, but this reintroduces the Windows UTF-8 failure.