AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

Production PostgreSQL Unavailable After Disk Exhaustion And Expansion; Crash Recovery Not Bringing Database Online

A disk-full event during a migration caused PostgreSQL to crash with SQLSTATE 53100. After expanding disk to 8GB, crash recovery checkpoint completed and read-only mode ended, but all connection methods (Supavisor, direct, REST, SQL Editor) remain timed out, indicating the PostgreSQL backend is not accepting connections or is stuck in a state that prevents normal operation.

criticalConfidence 65%Supabase

Origin Analysis

Disk exhaustion during a WAL write forced the startup process to exit and shut down PostgreSQL. After disk expansion, PostgreSQL entered crash recovery but likely encountered a missing or corrupt WAL segment required for redo, causing the startup process to hang or fail before opening the listen socket. Consequently, Supavisor and authentication services cannot reach the backend, resulting in timeouts across all connection methods.
1. Fill the PostgreSQL data volume during a write-heavy migration until SQLSTATE 53100 is raised. 2. Observe postmaster shutdown. 3. Expand disk to 8GB and restart database/project. 4. See 'crash-recovery checkpoint complete' and 'read-only off' in dashboard logs. 5. Attempt connections via Supavisor session/transaction, direct PostgreSQL, REST, and SQL Editor; all time out.

Fixing Code Block

Edge Case Audit

Do not run pg_resetwal, delete WAL files, or restore an older backup without explicit data-loss acceptance; these actions can corrupt the cluster or discard current WAL history. A forced restart may interrupt in-progress recovery and should be preceded by a filesystem snapshot. Rollback: if the hotfix worsens the situation, restore from the latest base backup and replay available WAL, but do not overwrite the current data directory unless absolutely necessary.

Ecosystem Topology