AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

Postgres Fails To Restart After Compute Resize; Restore-To-New-Project Stuck COMING_UP Due To Image Mismatch; False ACTIVE_HEALTHY Status

Production database on us-east-2 became unreachable after a compute resize, with no Postgres logs. Management API incorrectly reports ACTIVE_HEALTHY while the restore-to-new-project hangs in COMING_UP because a backup from image 17.6.1.104 was restored onto image 17.11.0.002. Urgent support tickets have not been answered for 12+ hours.

criticalConfidence 82%Supabase PlatformAffected V17.6.1.104Affected V17.11.0.002

Origin Analysis

Postgres failed to start after compute resize due to a platform orchestration defect: the old VM termination left a stale or unavailable data volume, the new instance came up without successfully initializing Postgres, and the health check only validated VM status rather than actual database connectivity, resulting in a false ACTIVE_HEALTHY. The restore-to-new-project hung because the platform lacks a version compatibility check, applying a physical backup from image 17.6.1.104 onto a project running 17.11.0.002 without preflight validation. Agent A (30%): The most probable root cause is a combination of disk/IO exhaustion leading to a crash during resize, followed by a failure to remount or reattach the Postgres data volume on the new VM, plus a monitoring gap that doesn't probe database-level liveness. Agent B (50%): Hotfix provided below. Agent C (20%): Risk warnings included.
1. Have a Supabase Pro project on t4g.nano with heavy disk I/O and memory pressure. 2. Initiate a compute resize (Nano to Small) from the dashboard. 3. Observe that the dashboard shows the new instance size but Postgres writes no new log lines. 4. Restart project; status becomes Unhealthy. 5. Attempt restore-to-new-project from a physical backup taken on image 17.6.1.104 while the new project is on image 17.11.0.002. 6. The new project remains in COMING_UP indefinitely, with only errors like `relation "realtime.subscription" does not exist`. 7. Management API `GET /v1/projects/{ref}` continues to report ACTIVE_HEALTHY for the original project despite database being unreachable.

Fixing Code Block

Edge Case Audit

Manual removal of postmaster.pid and starting Postgres can cause data corruption if another process is running or if the data directory is in an inconsistent state. Only run under platform engineer supervision with a snapshot of the volume before changes. The restore-to-new-project API may not officially support selecting an image version; the `postgres_version` field in the create payload may be ignored by the platform, resulting in another stuck restore. If the script fails, do not delete the original project or its disk; escalate to Supabase infrastructure on-call. Rollback: keep all original resources untouched, terminate any recovery project if it becomes unhealthy, and rely on support to manually recover the original VM.

Ecosystem Topology