AI & Agent Dev Bug Sandbox logo
AI & Agent Dev Bug Sandbox
Back to Radar

GoTrue Auth Silent Freeze On Pro Project — POST Endpoints Dead While /Health Responds OK

Supabase Pro project's GoTrue service became unresponsive, causing all auth POST endpoints (e.g., /auth/v1/token) to timeout while GET /auth/v1/health returned 200 quickly, preventing Kubernetes auto-restart and leaving users unable to sign in/up for 30+ minutes.

criticalConfidence 65%GoTrue

Origin Analysis

Likely a deadlock or goroutine starvation in GoTrue's authentication handler path. A shared mutex or critical section (e.g., rate limiter, session cache, or DB connection pool for auth) was held indefinitely, blocking all POST token requests. The /health endpoint does not check this critical section, so the liveness probe incorrectly reports the container as healthy.
1. POST to https://hapzzxnhedpdnaffsylc.supabase.co/auth/v1/token?grant_type=password with valid apikey and any body -> timeout >15s, no response. 2. GET https://hapzzxnhedpdnaffsylc.supabase.co/auth/v1/health with same apikey -> 401 in ~75ms (server responds). 3. GET https://hapzzxnhedpdnaffsylc.supabase.co/rest/v1/ -> 401 in ~90ms (PostgREST fine). 4. Dashboard Authentication Users page hangs indefinitely. 5. Restart project recovers service in ~2 min.

Fixing Code Block

Edge Case Audit

Using TryLock on the auth mutex may cause false positives if the mutex is legitimately held for longer than expected under high concurrency, leading to unnecessary pod restarts. The 503 response should be tuned (e.g., use a timeout instead of instant failure) to avoid flapping. This change only detects mutex-based deadlocks, not other types of hangs. To rollback, revert the health handler to its previous implementation that only returns 200.

Ecosystem Topology