INITIALIZING…
SYSTEMS
00
00
Tushaar Naagar
Back to writing

Debugging

Debugging PostgreSQL Connection Exhaustion

Jun 10, 20262 min read

A pool at default size, a traffic bump, and a primary that looks 'idle' while every app server waits on acquire().

The symptom

App CPU is low. Postgres CPU is low. Requests hang. pg_stat_activity isn't full of queries — it's full of nothing, because the bottleneck is the pool on the Node side, or Postgres hit max_connections and new handshakes fail. From the product it looks like 'the API is down.' From the database it looks like a quiet afternoon.

Why it happens

Each Next.js / Nest instance opens its own pool. Three instances × default 10 looks fine until you add SSR, cron, and a serverless spike. You pass max_connections. Or you don't, and acquire() waits until the client gives up. Serverless is the worst version of this: a new pool per isolate, no sharing, death by a thousand connections.

How to fix it

One pool per process, sized on purpose. PgBouncer in transaction mode in front of the primary if you have many app workers. Kill long idle transactions. For serverless, a proxy that multiplexes is not optional. Alert on pool wait time and remaining connection slots — not only on query latency.

Takeaway

Connection exhaustion is a capacity-planning bug that shows up as a timeout. Count connections the way you count memory. 'Default' is not a plan.