Fixing Connection Pool Exhaustion Before PgBouncer Runs Dry
Postgres was never designed to hold tens of thousands of idle connections open. Each one costs real memory on the server, which is why almost every production Postgres setup sits behind a pooler like PgBouncer, and why that pooler is usually the thing that falls over first when traffic spikes.
The symptom is always the same: requests start queueing, latency climbs, and the database itself looks fine in every metric you'd normally check. The cause is almost always a pool that was sized for average load, not peak load, or a client that's holding connections longer than it needs to.
The symptom: requests queueing behind a full pool
When every pooled connection is checked out, new requests wait in line rather than failing immediately. That queueing is often invisible until it isn't: latency creeps up gradually as the queue grows, then crosses a threshold and every request downstream starts timing out at once. By the time an alert fires, the queue has usually been building for several minutes.
This is why pool utilization deserves its own dashboard, not just database CPU and memory. A pool sitting at capacity for an extended stretch is a leading indicator of an outage, not a lagging one.
Session pooling versus transaction pooling
PgBouncer supports a few pooling modes, and the choice matters more than most teams realize. Session pooling hands a client a connection for the life of its session, which is safe for everything but doesn't share connections efficiently. Transaction pooling returns the connection to the pool the moment a transaction commits, which lets far fewer real Postgres connections serve far more concurrent clients.
The tradeoff is that transaction pooling breaks anything that depends on session state surviving between queries, like prepared statements tied to a session, session-level advisory locks, or certain extensions. Most application traffic doesn't need that state, which is why transaction pooling is the default choice for anything read and write heavy.
Sizing the pool without guessing
Start from the database's actual connection ceiling, not from a round number that felt safe. Postgres has a hard max_connections setting, and every other process on the box, replication, backups, monitoring, needs a share of it too. The pool should be sized so that peak concurrent load fits comfortably under that ceiling with room left for those other processes.
Then watch the pool under real load, not synthetic load. A pool that looks fine in a staging test with ten concurrent users can behave completely differently once a slow query starts holding connections open longer than expected during a traffic spike.
Work through the sizing in this order:
- Find the database's actual connection ceiling in its max_connections setting instead of starting from a round number.
- Set aside a share of that ceiling for replication, backups, monitoring, and every other process on the box.
- Size the pool so peak concurrent load fits comfortably under what remains, not average load.
- Watch pool utilization directly during real peak traffic, and treat a pool sitting near full as undersized.
- Keep a documented runbook for raising the pool size or the database's connection ceiling.
What breaks when an ORM assumes session pooling
Plenty of ORMs and drivers assume they can rely on session-level features: temporary tables, session variables, or a prepared statement cache tied to the connection. Point one of these at a transaction-mode pooler and you'll see errors that look random, because they depend on which physical connection happens to get reused underneath the abstraction.
When PgBouncer runs out of pooled connections, every queued request counts against the same downtime budget your uptime target already promised away1. Catching an ORM's pooling assumptions in a load test, before that happens in production, is a lot cheaper than catching it during an incident.
A checklist before you ship PgBouncer
Confirm which pooling mode each of your services actually needs, not which one is the default. Load test with realistic concurrency, not a handful of requests from your laptop. Alert on pool utilization directly, not just on downstream symptoms like request latency. And keep a documented runbook for raising the pool size or the database's connection ceiling, because that decision during an actual incident should take minutes, not require someone to relearn the tradeoffs from scratch.
It's also worth deciding in advance which services get priority if the database's connection ceiling is ever the binding constraint. A background export job and a customer-facing checkout flow shouldn't be competing for the same slice of the pool with no ordering between them.
What Good Looks Like
Good connection pool management means the pool is sized against a load test at real peak concurrency, and its utilization is a monitored metric on its own.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Why does Postgres need a connection pooler at all?
Each Postgres connection is a real operating system process that consumes memory whether it's doing work or sitting idle. Modern applications open far more connections than Postgres can hold efficiently, so a pooler multiplexes many client connections onto a much smaller number of real database connections.
Is transaction pooling always the right choice?
It's the right default for typical web application traffic, but not for code that relies on session-level features like advisory locks or session variables. Check what your ORM and any long-running background jobs actually depend on before switching modes.
How do we know if our pool is too small?
Watch pool utilization directly rather than waiting for request latency to climb. If the pool spends meaningful time at or near full capacity during normal peak traffic, it's undersized for the load you already have, not just for future growth.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
A Rollout Checklist for Swapping Models in Production
A rollout checklist for swapping AI models in production: evaluation gates, canary and shadow traffic, and fast rollback paths.
Why Inference Latency Creeps Up After You Ship
Where AI model-serving latency actually hides: tokenization, queueing, batching windows, and network hops, plus a worked example fix.
PgBouncer in Production: A Connection Pooling Checklist
Why Postgres runs out of connections before it runs out of CPU, and a rollout checklist for putting PgBouncer in front of it safely.
Sharding Your Feature Store as Inference Traffic Grows
When a single feature store or vector database starts limiting inference throughput, and the sharding approaches that fit a retrieval-heavy serving path.
Zero Trust for Machines Calling Your Model Endpoints
Why internal services calling your inference endpoints still need identity checks, and where CrowdStrike-style posture checks and Tenable-style scanning fit.
Fixing 'Too Many Connections' Without Just Raising the Limit
A troubleshooting walkthrough for too many connections errors: what's actually consuming your pool, and the fixes that hold up under real load.