AI Model Serving & Inference OptimizationPlaybook3 min readUpdated September 2026

Fixing Connection Pool Exhaustion Before PgBouncer Runs Dry

Postgres was never designed to hold tens of thousands of idle connections open. Each one costs real memory on the server, which is why almost every production Postgres setup sits behind a pooler like PgBouncer, and why that pooler is usually the thing that falls over first when traffic spikes.

The symptom is always the same: requests start queueing, latency climbs, and the database itself looks fine in every metric you'd normally check. The cause is almost always a pool that was sized for average load, not peak load, or a client that's holding connections longer than it needs to.

The symptom: requests queueing behind a full pool

When every pooled connection is checked out, new requests wait in line rather than failing immediately. That queueing is often invisible until it isn't: latency creeps up gradually as the queue grows, then crosses a threshold and every request downstream starts timing out at once. By the time an alert fires, the queue has usually been building for several minutes.

This is why pool utilization deserves its own dashboard, not just database CPU and memory. A pool sitting at capacity for an extended stretch is a leading indicator of an outage, not a lagging one.

Session pooling versus transaction pooling

PgBouncer supports a few pooling modes, and the choice matters more than most teams realize. Session pooling hands a client a connection for the life of its session, which is safe for everything but doesn't share connections efficiently. Transaction pooling returns the connection to the pool the moment a transaction commits, which lets far fewer real Postgres connections serve far more concurrent clients.

The tradeoff is that transaction pooling breaks anything that depends on session state surviving between queries, like prepared statements tied to a session, session-level advisory locks, or certain extensions. Most application traffic doesn't need that state, which is why transaction pooling is the default choice for anything read and write heavy.

Sizing the pool without guessing

Start from the database's actual connection ceiling, not from a round number that felt safe. Postgres has a hard max_connections setting, and every other process on the box, replication, backups, monitoring, needs a share of it too. The pool should be sized so that peak concurrent load fits comfortably under that ceiling with room left for those other processes.

Then watch the pool under real load, not synthetic load. A pool that looks fine in a staging test with ten concurrent users can behave completely differently once a slow query starts holding connections open longer than expected during a traffic spike.

Work through the sizing in this order:

  1. Find the database's actual connection ceiling in its max_connections setting instead of starting from a round number.
  2. Set aside a share of that ceiling for replication, backups, monitoring, and every other process on the box.
  3. Size the pool so peak concurrent load fits comfortably under what remains, not average load.
  4. Watch pool utilization directly during real peak traffic, and treat a pool sitting near full as undersized.
  5. Keep a documented runbook for raising the pool size or the database's connection ceiling.

What breaks when an ORM assumes session pooling

Plenty of ORMs and drivers assume they can rely on session-level features: temporary tables, session variables, or a prepared statement cache tied to the connection. Point one of these at a transaction-mode pooler and you'll see errors that look random, because they depend on which physical connection happens to get reused underneath the abstraction.

When PgBouncer runs out of pooled connections, every queued request counts against the same downtime budget your uptime target already promised away1. Catching an ORM's pooling assumptions in a load test, before that happens in production, is a lot cheaper than catching it during an incident.

A checklist before you ship PgBouncer

Confirm which pooling mode each of your services actually needs, not which one is the default. Load test with realistic concurrency, not a handful of requests from your laptop. Alert on pool utilization directly, not just on downstream symptoms like request latency. And keep a documented runbook for raising the pool size or the database's connection ceiling, because that decision during an actual incident should take minutes, not require someone to relearn the tradeoffs from scratch.

It's also worth deciding in advance which services get priority if the database's connection ceiling is ever the binding constraint. A background export job and a customer-facing checkout flow shouldn't be competing for the same slice of the pool with no ordering between them.

Executive Capability Standard

What Good Looks Like

Good connection pool management means the pool is sized against a load test at real peak concurrency, and its utilization is a monitored metric on its own.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Read up on the difference between session and transaction pooling, and check which one each of your services actually assumes.
2. Do Manually:Watch pool utilization by hand during your next known traffic spike, such as a launch or a marketing send, to see how close you run to the ceiling.
3. Delegate:Give one engineer ownership of the pooling configuration and the runbook for adjusting it under load.
4. Automate:Add pool utilization as a first-class alert, separate from database CPU and memory, so queueing shows up before requests start timing out.
5. Buy:Move to a managed Postgres provider with built-in, tunable connection pooling if operating PgBouncer yourself isn't where your team's time is best spent.

How to Get Started

Frequently Asked Questions

Why does Postgres need a connection pooler at all?

Each Postgres connection is a real operating system process that consumes memory whether it's doing work or sitting idle. Modern applications open far more connections than Postgres can hold efficiently, so a pooler multiplexes many client connections onto a much smaller number of real database connections.

Is transaction pooling always the right choice?

It's the right default for typical web application traffic, but not for code that relies on session-level features like advisory locks or session variables. Check what your ORM and any long-running background jobs actually depend on before switching modes.

How do we know if our pool is too small?

Watch pool utilization directly rather than waiting for request latency to climb. If the pool spends meaningful time at or near full capacity during normal peak traffic, it's undersized for the load you already have, not just for future growth.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides