Production RAG & Vector Data ArchitecturePlaybook3 min readUpdated September 2026

The Connection Pool Checklist Most Teams Skip Until an Outage

Connection pool exhaustion is one of those failures that looks like a database problem, gets escalated to the database team, and turns out an hour later to be an application misconfiguration. The database was fine the whole time. The application just asked for more connections than it ever gave back.

This checklist covers the settings and habits that prevent that hour of wasted triage, organized around the mistakes that actually cause it.

How should you size a database connection pool?

A larger pool isn't automatically safer. Every connection your application holds open is also a connection the database has to manage, and Postgres or MySQL instances have their own hard ceiling on total connections across every client. Size each service's pool as a fraction of that ceiling, accounting for every other service that also connects to the same database, and leave headroom for administrative and monitoring connections. A common pitfall: setting a generous pool size per instance, then scaling the instance count, and quietly exceeding the database's connection ceiling during a traffic spike. Revisit pool sizing any time you change instance count, not only when you first configure it, since autoscaling policies drift away from the pooling assumptions that were true when the numbers were first chosen.

Set a connection timeout that fails fast

Without a timeout, a request waiting for a pool connection will hang indefinitely if the pool is exhausted, which turns a transient spike into a pile of stuck requests and, eventually, a cascading outage. Set an explicit acquire timeout, short enough that a stuck request fails and frees up resources rather than queuing forever, and make sure that failure surfaces as a clear error your monitoring can catch, not a silent hang. A fast, visible failure on one request is almost always a better outcome for the rest of your users than a slow, silent one that ties up resources while it waits.

Why do connections never get released back to the pool?

The most common cause of pool exhaustion isn't traffic, it's a code path that opens a connection and doesn't reliably return it, often in an exception handler that skips the cleanup step. Use your framework's connection pooling library rather than managing connections by hand, and add a test that deliberately triggers an error mid-transaction to confirm the connection still gets released afterward.

Match pool lifetime to actual network behavior

Connections that sit idle for a long time can get silently dropped by a load balancer, firewall, or the database itself, and the next request to grab that dead connection from the pool will fail. Set a maximum connection lifetime and an idle timeout shorter than any intermediate network device's own timeout, so your pool proactively recycles connections before they go stale rather than handing out ones that are already dead. This is especially easy to miss when a managed database sits behind a proxy layer you don't directly control, since the proxy's own idle timeout is often undocumented until it's the thing causing intermittent failures.

Use a separate, smaller pool for background jobs

A long-running batch job or report query that holds a connection open for minutes can starve the request-handling pool if they share the same connection budget. Give background work its own pool with its own, smaller connection allocation, so a slow report doesn't take down checkout because both were quietly drawing from the same limited pool of database connections.

A worked example: the report that looked like a database outage

Say an internal dashboard kicks off a heavy analytics query that holds its connection open while it scans a large table. If that dashboard shares a pool with the checkout service, every checkout request behind it in the queue eventually times out waiting for a free connection, and the on-call engineer sees checkout failing and starts investigating checkout code. The actual cause is a slow query in a completely different feature, quietly starving the shared pool. Separate pools per workload turn that hour of misdirected debugging into an immediate, obvious signal.

Run through this pre-flight list for each service:

  • Is the pool sized as a fraction of the database's connection ceiling, with headroom for other services and monitoring, and rechecked whenever instance count changes?
  • Is there an explicit acquire timeout, short enough that a stuck request fails instead of queuing forever?
  • Does a test trigger an error mid-transaction to confirm the connection still returns to the pool?
  • Are maximum lifetime and idle timeout shorter than any load balancer or firewall timeout between the app and the database?
  • Do background jobs and long reports use their own smaller pool, separate from request handling?
Executive Capability Standard

What Good Looks Like

A well-configured connection pool is sized against the database's actual connection ceiling, times out fast on acquisition, recycles idle connections before the network drops them, and keeps background work on a separate pool from request handling.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Read your database's documentation for its maximum connection limit and how that's shared across every client that connects.
2. Do Manually:Manually audit every service's pool size against that shared limit to check nothing is oversubscribed.
3. Delegate:Assign one engineer to own the pooling configuration standard so every new service starts from the same tested settings.
4. Automate:Add alerting on pool utilization so a team gets warned as usage approaches the configured maximum, not after requests start failing.
5. Buy:Use a managed connection pooler that sits in front of the database and enforces limits centrally, instead of trusting every service to self-limit.

How to Get Started

Frequently Asked Questions

What's a reasonable starting pool size for a small application?

There's no universal number, since it depends on your database's connection ceiling and how many other services share it. A common starting point is sizing each service's pool to a small fraction of the database's total connection limit, then adjusting based on observed queueing under real load rather than guessing upward.

How do we know we're actually hitting pool exhaustion in production?

Look for a spike in requests timing out while acquiring a database connection specifically, as opposed to timing out on the query itself. Most connection pooling libraries expose a metric for connections in use versus available, and that metric climbing to the pool's maximum right before errors start is the clearest signal.

Do we need connection pooling if we're already using a managed database?

Yes. A managed database still has a maximum connection limit, and pooling at the application layer is what keeps your services from individually or collectively exceeding it, especially as you scale instance count rather than instance size.

Why does pool exhaustion often look like a database performance problem at first?

Requests waiting for a connection and requests waiting on a slow query both show up as elevated response times, so the initial symptom looks identical. The distinguishing signal is where the time is actually spent: check whether latency is happening while acquiring a connection or while the query itself is running, since those point to completely different fixes.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides