Distributed Systems & Enterprise ResiliencePlaybook3 min readUpdated September 2026

Diagnosing and Living With Postgres Replica Lag

Replica lag becomes a problem the moment someone writes data and then immediately reads it back from a replica that hasn't caught up yet. The fix isn't chasing zero lag, which is rarely achievable and rarely necessary. It's knowing your current lag as a number, alerting before it becomes visible to a user, and deciding in advance which reads are allowed to be stale and which ones aren't.

Where the lag actually comes from

Postgres streams write-ahead log records from the primary to each replica, and the replica replays them before the data is visible to readers. Lag shows up when replay falls behind streaming: a long-running write transaction on the primary, a replica under-provisioned for its write volume, or network latency between regions if the replica sits somewhere geographically distant from the primary.

A replica that's consistently behind under normal load, not just during a spike, usually means it's undersized for the write volume it's replaying, not that something is broken. Check its I/O throughput and CPU headroom against the primary's before assuming the network is the culprit.

Measuring lag as a number you can alert on

Query pg_stat_replication on the primary, or the equivalent view your managed provider exposes, and track replay_lag directly rather than inferring it from application symptoms like stale-looking dashboards. Set an alert threshold based on what your application can actually tolerate, not on an arbitrary round number.

A reporting dashboard that reads from a replica can usually tolerate more staleness than a page that shows a user their own just-submitted data. Set separate thresholds for each use case rather than one number for the whole replica pool, and alert on the threshold that matters for your most lag-sensitive consumer.

Routing reads that can tolerate lag versus reads that can't

Not every read needs the primary. A common pattern is to route a user's own session to the primary for a short window right after they write, then fall back to replicas for everything else. That avoids the classic "I just saved this and it's not there" report without sending your entire read volume to the primary.

This only works if the routing rule is explicit and applied consistently, not left to individual engineers to remember per query. Bake it into your data access layer once, rather than asking every new endpoint to reinvent it, and treat any endpoint that skips the rule as a bug, not a one-off exception.

Decide read routing ahead of time with this short checklist:

  • List the reads where a stale answer would confuse a user, such as anything shown right after that user saves a change, and mark them primary-only.
  • Route a user's own session to the primary for a short window after each write, then fall back to replicas.
  • Let reporting dashboards and other reads that tolerate staleness go to replicas by default.
  • Build the routing rule into your data access layer once, so every engineer gets it without having to remember it.

The failover trap: promoting a lagging replica

When a primary fails and you promote a replica to take over, any writes that hadn't replicated yet are gone. If you're holding yourself to a high uptime target, promoting a replica that's noticeably behind risks losing writes in a way that eats into that availability budget fast, on top of the outage itself1.

Check current lag before promoting, not after, and use synchronous replication for the specific writes you genuinely can't afford to lose, such as a payment confirmation, rather than trying to make every write synchronous and paying that latency cost on every transaction regardless of how much it matters.

The mistake underneath the mistake: balancing by load, not by lag

A load balancer sitting in front of a replica pool that picks the least busy replica for each query sounds reasonable, but connection count and replication lag are different things. A replica can have few active connections and still be badly behind if it's fallen behind on WAL replay for other reasons, such as a long-running analytical query competing for the same I/O.

Route by measured lag, or at minimum exclude a replica once it crosses your alert threshold, instead of assuming low connection count means it's caught up. The two numbers move independently often enough that assuming one from the other will eventually send a sensitive read to the replica furthest behind.

Watching lag trend, not just its current value

A single lag reading tells you where a replica stands right now, not whether it's catching up or falling further behind. Graph it over time alongside your write volume, and you'll usually see lag climb during known heavy-write periods, such as a nightly batch import, and recover afterward.

A replica that doesn't recover after the heavy-write period ends is the pattern worth paging on. That's a structural capacity problem, not a transient blip, and it will keep getting worse as write volume grows unless the replica is resized or the write pattern that's causing it changes.

Executive Capability Standard

What Good Looks Like

Good replica-lag handling means you can state your current lag in concrete terms, alert before it becomes visible to users, and know exactly which reads are allowed to be stale.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Read your database engine's replication documentation and query its replication status view to see your current lag.
2. Do Manually:Manually check replica lag before any deploy that touches a write-heavy table, and route sensitive reads to the primary by hand until routing is automated.
3. Delegate:Give a database-focused engineer ownership of both the lag alerting thresholds and the read-routing rules as one piece of work.
4. Automate:Build lag-aware read routing into your connection pooler or data access layer so it applies per request, not per developer memory.
5. Buy:Bring in a database reliability consultant or move to a managed provider with built-in lag monitoring if incidents are outpacing your team's ability to instrument this.

How to Get Started

Frequently Asked Questions

How much replica lag is normal?

It depends entirely on your write volume and what your reads can tolerate, so there's no universal safe number. Set your threshold from what breaks for users, such as a just-saved record not appearing, rather than picking an arbitrary figure and assuming it's safe.

Can I force a read-after-write to always hit the primary?

Yes, and it's a common pattern: route a user's own session to the primary for a short window right after they write, then fall back to replicas afterward. Build this into your data access layer once so it applies consistently rather than depending on each engineer remembering it.

Does adding more replicas reduce lag?

Not directly. More replicas spread out read load, but each one still has to replay the same write-ahead log stream from the primary independently. A replica falls behind because of its own replay capacity or network path, and adding siblings doesn't change that for it.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides