Diagnosing Stale Reads From a Lagging Read Replica
Stale reads from a read replica happen when a write lands on the primary but the next read is served by a replica that hasn't caught up, so the fix is measuring lag and routing lag-sensitive reads to the primary. For example, a user submits a form and the confirmation page shows the old data.
This walks through measuring lag correctly, finding what's actually causing it, and deciding which specific reads need to bypass the replica entirely rather than trying to eliminate lag everywhere.
The Symptom: A User Updates Something and Can't See It
The classic manifestation is a write followed immediately by a read of the same data, a profile update followed by the profile page reloading, a comment posted followed by the comment thread refreshing, where the read hits a replica that hasn't yet applied the write. The user sees stale data and reasonably assumes their action failed, even though the write succeeded fine on the primary.
This is specifically a read-after-write problem, not a general replica health problem. A replica can be perfectly healthy and still show this symptom under normal, expected replication delay, so the fix isn't necessarily reducing lag, it's routing this specific class of read differently.
Measuring Lag Correctly
A green replication status doesn't tell you how far behind the replica actually is in wall-clock time, only that replication hasn't broken outright. Measure lag directly: the difference between the primary's current write position and what the replica has actually applied, converted to a time delay, not just a byte or transaction count that's hard to reason about on its own.
Track this continuously, not just when someone reports a symptom. A replica that's usually a fraction of a second behind but occasionally spikes to many seconds behind under load has a different problem than one that's consistently moderately behind, and the fix for each is different.
The Usual Causes: Long Queries, Vacuum, and Network
A long-running query on the replica itself can block the application of incoming replication changes, depending on your replication configuration, creating lag that has nothing to do with the primary's write volume. Autovacuum running on the primary during a heavy write period can also slow replication indirectly by competing for the same I/O the replication stream needs.
Network latency and bandwidth between primary and replica matter more than most teams initially assume, especially for replicas in a different availability zone or region. If lag correlates cleanly with your write volume, the cause is likely query or vacuum related. If it correlates with network conditions instead, geographic placement or bandwidth is the more likely culprit.
Routing Reads That Can't Tolerate Any Lag
Rather than trying to eliminate lag everywhere, which usually means over-provisioning replicas for a problem that only affects a specific subset of reads, identify the read-after-write cases specifically and route just those to the primary. The common pattern: route a read to the primary for a short window immediately following that same session's write, then fall back to replicas for everything else.
This is a targeted fix, not a blanket one. Most reads in a typical application don't need read-after-write consistency, a dashboard, a report, a search result, and routing all of them to the primary just to solve the narrow confirmation-page case defeats the purpose of having replicas in the first place.
What to Alert On
Set alerts around the specific failure modes that actually matter, not just a single lag threshold:
- Lag exceeding a threshold set from your own historical baseline, not a round number, since normal baseline lag varies meaningfully by workload and replica placement.
- A sudden lag spike correlated with a specific deploy or migration, which usually points at a new expensive query rather than a general capacity problem.
- Replication that's stopped entirely, which is a different and more urgent alert than elevated lag, since a stopped replica will eventually diverge enough to require a full resync.
- Read-after-write routing failures, if you've built that routing layer, since a bug there silently reintroduces the exact stale-read symptom you built it to prevent.
Review the alert thresholds after any significant change to write volume or replica topology, since a threshold tuned for an earlier traffic pattern can go stale quietly.
What Good Looks Like
Good replica lag management means lag is measured as an actual time delay and tracked continuously, with only the specific reads that need read-after-write consistency routed around it.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Does a healthy replica mean lag isn't a problem?
Not necessarily. A replica can show healthy replication status while still being far enough behind in wall-clock time to cause visible stale reads. Measure lag as an actual time delay, not just a pass or fail health check, and track it continuously rather than only when someone reports a symptom.
Should we route all reads to the primary to avoid this problem entirely?
That defeats the purpose of having replicas. Most reads, dashboards, reports, search, don't need read-after-write consistency. Identify the specific reads that immediately follow a write in the same session and route just those to the primary, leaving everything else on replicas.
What usually causes a sudden lag spike after a deploy?
Most often a new query introduced by the deploy that's expensive enough to slow how fast the replica can apply incoming changes, or a migration running on the primary that competes with replication for I/O. Correlate lag spikes with deploy timestamps to confirm the cause before assuming it's a capacity issue.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Read Replica Lag: Catching It Before Customers Do
A user updates their profile, reloads, and sees the old data because the read hit a lagging replica. How to monitor lag and route around it.
The Stale Read Bug Replication Lag Causes, and the Fix
Why Postgres read replicas fall behind the primary, the stale read bug that shows up right after a write, and when lag means you've outgrown one primary.
Build Your Own Replica Lag Guardrails, or Buy a Managed One?
Whether to build custom replication lag monitoring and read routing yourself or rely on a managed database's built-in guardrails, and how to decide.
Debugging Stale Reads From a Postgres Replica
A walkthrough of why read replicas fall behind, how to measure lag correctly, and the read-your-own-write pattern that fixes the most common symptom.
Postgres Replica Lag: Where Reads Go Stale and How to Handle It
Why Postgres replicas fall behind under write-heavy load, how to monitor lag properly, and which read patterns need the primary instead.
Read Replica Lag: The Bugs It Causes and How to Route Around Them
A checklist for diagnosing replication lag causes, fixing read-after-write bugs, and deciding between synchronous and asynchronous replication.