Debugging Stale Reads From a Postgres Replica
The bug report usually reads the same way: a user updates their profile, refreshes the page, and sees their old data. Nothing is actually broken. The write went to the primary database successfully. The read, a moment later, hit a replica that hasn't caught up yet, and replication lag, not a real bug, is what the user is seeing.
Understanding why lag happens and how to measure it correctly is most of the way to fixing it.
Why Replicas Fall Behind in the First Place
A replica applies changes from the primary's write-ahead log in order, and it can only apply them as fast as its own hardware and current load allow. Under normal traffic, this happens fast enough that lag is negligible, often well under a second. Under a burst of heavy writes, a large batch import, a bulk update, or just elevated traffic, the replica's apply process can fall behind the rate new changes arrive, and the gap between primary and replica grows until write volume drops back down enough for the replica to catch up.
Measuring Lag the Right Way
A lot of teams check lag using byte-based metrics, how far behind the replica is in write-ahead log position, which is useful for diagnosing replication health but doesn't directly tell you what a user experiences. What actually matters for the stale-read symptom is time-based lag: how many seconds behind the primary the replica's most recently applied transaction is. Most managed database platforms expose this directly; if yours doesn't, you can query it by comparing the replica's last replayed transaction timestamp against the current time. Alert on the time-based number, since that's the one that maps to what a user actually notices.
The Fix Most Teams Reach for First: Read Your Own Writes
The most common fix isn't eliminating lag entirely, it's routing around it for the specific case that causes visible bugs: a user reading data they themselves just wrote. After a write, route that same user's next read for the affected data to the primary, or to whichever replica you can confirm has caught up, instead of a random replica that might not have. This targeted approach avoids sending your entire read load to the primary, which would defeat the point of having replicas, while still fixing the symptom that actually generates support tickets.
A Worked Example of the Read-Your-Own-Writes Pattern
Say a user updates their shipping address at checkout, and the confirmation page immediately queries for their profile to display it back to them. If that read hits a lagging replica, the confirmation page shows the old address, which looks like the update silently failed even though it didn't. Tagging that specific post-write read to go to the primary, just for a short window after the user's own write, fixes exactly this case without changing where the rest of your read traffic goes. The pattern only needs to apply to reads that could plausibly reflect a user's own very recent write, not to all reads generally.
When Lag Points to a Deeper Capacity Problem
Occasional, brief lag spikes during traffic bursts are normal and usually self-correct. Lag that's persistently elevated, or that keeps growing instead of recovering between bursts, is a signal that your replica's hardware can't keep up with your actual write volume, not a transient blip to route around. At that point, the fix isn't smarter routing, it's addressing the underlying capacity: a faster replica instance, reducing unnecessary write volume, or, if you're running many replicas, checking whether something specific to one replica, not the write volume itself, is the actual cause.
Don't Overlook Long-Running Queries on the Replica Itself
A replica isn't just applying writes, it's usually also serving read queries, and a slow analytical query holding a long-running transaction on the replica can itself delay how quickly that replica can apply incoming changes from the primary. If your lag spikes correlate more closely with a scheduled reporting job than with write volume on the primary, the fix is isolating that workload, moving heavy analytical queries to a dedicated replica reserved for that purpose, rather than tuning anything about replication itself. Checking what else is running on a lagging replica, not just how fast it's replicating, catches a cause that pure replication metrics won't show you.
When users report stale reads, work through these checks:
- Measure time-based lag in seconds behind the primary, not only byte position in the write-ahead log, since seconds match what users experience.
- Route a user's next read after their own write to the primary, or to a replica you can confirm has caught up.
- Check whether lag stays elevated or keeps growing between bursts, which points to a replica that cannot keep up with your write volume.
- Look for long-running analytical queries on the replica that delay applying changes, and isolate that reporting work elsewhere.
What Good Looks Like
A healthy replica setup measures lag in time, not just log position, routes a user's own post-write reads around lag automatically, and treats persistently growing lag as a capacity signal rather than something to route around indefinitely.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How much replication lag is considered normal?
Sub-second lag under normal traffic is typical for most workloads. A few seconds during a heavy write burst is common and usually self-corrects once the burst passes. Lag measured in minutes, or lag that keeps climbing rather than recovering, points to a genuine capacity problem rather than normal variance.
Should we just send every read to the primary to avoid this entirely?
That eliminates the stale-read symptom but also eliminates the reason you added replicas: spreading read load so the primary isn't a bottleneck. Targeted read-your-own-writes routing fixes the specific symptom that causes visible bugs without giving up the capacity benefit replicas provide for the rest of your traffic.
Can application-level caching make replication lag worse?
It can mask it in a confusing way rather than making the underlying lag worse. If a cache serves stale data independently of replica lag, you can end up debugging two overlapping staleness sources at once. Keep cache invalidation and replica-lag handling as separate, clearly understood layers so a stale-data bug report doesn't send you down the wrong one.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Read Replica Lag: Catching It Before Customers Do
A user updates their profile, reloads, and sees the old data because the read hit a lagging replica. How to monitor lag and route around it.
Build Your Own Replica Lag Guardrails, or Buy a Managed One?
Whether to build custom replication lag monitoring and read routing yourself or rely on a managed database's built-in guardrails, and how to decide.
Read Replica Lag: The Bugs It Causes and How to Route Around Them
A checklist for diagnosing replication lag causes, fixing read-after-write bugs, and deciding between synchronous and asynchronous replication.
Postgres Replica Lag: Where Reads Go Stale and How to Handle It
Why Postgres replicas fall behind under write-heavy load, how to monitor lag properly, and which read patterns need the primary instead.
Build vs. Buy for Handling Postgres Replica Lag Safely
A build-versus-buy guide to handling read replica lag in Postgres, covering when a simple wait-and-check approach is enough and when you need more.
Diagnosing Stale Reads From a Lagging Read Replica
A troubleshooting walkthrough for stale reads from a lagging replica: how to measure lag, find the cause, and route reads that can't tolerate it.