Cloud FinOps & Infrastructure ScalingPlaybook3 min readUpdated September 2026

Debugging Stale Reads From a Postgres Replica

The bug report usually reads the same way: a user updates their profile, refreshes the page, and sees their old data. Nothing is actually broken. The write went to the primary database successfully. The read, a moment later, hit a replica that hasn't caught up yet, and replication lag, not a real bug, is what the user is seeing.

Understanding why lag happens and how to measure it correctly is most of the way to fixing it.

Why Replicas Fall Behind in the First Place

A replica applies changes from the primary's write-ahead log in order, and it can only apply them as fast as its own hardware and current load allow. Under normal traffic, this happens fast enough that lag is negligible, often well under a second. Under a burst of heavy writes, a large batch import, a bulk update, or just elevated traffic, the replica's apply process can fall behind the rate new changes arrive, and the gap between primary and replica grows until write volume drops back down enough for the replica to catch up.

Measuring Lag the Right Way

A lot of teams check lag using byte-based metrics, how far behind the replica is in write-ahead log position, which is useful for diagnosing replication health but doesn't directly tell you what a user experiences. What actually matters for the stale-read symptom is time-based lag: how many seconds behind the primary the replica's most recently applied transaction is. Most managed database platforms expose this directly; if yours doesn't, you can query it by comparing the replica's last replayed transaction timestamp against the current time. Alert on the time-based number, since that's the one that maps to what a user actually notices.

The Fix Most Teams Reach for First: Read Your Own Writes

The most common fix isn't eliminating lag entirely, it's routing around it for the specific case that causes visible bugs: a user reading data they themselves just wrote. After a write, route that same user's next read for the affected data to the primary, or to whichever replica you can confirm has caught up, instead of a random replica that might not have. This targeted approach avoids sending your entire read load to the primary, which would defeat the point of having replicas, while still fixing the symptom that actually generates support tickets.

A Worked Example of the Read-Your-Own-Writes Pattern

Say a user updates their shipping address at checkout, and the confirmation page immediately queries for their profile to display it back to them. If that read hits a lagging replica, the confirmation page shows the old address, which looks like the update silently failed even though it didn't. Tagging that specific post-write read to go to the primary, just for a short window after the user's own write, fixes exactly this case without changing where the rest of your read traffic goes. The pattern only needs to apply to reads that could plausibly reflect a user's own very recent write, not to all reads generally.

When Lag Points to a Deeper Capacity Problem

Occasional, brief lag spikes during traffic bursts are normal and usually self-correct. Lag that's persistently elevated, or that keeps growing instead of recovering between bursts, is a signal that your replica's hardware can't keep up with your actual write volume, not a transient blip to route around. At that point, the fix isn't smarter routing, it's addressing the underlying capacity: a faster replica instance, reducing unnecessary write volume, or, if you're running many replicas, checking whether something specific to one replica, not the write volume itself, is the actual cause.

Don't Overlook Long-Running Queries on the Replica Itself

A replica isn't just applying writes, it's usually also serving read queries, and a slow analytical query holding a long-running transaction on the replica can itself delay how quickly that replica can apply incoming changes from the primary. If your lag spikes correlate more closely with a scheduled reporting job than with write volume on the primary, the fix is isolating that workload, moving heavy analytical queries to a dedicated replica reserved for that purpose, rather than tuning anything about replication itself. Checking what else is running on a lagging replica, not just how fast it's replicating, catches a cause that pure replication metrics won't show you.

When users report stale reads, work through these checks:

  1. Measure time-based lag in seconds behind the primary, not only byte position in the write-ahead log, since seconds match what users experience.
  2. Route a user's next read after their own write to the primary, or to a replica you can confirm has caught up.
  3. Check whether lag stays elevated or keeps growing between bursts, which points to a replica that cannot keep up with your write volume.
  4. Look for long-running analytical queries on the replica that delay applying changes, and isolate that reporting work elsewhere.
Executive Capability Standard

What Good Looks Like

A healthy replica setup measures lag in time, not just log position, routes a user's own post-write reads around lag automatically, and treats persistently growing lag as a capacity signal rather than something to route around indefinitely.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Check whether your monitoring currently tracks time-based replication lag, not just byte-based log position, for every replica.
2. Do Manually:Manually route the highest-complaint post-write read path, like a profile update confirmation, to the primary as a quick fix.
3. Delegate:Have a backend engineer implement a general read-your-own-writes pattern for any read that could reflect a very recent write.
4. Automate:Add automated alerting on time-based replica lag so a growing gap gets caught before it generates support tickets.
5. Buy:Bring in a database specialist if lag persists after routing fixes are in place, since that usually points to a capacity problem needing infrastructure changes, not more routing logic.

How to Get Started

Frequently Asked Questions

How much replication lag is considered normal?

Sub-second lag under normal traffic is typical for most workloads. A few seconds during a heavy write burst is common and usually self-corrects once the burst passes. Lag measured in minutes, or lag that keeps climbing rather than recovering, points to a genuine capacity problem rather than normal variance.

Should we just send every read to the primary to avoid this entirely?

That eliminates the stale-read symptom but also eliminates the reason you added replicas: spreading read load so the primary isn't a bottleneck. Targeted read-your-own-writes routing fixes the specific symptom that causes visible bugs without giving up the capacity benefit replicas provide for the rest of your traffic.

Can application-level caching make replication lag worse?

It can mask it in a confusing way rather than making the underlying lag worse. If a cache serves stale data independently of replica lag, you can end up debugging two overlapping staleness sources at once. Keep cache invalidation and replica-lag handling as separate, clearly understood layers so a stale-data bug report doesn't send you down the wrong one.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides