Postgres Replica Lag: Where Reads Go Stale and How to Handle It
A read replica exists to take load off your primary database, and it works right up until the replica falls far enough behind that a user reads stale data immediately after writing it, like submitting a form and not seeing their own change on the next page.
Replica lag isn't a bug, it's a property of asynchronous replication, and the fix is routing reads intelligently, not eliminating lag entirely, since eliminating it means giving up the performance benefit replicas exist for.
What actually causes lag to grow
A replica applies changes from the primary's write-ahead log in order, and it falls behind when it can't apply changes as fast as the primary generates them: a burst of heavy writes, a long-running query on the replica blocking log application, or the replica simply running on less capable hardware than the primary.
Lag isn't constant, it spikes with write volume and with anything that competes with replication for the replica's resources, which is why monitoring a single average lag number hides the moments that actually cause user-visible problems.
Monitoring lag properly: track the spikes, not the average
Postgres's own replication status views give you real lag data, and the number to alert on is a sustained spike or a steady upward trend, not a momentary blip that resolves in a second or two.
A replica that's usually near zero but spikes during your nightly batch job is a different problem than one that's steadily drifting further behind every day, and the two need different fixes.
Deciding which reads can tolerate lag and which can't
- Read-after-write from the same user session, like a user reloading a page right after submitting a form, usually can't tolerate lag and should read from the primary or use session-level consistency tracking to route around a replica that hasn't caught up yet.
- Analytics, reporting, and dashboards almost always can tolerate some lag, since nobody's making a real-time decision off a dashboard number that's a moment stale.
- Anything involving a financial balance or inventory count needs the primary, or at minimum an explicit lag check before trusting a replica's answer, since acting on a stale number here has real consequences.
A pattern for routing reads without hardcoding every decision
Rather than deciding case by case in application code which queries go to the primary, tag queries by consistency requirement, strict, eventual, or analytics, at the point they're written, and let a data access layer route based on that tag plus the replica's current measured lag. This keeps the decision close to the code that actually knows what the read needs, instead of a separate team guessing at read patterns from the outside.
A worked example: a user who doesn't see their own update
Say a user updates their billing address and is immediately redirected to a confirmation page that reads the address back from a replica that hasn't caught up yet, showing the old address for a moment before a page refresh finally reflects the change. Support gets a ticket that reads like a data-loss bug even though nothing was actually lost, just briefly stale.
The fix isn't routing every read to the primary defensively. It's tagging that specific read, the address as read back to the same user immediately after their own write, as one that needs strict consistency, while every other address-related read elsewhere in the app keeps using the replica as intended.
When replica lag signals a capacity problem, not just a monitoring gap
Occasional lag spikes during a known batch job are a scheduling and resource-contention issue. Lag that's trending upward over weeks, even outside of any specific batch job, usually means write volume has outgrown what a single replica can keep pace with, and no amount of read routing sophistication fixes that; it needs either a beefier replica, a second replica to spread reads further, or a look at whether write volume itself can be reduced.
Distinguishing the two matters because they call for different people to get involved. A batch-job spike is usually a scheduling fix an application engineer can make alone, shifting the job to a quieter window or breaking it into smaller chunks. A steady upward trend is a capacity conversation that needs whoever owns the database infrastructure at the table, since the fix is likely to cost real money or real re-architecture, not just a config change.
What Good Looks Like
A good replica lag strategy monitors for spikes and trends, not just an average, and routes reads to the primary or a replica based on how much staleness that specific read can actually tolerate.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How much replica lag is acceptable?
It depends entirely on what the read is for. A read-after-write from the same user session usually can't tolerate more than a moment; an analytics dashboard can usually tolerate some delay without anyone noticing. Set your alert threshold per use case, not one number for every replica-backed query.
Why does lag spike during our nightly batch job specifically?
A long-running query or heavy write burst competes with the replica for the same resources it needs to apply changes from the primary's write-ahead log. If lag reliably spikes at the same time as a known batch job, that's usually resource contention on the replica, not a general capacity problem.
Should read-after-write queries always go to the primary?
For anything where a user would notice stale data immediately, like reloading a page right after submitting a form, yes, or use session-level consistency tracking that routes to the primary only until the replica has caught up. Sending every read-after-write query to the primary permanently defeats the point of having a replica.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Read Replica Lag: Catching It Before Customers Do
A user updates their profile, reloads, and sees the old data because the read hit a lagging replica. How to monitor lag and route around it.
Build Your Own Replica Lag Guardrails, or Buy a Managed One?
Whether to build custom replication lag monitoring and read routing yourself or rely on a managed database's built-in guardrails, and how to decide.
Debugging Stale Reads From a Postgres Replica
A walkthrough of why read replicas fall behind, how to measure lag correctly, and the read-your-own-write pattern that fixes the most common symptom.
Read Replica Lag: The Bugs It Causes and How to Route Around Them
A checklist for diagnosing replication lag causes, fixing read-after-write bugs, and deciding between synchronous and asynchronous replication.
Build vs. Buy for Handling Postgres Replica Lag Safely
A build-versus-buy guide to handling read replica lag in Postgres, covering when a simple wait-and-check approach is enough and when you need more.
Diagnosing Stale Reads From a Lagging Read Replica
A troubleshooting walkthrough for stale reads from a lagging replica: how to measure lag, find the cause, and route reads that can't tolerate it.