Read Replica Lag: The Bugs It Causes and How to Route Around Them
A user updates their profile, refreshes the page, and sees their old data because the read that served the refresh hit a replica that hadn't caught up yet. This is the single most common bug read replicas introduce, and it's rarely caused by replication being broken, it's caused by nobody accounting for the lag that's normal and expected.
Here's a checklist for finding the actual cause of lag in your setup and routing around the bugs it creates.
Know Your Actual Lag in Milliseconds, Not "Replication Is On"
Most managed database platforms expose a replication lag metric; if you're not graphing it and alerting on a threshold, you don't actually know whether your replicas are typically milliseconds or seconds behind, and that number determines how big a problem read-after-write inconsistency actually is for your application. Start here before making any architectural changes: a lag problem measured in single-digit milliseconds needs a very different fix than one measured in multiple seconds.
Long-Running Transactions on the Primary Are a Common Root Cause
A long transaction on the primary, a slow batch update, a reporting query that shouldn't be running there in the first place, can hold locks or generate a burst of write-ahead log that the replica has to work through sequentially, widening lag well beyond its normal baseline. Postgres replicas apply the write-ahead log largely in order on a single thread, so a burst of writes from one long transaction creates a backlog the replica has to work through before it catches up, even if the replica's hardware is otherwise idle.
Route a User's Own Writes Back to the Primary Temporarily
The most reliable fix for the refresh-and-see-stale-data bug is routing reads for a short window immediately following a user's own write back to the primary, rather than trying to guarantee low enough lag that it never matters. This can be implemented with a short-lived cookie or session flag set after a write, checked by the routing layer to decide primary versus replica for that user's next few reads, and it sidesteps the problem without requiring synchronous replication everywhere.
Synchronous Replication Trades Latency for Consistency, Deliberately
Synchronous replication, where a write doesn't commit until at least one replica confirms it received it, eliminates the lag problem for reads from that replica but adds write latency and a new failure mode: what happens to writes if the synchronous replica becomes unreachable. This tradeoff is usually only worth it for a small subset of your data, financial transactions, anything where stale reads are genuinely unacceptable, rather than applied uniformly across an entire database, where the added write latency costs more than it's worth for most tables.
Watch for Lag During Deploys and Schema Migrations Specifically
A schema migration or a large backfill on the primary is one of the most common times lag spikes unexpectedly, since it generates a burst of write-ahead log volume outside normal traffic patterns. Treat replica lag as a metric to watch during any migration rollout the same way you'd watch error rate during a deploy, and consider pausing read traffic to replicas or routing critical reads to the primary temporarily during a known high-write-volume operation.
Load Balancing Reads Across Multiple Replicas Adds Its Own Wrinkle
If you're spreading read traffic across more than one replica, each one can lag by a different amount depending on its own hardware and network path, which means a user's request can hit a more current replica on one page load and a more stale one on the next. Monitor lag per replica individually, not just an average across the fleet, since an average can look fine while one specific replica is meaningfully behind and quietly serving stale data to whichever requests happen to land on it.
Work through these checks when lag causes stale reads:
- Graph replication lag and alert on a threshold, so you know whether replicas run milliseconds or seconds behind.
- Keep long-running transactions and reporting queries off the primary, since they can widen lag.
- Route a user's reads back to the primary for a short window after their own write.
- Watch lag during deploys, schema migrations, and large backfills.
- Compare lag across replicas if you balance reads across several, and treat persistent, growing lag as a possible capacity problem.
When Lag Is a Symptom of a Deeper Capacity Problem
Persistent, growing lag that doesn't correlate with any specific migration or batch job is often a sign the replica's own hardware can't keep up with the primary's steady-state write volume, not a transient issue that will resolve on its own. Treat a lag trend that keeps climbing over weeks, rather than spiking and recovering, as a capacity planning problem, upgrading the replica's resources or reducing write volume, since routing logic alone can't fix a replica that's structurally falling further behind every day.
What Good Looks Like
You should know, in milliseconds, how far behind your replicas run right now, not just that replication is turned on.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How much replica lag is actually normal?
It varies by workload and hardware, but single-digit milliseconds under normal load is common for a well-provisioned setup. Consistently seeing lag in the hundreds of milliseconds or seconds under normal, non-migration traffic is worth investigating rather than treating as expected.
Should read-heavy analytics queries run against a replica or the primary?
A replica, almost always, and ideally a dedicated one used only for analytics. A long analytics query can contribute to lag or lock contention that affects other readers of the same replica, so keep it apart from application read traffic.
Is routing writes-then-reads back to the primary a permanent architectural pattern, or a workaround?
It's a pragmatic, durable pattern, not just a stopgap; most systems that use read replicas at scale use some version of this for read-after-write consistency rather than trying to eliminate lag entirely, which is rarely achievable and expensive to chase.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Read Replica Lag: Catching It Before Customers Do
A user updates their profile, reloads, and sees the old data because the read hit a lagging replica. How to monitor lag and route around it.
Build Your Own Replica Lag Guardrails, or Buy a Managed One?
Whether to build custom replication lag monitoring and read routing yourself or rely on a managed database's built-in guardrails, and how to decide.
Debugging Stale Reads From a Postgres Replica
A walkthrough of why read replicas fall behind, how to measure lag correctly, and the read-your-own-write pattern that fixes the most common symptom.
Living With Replication Lag Instead of Pretending It Doesn't Exist
A comparison of ways to handle Postgres read replica lag, from routing reads by freshness requirement to synchronous replication, and their real tradeoffs.
Build vs. Buy for Handling Postgres Replica Lag Safely
A build-versus-buy guide to handling read replica lag in Postgres, covering when a simple wait-and-check approach is enough and when you need more.
Postgres Replica Lag: Where Reads Go Stale and How to Handle It
Why Postgres replicas fall behind under write-heavy load, how to monitor lag properly, and which read patterns need the primary instead.