Engineering Leadership & Technical HiringPlaybook3 min readUpdated September 2026

Read Replica Lag: The Bugs It Causes and How to Route Around Them

A user updates their profile, refreshes the page, and sees their old data because the read that served the refresh hit a replica that hadn't caught up yet. This is the single most common bug read replicas introduce, and it's rarely caused by replication being broken, it's caused by nobody accounting for the lag that's normal and expected.

Here's a checklist for finding the actual cause of lag in your setup and routing around the bugs it creates.

Know Your Actual Lag in Milliseconds, Not "Replication Is On"

Most managed database platforms expose a replication lag metric; if you're not graphing it and alerting on a threshold, you don't actually know whether your replicas are typically milliseconds or seconds behind, and that number determines how big a problem read-after-write inconsistency actually is for your application. Start here before making any architectural changes: a lag problem measured in single-digit milliseconds needs a very different fix than one measured in multiple seconds.

Long-Running Transactions on the Primary Are a Common Root Cause

A long transaction on the primary, a slow batch update, a reporting query that shouldn't be running there in the first place, can hold locks or generate a burst of write-ahead log that the replica has to work through sequentially, widening lag well beyond its normal baseline. Postgres replicas apply the write-ahead log largely in order on a single thread, so a burst of writes from one long transaction creates a backlog the replica has to work through before it catches up, even if the replica's hardware is otherwise idle.

Route a User's Own Writes Back to the Primary Temporarily

The most reliable fix for the refresh-and-see-stale-data bug is routing reads for a short window immediately following a user's own write back to the primary, rather than trying to guarantee low enough lag that it never matters. This can be implemented with a short-lived cookie or session flag set after a write, checked by the routing layer to decide primary versus replica for that user's next few reads, and it sidesteps the problem without requiring synchronous replication everywhere.

Synchronous Replication Trades Latency for Consistency, Deliberately

Synchronous replication, where a write doesn't commit until at least one replica confirms it received it, eliminates the lag problem for reads from that replica but adds write latency and a new failure mode: what happens to writes if the synchronous replica becomes unreachable. This tradeoff is usually only worth it for a small subset of your data, financial transactions, anything where stale reads are genuinely unacceptable, rather than applied uniformly across an entire database, where the added write latency costs more than it's worth for most tables.

Watch for Lag During Deploys and Schema Migrations Specifically

A schema migration or a large backfill on the primary is one of the most common times lag spikes unexpectedly, since it generates a burst of write-ahead log volume outside normal traffic patterns. Treat replica lag as a metric to watch during any migration rollout the same way you'd watch error rate during a deploy, and consider pausing read traffic to replicas or routing critical reads to the primary temporarily during a known high-write-volume operation.

Load Balancing Reads Across Multiple Replicas Adds Its Own Wrinkle

If you're spreading read traffic across more than one replica, each one can lag by a different amount depending on its own hardware and network path, which means a user's request can hit a more current replica on one page load and a more stale one on the next. Monitor lag per replica individually, not just an average across the fleet, since an average can look fine while one specific replica is meaningfully behind and quietly serving stale data to whichever requests happen to land on it.

Work through these checks when lag causes stale reads:

  • Graph replication lag and alert on a threshold, so you know whether replicas run milliseconds or seconds behind.
  • Keep long-running transactions and reporting queries off the primary, since they can widen lag.
  • Route a user's reads back to the primary for a short window after their own write.
  • Watch lag during deploys, schema migrations, and large backfills.
  • Compare lag across replicas if you balance reads across several, and treat persistent, growing lag as a possible capacity problem.

When Lag Is a Symptom of a Deeper Capacity Problem

Persistent, growing lag that doesn't correlate with any specific migration or batch job is often a sign the replica's own hardware can't keep up with the primary's steady-state write volume, not a transient issue that will resolve on its own. Treat a lag trend that keeps climbing over weeks, rather than spiking and recovering, as a capacity planning problem, upgrading the replica's resources or reducing write volume, since routing logic alone can't fix a replica that's structurally falling further behind every day.

Executive Capability Standard

What Good Looks Like

You should know, in milliseconds, how far behind your replicas run right now, not just that replication is turned on.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Read your database's replication lag metric for a week before you route any traffic decisions based on it.
2. Do Manually:Manually verify a read-after-write bug report against actual replica lag before assuming it's an application code issue.
3. Delegate:Give one engineer ownership of the routing rules that decide which reads go to a replica versus the primary.
4. Automate:Route a user's own writes back to the primary automatically for a short window after they write.
5. Buy:Consider a managed database platform with built-in read-after-write routing once hand-rolled routing logic spans too many services.

How to Get Started

Frequently Asked Questions

How much replica lag is actually normal?

It varies by workload and hardware, but single-digit milliseconds under normal load is common for a well-provisioned setup. Consistently seeing lag in the hundreds of milliseconds or seconds under normal, non-migration traffic is worth investigating rather than treating as expected.

Should read-heavy analytics queries run against a replica or the primary?

A replica, almost always, and ideally a dedicated one used only for analytics. A long analytics query can contribute to lag or lock contention that affects other readers of the same replica, so keep it apart from application read traffic.

Is routing writes-then-reads back to the primary a permanent architectural pattern, or a workaround?

It's a pragmatic, durable pattern, not just a stopgap; most systems that use read replicas at scale use some version of this for read-after-write consistency rather than trying to eliminate lag entirely, which is rarely achievable and expensive to chase.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides