Developer Productivity & Platform EngineeringPlaybook3 min readUpdated September 2026

Read Replica Lag: Catching It Before Customers Do

A user updates their profile, reloads the page immediately, and sees the old data, because the read landed on a replica that hadn't caught up to the write yet. Nothing is actually broken; asynchronous replication just means the primary doesn't wait for a replica to confirm before acknowledging a write.

The fix isn't eliminating lag entirely, which is rarely practical. It's monitoring it directly and routing the reads that can't tolerate it back to the primary.

Why Replicas Lag At All

Asynchronous replication means the primary acknowledges a write and moves on without waiting for a replica to apply it. Lag grows under write-heavy load, when a long-running query on the replica itself competes for the same resources the replication stream needs, or during a network blip between primary and replica. It's a normal, expected behavior of the architecture, not a malfunction.

The Read-Your-Own-Writes Problem

The most common symptom is a user not seeing their own recent change reflected back to them. The practical fix most teams reach for isn't eliminating lag, it's routing the read immediately following a write, for that same user or session, back to the primary for a short window, then letting subsequent reads fall back to a replica once enough time has passed.

Monitoring Lag as a First-Class Metric

Track replication lag directly, in seconds or bytes behind the primary, rather than inferring database health from replica CPU alone. Alert on a threshold tied to what your application can actually tolerate for each read path, not an arbitrary round number picked because it looked reasonable at the time.

Deciding Which Reads Can Tolerate Lag

  • A public read-only dashboard or an internal analytics query can usually tolerate several seconds of lag with no noticeable user impact.
  • Anything reading data right after that same user's own write needs the primary, or an explicit read-your-writes routing rule, at least for a short window.
  • A background report or export job can typically tolerate minutes of lag without anyone noticing at all.

For example, an e-commerce team might route the order confirmation page to the primary for a short window after checkout, because a customer who just paid expects to see the order immediately. The product catalog, by contrast, can read from any replica without anyone noticing a delay. A common mistake is treating the whole application as one read path with a single lag tolerance. The fix is to list your read paths, label each one as tolerant or intolerant of stale data, and route accordingly. Revisit that list whenever a new feature ships, since a new read that follows a write is easy to miss.

When to Add More Replicas vs Fix the Query Causing Lag

A single slow analytical query running against a replica can itself be the cause of lag, by holding locks or consuming the I/O bandwidth the replication stream needs to keep up. Before provisioning more replica capacity to solve a lag problem, check whether one specific query is the actual culprit; moving that query off the affected replica often fixes the lag faster and cheaper than scaling out.

Isolating a Replica for Heavy Analytical Queries

A dedicated replica set aside specifically for long-running reporting or analytical queries, kept separate from the replicas serving your application's normal read traffic, contains the damage a single expensive query can do. If that reporting replica falls behind, it only affects reports, not the read paths customers actually depend on for a responsive product experience.

This separation also makes lag easier to reason about, since you can set a more relaxed lag threshold on the analytical replica, where staleness of a few minutes is genuinely fine, and a tighter one on the application-facing replicas, instead of trying to find one threshold that satisfies every workload at once.

A Quick Diagnostic When Lag Spikes Unexpectedly

Check three things in order: whether write volume on the primary actually increased, whether a specific long-running query is active on the lagging replica right now, and whether there's a network issue between the primary and that replica specifically rather than a general problem. Most unexpected lag spikes trace back to one of these three, and knowing which one saves a lot of time compared to guessing at a fix.

Keep a short runbook for this diagnostic, not just the knowledge in one engineer's head. Lag spikes tend to happen at inconvenient times, and whoever is on call when it fires benefits from a written starting point rather than reconstructing the same three checks from scratch under pressure.

Executive Capability Standard

What Good Looks Like

Good replica lag management means lag is monitored directly as its own metric, and the reads that genuinely can't tolerate it, like a user's own recent write, are routed to the primary instead of quietly showing stale data.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Check whether replication lag is currently monitored as its own metric, separate from general replica health checks.
2. Do Manually:Manually identify the two or three read paths most likely to show a user their own stale write and route those to the primary.
3. Delegate:Assign one engineer to own lag alert thresholds so they reflect what each read path can actually tolerate.
4. Automate:Automate routing for read-your-writes cases so it doesn't depend on every engineer remembering the rule for every new endpoint.
5. Buy:Bring in a database specialist to review replica sizing if a single slow query is a recurring cause of lag spikes.

How to Get Started

Frequently Asked Questions

Can you eliminate replication lag entirely?

Not in a purely asynchronous replication setup, and forcing synchronous replication for every write usually costs more in write latency than it's worth for most applications. The practical approach is monitoring lag and routing the specific reads that can't tolerate it, rather than chasing zero lag everywhere.

How do you handle a user who wants to see their own update immediately?

Route the read that immediately follows a write, for that specific user or session, back to the primary for a short window rather than a replica. Once enough time has passed for replication to reasonably catch up, subsequent reads can safely fall back to a replica without the user noticing any inconsistency.

What's a reasonable replication lag alert threshold?

It depends entirely on what your application's read paths can tolerate, not a generic number. A dashboard that's fine with several seconds of staleness needs a very different threshold than a read path serving data right after a user's own write. Set the threshold per use case, not once for the whole database.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides