Model Context Protocol & Agentic ArchitecturePlaybook3 min readUpdated September 2026

Diagnosing Stale Reads From a Lagging Read Replica

Stale reads from a read replica happen when a write lands on the primary but the next read is served by a replica that hasn't caught up, so the fix is measuring lag and routing lag-sensitive reads to the primary. For example, a user submits a form and the confirmation page shows the old data.

This walks through measuring lag correctly, finding what's actually causing it, and deciding which specific reads need to bypass the replica entirely rather than trying to eliminate lag everywhere.

The Symptom: A User Updates Something and Can't See It

The classic manifestation is a write followed immediately by a read of the same data, a profile update followed by the profile page reloading, a comment posted followed by the comment thread refreshing, where the read hits a replica that hasn't yet applied the write. The user sees stale data and reasonably assumes their action failed, even though the write succeeded fine on the primary.

This is specifically a read-after-write problem, not a general replica health problem. A replica can be perfectly healthy and still show this symptom under normal, expected replication delay, so the fix isn't necessarily reducing lag, it's routing this specific class of read differently.

Measuring Lag Correctly

A green replication status doesn't tell you how far behind the replica actually is in wall-clock time, only that replication hasn't broken outright. Measure lag directly: the difference between the primary's current write position and what the replica has actually applied, converted to a time delay, not just a byte or transaction count that's hard to reason about on its own.

Track this continuously, not just when someone reports a symptom. A replica that's usually a fraction of a second behind but occasionally spikes to many seconds behind under load has a different problem than one that's consistently moderately behind, and the fix for each is different.

The Usual Causes: Long Queries, Vacuum, and Network

A long-running query on the replica itself can block the application of incoming replication changes, depending on your replication configuration, creating lag that has nothing to do with the primary's write volume. Autovacuum running on the primary during a heavy write period can also slow replication indirectly by competing for the same I/O the replication stream needs.

Network latency and bandwidth between primary and replica matter more than most teams initially assume, especially for replicas in a different availability zone or region. If lag correlates cleanly with your write volume, the cause is likely query or vacuum related. If it correlates with network conditions instead, geographic placement or bandwidth is the more likely culprit.

Routing Reads That Can't Tolerate Any Lag

Rather than trying to eliminate lag everywhere, which usually means over-provisioning replicas for a problem that only affects a specific subset of reads, identify the read-after-write cases specifically and route just those to the primary. The common pattern: route a read to the primary for a short window immediately following that same session's write, then fall back to replicas for everything else.

This is a targeted fix, not a blanket one. Most reads in a typical application don't need read-after-write consistency, a dashboard, a report, a search result, and routing all of them to the primary just to solve the narrow confirmation-page case defeats the purpose of having replicas in the first place.

What to Alert On

Set alerts around the specific failure modes that actually matter, not just a single lag threshold:

  • Lag exceeding a threshold set from your own historical baseline, not a round number, since normal baseline lag varies meaningfully by workload and replica placement.
  • A sudden lag spike correlated with a specific deploy or migration, which usually points at a new expensive query rather than a general capacity problem.
  • Replication that's stopped entirely, which is a different and more urgent alert than elevated lag, since a stopped replica will eventually diverge enough to require a full resync.
  • Read-after-write routing failures, if you've built that routing layer, since a bug there silently reintroduces the exact stale-read symptom you built it to prevent.

Review the alert thresholds after any significant change to write volume or replica topology, since a threshold tuned for an earlier traffic pattern can go stale quietly.

Executive Capability Standard

What Good Looks Like

Good replica lag management means lag is measured as an actual time delay and tracked continuously, with only the specific reads that need read-after-write consistency routed around it.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Set up continuous measurement of actual replication lag in wall-clock time, not just a binary replication health check.
2. Do Manually:Identify the handful of user flows in your application where a read-after-write mismatch would actually be visible to a user, before building any routing logic.
3. Delegate:Have whoever owns the database layer own lag monitoring and alert thresholds, reviewed after any significant change in write volume.
4. Automate:Build a routing layer that sends a session's reads to the primary for a short window after that session's own write, falling back to replicas otherwise.
5. Buy:Use a managed Postgres provider with built-in read-after-write routing if your application has many scattered instances of this pattern, rather than building and maintaining the routing logic yourself.

How to Get Started

Frequently Asked Questions

Does a healthy replica mean lag isn't a problem?

Not necessarily. A replica can show healthy replication status while still being far enough behind in wall-clock time to cause visible stale reads. Measure lag as an actual time delay, not just a pass or fail health check, and track it continuously rather than only when someone reports a symptom.

Should we route all reads to the primary to avoid this problem entirely?

That defeats the purpose of having replicas. Most reads, dashboards, reports, search, don't need read-after-write consistency. Identify the specific reads that immediately follow a write in the same session and route just those to the primary, leaving everything else on replicas.

What usually causes a sudden lag spike after a deploy?

Most often a new query introduced by the deploy that's expensive enough to slow how fast the replica can apply incoming changes, or a migration running on the primary that competes with replication for I/O. Correlate lag spikes with deploy timestamps to confirm the cause before assuming it's a capacity issue.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides