AI Model Serving & Inference OptimizationPlaybook3 min readUpdated September 2026

The Stale Read Bug Replication Lag Causes, and the Fix

A read replica is supposed to be invisible to the user, a way to scale reads without anyone noticing which database actually answered the query. Replication lag is what breaks that illusion: a user submits something, the page redirects, and the read that follows lands on a replica that hasn't caught up yet.

The fix isn't to stop using replicas. It's knowing which reads actually need to see a write immediately, and routing only those around the lag.

Why a replica falls behind in the first place

Standard Postgres replication streams write-ahead log records to the replica asynchronously, and the replica applies them in order. Under normal write volume that keeps up easily. Under a burst of writes, a long-running query on the replica that blocks WAL apply, or extra network distance for a cross-region replica, the gap between primary and replica widens.

None of those causes are exotic, which is exactly why lag is easy to ignore until it isn't. A replica that's usually caught up within milliseconds can fall seconds or more behind during a batch job or a traffic spike, and if nothing in your application accounts for that, it will show stale data at exactly the moment users are most active.

The bug: read your own write, except you can't

The classic version of this bug goes: a user fills out a form, the app writes to the primary, redirects to a confirmation or detail page, and that page reads from a replica that hasn't replayed the write yet. The user sees their own submission missing, or worse, an old version of the record they just changed.

It's a confusing bug to debug from a support ticket, because it's intermittent by nature, it only shows up when lag happens to be nonzero at that exact moment, which means it often passes every manual test a developer runs slowly enough to give the replica time to catch up.

Read your writes without sending everyone to the primary

The simplest fix is sticky routing: after a user writes something, route that user's next few reads to the primary for a short window, then let them fall back to replicas. It's crude but effective, and it doesn't require tracking anything more precise than "this user just wrote."

A more precise approach tracks the log sequence number of the write and has the read wait until a replica reports it has caught up past that point, or falls back to the primary if it hasn't within a short timeout. That's more work to build, but it scopes the primary traffic to genuinely affected reads instead of an entire user session.

Three practical ways to protect a user's read after their own write:

  • Use sticky routing: after a user writes something, send that user's next few reads to the primary for a short window, then let them fall back to replicas.
  • Track the log sequence number of the write and make the follow-up read wait until a replica has replayed past it, which is more precise than a time window.
  • Send only the reads that truly need to see a just-made write to the primary, and leave everything else on replicas so the primary is not overloaded.

Watching lag before your users do

Postgres exposes replication lag as a metric you can query directly, and it belongs on a dashboard with an alert threshold, not something you only think to check after a stale-read bug report comes in. A lag that's normally near zero and briefly spikes during a known batch job is expected. A lag that's slowly, steadily growing under steady-state load is a different signal entirely.

That second pattern means the replica can't keep up with sustained write volume, not a transient blip, and it tends to get worse, not better, as the write volume grows. Catching that trend early is the difference between a planned change and an emergency one.

When lag means it's time to shard, not tune

If lag keeps growing despite reasonable tuning, shorter transactions on the replica, adjusted checkpoint settings, more replica capacity, and the root cause is sustained write volume rather than a one-off event, that's the database topology telling you it's outgrown a single primary with replicas hanging off it.

At that point the conversation shifts from tuning replication to partitioning writes, whether that's sharding by tenant, splitting write-heavy tables onto their own database, or a managed horizontally scalable database built for that pattern. Adding another replica doesn't fix a write-volume problem, since every replica has to apply the same growing stream of writes regardless of how many of them exist.

Executive Capability Standard

What Good Looks Like

Good replica handling means a user can write something and immediately see it reflected, without every read in your system being forced onto the primary just to be safe.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Learn how your database exposes replication lag as a metric, and confirm you're actually watching it instead of assuming replicas are always caught up.
2. Do Manually:Add a manual rule that routes reads to the primary right after a write for the handful of screens where a stale read is visibly wrong.
3. Delegate:Have someone own a small read-routing layer that decides primary versus replica per query, instead of leaving that judgment call scattered across the codebase.
4. Automate:Automate replication lag alerts tied to the queries most likely to show a stale read, so a growing lag pages someone before users notice.
5. Buy:If write volume keeps outrunning a single primary despite tuning, that's a case for a managed, horizontally scalable database rather than adding more replicas.

How to Get Started

Frequently Asked Questions

How much replication lag is normal?

It depends heavily on write volume and network distance to the replica. What matters more than any single number is the trend: lag that briefly spikes during a known event and recovers is normal, lag that grows steadily under steady load is not.

Can I just always read from the primary to avoid this?

You can, but it defeats the purpose of having read replicas, which is scaling read capacity away from the primary. A targeted approach, routing only the reads that need to see a recent write immediately, keeps most of that benefit.

Does synchronous replication fix this?

It removes the lag problem, since the primary waits for the replica to confirm before committing, but it does so at the cost of write latency and availability if that replica becomes unreachable, a real tradeoff, not a free upgrade.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides