Production RAG & Vector Data ArchitecturePlaybook3 min readUpdated September 2026

Build vs. Buy for Handling Postgres Replica Lag Safely

Postgres read replica lag is the delay between a write on the primary and that write appearing on a replica, and it becomes a visible bug when a user reads their own change back from a replica that hasn't caught up. The fix doesn't require abandoning replicas. It requires deciding how much engineering effort read-after-write consistency is worth.

The simplest fix: route read-after-write to the primary

For the specific case of a user reading their own just-written data, the simplest and often sufficient fix is routing that particular read to the primary database instead of a replica, for a short window after the write. This avoids the whole class of lag-related bugs for the case that actually matters most to users, at the cost of some extra load on the primary, which is usually a fine tradeoff since it only applies to a narrow set of reads. Implementing this rarely requires new infrastructure, just a routing decision your existing data access layer can make based on whether a write just happened.

A step up: check replica lag before routing to it

Most database systems expose a lag metric you can query, and a slightly more sophisticated approach checks that metric before deciding whether a given read can safely go to a replica, falling back to the primary if lag exceeds an acceptable threshold. This handles more cases than the simple read-after-write fix alone, including a replica that's fallen unusually far behind for reasons unrelated to any specific user's own write, but it adds a check on the read path that has its own latency cost, which is worth measuring rather than assuming away.

A common mistake with the lag check is choosing the threshold by guesswork. For example, a team might route to the primary whenever lag exceeds a small fraction of a second, then discover that during a busy write burst nearly every read falls back, so the primary absorbs the load the replicas were meant to take. Start by measuring your typical lag under normal and peak write volume, set the threshold just above what normal looks like, and watch how often reads fall back. If fallbacks are frequent, treat that as a signal about replica capacity rather than raising the threshold until the alarm goes quiet.

The heavier option: causal consistency tracking

For applications where read-after-write correctness matters across many different flows, not just the obvious ones, some teams build or adopt a system that tracks a causal token representing what a client has already seen, and routes reads to whichever replica has caught up to that token. This is meaningfully more correct across a wider range of cases, and it's also real infrastructure to build, test, and maintain, which is a cost worth weighing honestly against how much the simpler fixes above would actually leave uncovered.

When the simple fix is genuinely enough

For most products, the specific case of a user seeing their own recent write is the overwhelming majority of real-world lag complaints, and the primary-routing fix above handles it directly. Before investing in a heavier consistency system, check your actual support tickets or bug reports for evidence that a broader class of lag-related inconsistency is actually costing you something, rather than building for a theoretical case nobody has reported.

When it's worth paying for more

If your product coordinates across services that each read from their own replica, and inconsistency between those views causes a genuinely costly problem, such as two parts of an order processing flow seeing different states of the same order, the heavier causal consistency approach earns its cost. The signal to watch for is a bug that a simple read-after-write fix wouldn't have caught, because it involves more than one client's own recent write. Write that signal down explicitly before starting the project, so the team building it has a concrete test for whether the investment actually paid off.

Work through the options in this order, and stop as soon as the problem is solved:

  1. Route a user's read of their own just-written data to the primary for a short window, using a rule in your existing data access layer.
  2. Check the replica's lag metric before sending a read to it, and fall back to the primary when lag exceeds your tolerance, measuring the latency cost of that check.
  3. Review support tickets and bug reports for evidence that lag inconsistencies go beyond a user's own recent write before building anything heavier.
  4. Adopt causal consistency tracking only when several services reading from separate replicas cause costly disagreements that a simple routing rule would miss.

A worked example: the profile update that looked reverted

Say a user updates their profile, gets redirected to their profile page, and the page loads by reading from a replica that hasn't yet applied the update, showing their old data. The user assumes the save failed and tries again, sometimes creating a duplicate submission in the process. Routing that specific post-save read to the primary for a short window eliminates this exact, common complaint without requiring the team to build anything more elaborate than a short-lived routing rule keyed to the write.

Executive Capability Standard

What Good Looks Like

A well-handled replica lag strategy matches the fix to the actual cost of inconsistency, typically routing a user's own post-write reads to the primary, and only investing in heavier causal consistency tracking once a broader class of real inconsistency bugs has actually shown up.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Understand how your specific database reports replica lag and what your application's actual tolerance for it is.
2. Do Manually:Manually reproduce a read-after-write scenario against a lagging replica to see exactly what a user would experience.
3. Delegate:Assign ownership of read routing rules to whichever team owns the data layer, so the logic doesn't get duplicated inconsistently across services.
4. Automate:Automate lag-based fallback routing so reads move to the primary automatically when a replica falls behind an acceptable threshold.
5. Buy:Use a managed database service with built-in read-after-write consistency options rather than building custom routing logic from scratch.

How to Get Started

Frequently Asked Questions

How much lag is normal for a Postgres read replica?

It varies with write volume and network conditions between primary and replica, and there's no single normal number. What matters more than the typical value is monitoring it continuously and understanding your application's actual tolerance, since even a small amount of lag is enough to cause a visible bug in a read-after-write scenario.

Does routing reads to the primary after a write defeat the purpose of having replicas?

Not in practice, since this routing only applies to a narrow, short-lived window right after a specific write, not to your read traffic generally. The vast majority of reads, which aren't immediately following that user's own write, still go to replicas exactly as intended.

Is it worth monitoring replica lag even if we've fixed read-after-write?

Yes. Lag that grows unexpectedly is often an early sign of a different problem, such as a replica falling behind due to resource contention or an expensive query, and catching that trend early is valuable independent of any consistency fix you've already put in place.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides