Making Your Data Pipeline Safe to Rerun
A nightly pipeline fails halfway through. Someone reruns it, and by morning half of yesterday's revenue rows are counted twice in the warehouse, because nothing about the pipeline was designed to be replayed safely.
At-least-once delivery, a message redelivered, a batch retried after a partial failure, is the normal case for most pipelines, not an edge case. Idempotency isn't optional if you ever plan to rerun a job, which you will.
Why 'Just Rerun It' Is the Failure Mode, Not the Fix
Retries, redelivered messages, and reruns after a partial failure happen constantly in any real pipeline, whether you planned for them or not. If your writes aren't idempotent, every one of those normal events is a chance to duplicate or corrupt data. Treating idempotency as a nice-to-have instead of a correctness requirement is how a routine retry turns into a data quality incident.
The Core Pattern: A Stable Key and an Upsert
Give every record a deterministic key derived from the source data itself, a source record id combined with an event type or version, rather than relying on an auto-generated id assigned at write time. Write with an upsert against that key instead of a plain insert, so replaying the same batch twice produces the same end state instead of two copies of every row.
Where This Gets Harder Than It Looks
Raw ingestion is the easy part. Aggregation steps are where idempotency quietly breaks: a naive "add this batch's total to the running total" step double counts on rerun even if the raw rows underneath it are perfectly deduplicated, because the increment itself isn't idempotent. The fix is recomputing the aggregate from the already-idempotent raw table on every run, rather than incrementing a running number.
A Worked Example
Consider an order-events pipeline. Raw events land in a staging table keyed by order id, event type, and event timestamp, written with an upsert, so a redelivered event overwrites itself instead of duplicating. A daily revenue rollup is then recomputed from that staging table on every run, rather than incremented from the previous day's total. If the job fails halfway through and gets rerun, the staging table ends up in the same state either way, and the rollup recomputed from it produces the identical number, not a doubled one.
A Checklist Before You Trust a Pipeline to Rerun Safely
- Every write step uses a deterministic key derived from source data, not an auto-generated one assigned at write time.
- Every aggregate is recomputed from idempotent raw data on each run, not incremented from a previous total.
- A partial failure halfway through a batch doesn't leave the destination table in a state a rerun can't reconcile cleanly.
- You've actually tested a rerun after a deliberate partial failure, not just assumed the design will hold up.
Handling Late-Arriving and Out-of-Order Data
A record that arrives a day late, common when an upstream system retries its own failed delivery, is a special case of the same idempotency problem: your pipeline needs to accept it into the correct place without disturbing everything computed since. An upsert keyed on the record's own stable identity handles this the same way it handles a straightforward rerun, since the pipeline doesn't actually need to know the record arrived late, only that it needs to land in the same place it would have on time.
Aggregates that were already computed for the affected period still need to be recomputed once the late record lands, which is exactly why the recompute-from-raw-data pattern matters more than an incremental one: a recompute naturally picks up the late arrival, while an incremental running total has already moved on and has no mechanism to go back and correct itself.
Testing Idempotency Like You'd Test Anything Else
Most teams test a pipeline's happy path and stop, so a rerun scenario is never actually exercised until it happens for real, usually during an incident. Add a test that runs the pipeline once, kills it partway through a second run of the same input, reruns it again, and asserts the final output matches a single clean run byte for byte, not just approximately.
Run that same test after any schema change to the destination table, not only when the pipeline logic itself changes. A new column or a changed constraint can quietly break the upsert key's uniqueness guarantee even when the pipeline code hasn't been touched at all.
What Good Looks Like
Good pipeline idempotency means a rerun after a partial failure produces exactly the same result as if the job had run once cleanly, with no manual cleanup step required afterward.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Is exactly-once delivery from the message queue the real fix?
No, because exactly-once delivery is hard to guarantee end to end across real systems. Even where a queue offers it, a bug in your own processing logic can still duplicate a write. Building idempotent writes at the destination is a more reliable safeguard than trusting delivery semantics from every upstream component.
How do you make an aggregation step idempotent when the raw ingestion already is?
Recompute the aggregate from the idempotent raw table on every run instead of incrementing a running total. Incrementing is what breaks on a rerun, because the same batch gets added twice even if the underlying rows were correctly deduplicated. A full recompute over deduplicated raw data always lands on the same answer.
What if the source system doesn't provide a natural stable key?
Construct one from a combination of fields that are stable across redeliveries, such as a source id plus an event timestamp or a content hash of the record itself. The goal is a key that produces the same value every time the same logical event is processed, even if it's delivered more than once.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Making a Data Ingestion Pipeline Safe to Retry Without Duplicating Records
How to design idempotency keys and deduplication so a retried or replayed ingestion job never double counts or double writes a record.
Making a Data Pipeline Safe to Replay
A worked example of tracing a batch through a pipeline to find every place a retry could duplicate it, and the idempotency key pattern that fixes it.
The Data Pipeline Bug That Only Shows Up After a Retry
Why a retried job silently duplicates data in most pipelines, and the idempotency key pattern that makes a pipeline safe to rerun from any failure point.
Why Your Data Pipeline Needs to Survive Being Run Twice
How to design data ingestion jobs so a retry, a replay, or a duplicate message never produces duplicate rows, with concrete patterns for common failure points.
Why Your Ingestion Pipeline Needs to Survive Being Rerun
A guide to idempotent data pipelines, why retries and reruns are inevitable, and the specific patterns that keep a rerun from duplicating or corrupting data.
Making an Ingestion Pipeline Retry-Safe: A Walkthrough With Idempotency Keys
A worked example of a duplicate-row bug in a webhook ingestion pipeline, and how idempotency keys with an upsert actually fix it, versus fixes that don't.