Why Your Ingestion Pipeline Needs to Survive Being Rerun
Every data pipeline eventually gets rerun: a job times out partway through and the scheduler retries it, an engineer reprocesses a day of data after fixing a bug, or a network blip causes a message to be redelivered. If the pipeline isn't built to handle that gracefully, the result is duplicate rows, doubled totals, or records processed twice with subtly different outcomes each time.
Idempotency means running the same operation twice produces the same result as running it once. It sounds like an abstract property, but the absence of it is one of the more common causes of a data quality incident that takes days to fully untangle.
Why reruns are inevitable, not exceptional
Treating a pipeline rerun as a rare edge case is the root of most idempotency bugs, because in practice reruns happen constantly: a scheduler retry after a transient failure, a manual backfill after a bug fix, an at-least-once message queue redelivering a message it isn't sure was processed. Design for reruns as the normal case, not a special one, and a large category of data quality incidents simply stops happening.
Pattern: use a natural or generated key, and upsert instead of insert
Instead of blindly inserting every record a pipeline processes, write to the destination using a stable key, whether that's a natural identifier from the source system or one generated deterministically from the record's contents, and upsert rather than insert. Running the same batch twice then produces the same final state instead of duplicate rows, because the second run's writes simply overwrite the first run's writes for the same key.
Pattern: make aggregation idempotent, not additive
A pipeline that increments a running total each time it processes a batch will double-count on a rerun, since the increment happens again. Recompute the aggregate from the underlying records instead of accumulating an increment on top of the last known value, so a rerun produces the same total regardless of how many times it's been triggered. This costs more compute than an incremental update, and that cost is almost always worth paying for the correctness guarantee.
Pattern: track processed state explicitly
For pipelines where recomputing from scratch on every run isn't practical, track exactly what's been processed, down to the specific batch or offset, in a durable store the pipeline checks before doing any work. On a rerun, the pipeline can then skip anything already marked complete and pick up cleanly from where it actually left off, rather than guessing based on a timestamp that might not reflect what was truly written to the destination. Mark a batch complete only after its write is confirmed durable, not before, so a crash between the write and the marker still leaves the pipeline able to safely redo that one batch.
A common mistake: idempotent writes with non-idempotent side effects
A pipeline can get the database write exactly right, using upserts and stable keys, while still triggering a side effect like a customer notification email on every run. The database ends up correct. The customer still gets three copies of the same email because the side effect itself was never made idempotent. Any action a pipeline takes outside its own data store, not just the write to it, needs the same rerun-safe treatment, whether that means checking a sent-status flag before firing a notification or routing the notification itself through a deduplicating queue.
Make a pipeline safe to rerun with these patterns:
- Write to the destination with a stable key, natural or deterministically generated, and upsert rather than insert.
- Recompute aggregates from the underlying records instead of adding an increment onto the last known total.
- Track processed batches or offsets in a durable store that the pipeline checks before doing any work.
- Make side effects outside the data store, such as customer emails, idempotent too, not only the database writes.
- Test by running the same batch through twice and asserting the final state matches the state after one run.
A worked example: the backfill that doubled revenue
Say an analytics pipeline recalculates yesterday's revenue by summing new transaction records and adding that sum onto an existing running total in a dashboard table. An engineer reruns the job to backfill a day that failed silently, and the sum gets added a second time onto a total that already included it, quietly doubling reported revenue for that day. Nobody notices until finance reconciles at month end. Recomputing the total from the full set of underlying transactions each run, rather than adding to a stored running number, would have made the exact same backfill completely safe to rerun.
What Good Looks Like
A properly idempotent pipeline produces the same final state whether it runs once or is rerun several times, using stable keys for writes, recomputed rather than accumulated aggregates, and rerun-safe handling for any external side effects.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Does idempotency mean a pipeline can never use auto-incrementing IDs?
It means auto-incrementing IDs shouldn't be the thing that determines whether a record is new or a duplicate. You can still use them as a storage detail, but the pipeline needs a separate, stable key, from the source system or deterministically generated, to decide whether a given record has already been written.
Is making a pipeline idempotent always worth the extra engineering effort?
For anything that feeds a number people make decisions from, like revenue or usage metrics, yes, since the cost of a silent double-count discovered weeks later is almost always higher than the upfront cost of designing for reruns. For a low-stakes internal report, the tradeoff is more of a judgment call.
How do we test that a pipeline is actually idempotent?
Run the same batch of input data through the pipeline twice in a test environment and assert the final state is identical after both runs, not just after the first. A pipeline that looks correct after a single run can still fail this test the moment it's rerun, which is exactly the scenario worth catching before production.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
A Go-Live Checklist for Shipping a RAG Pipeline
The specific steps to check before a RAG pipeline goes live: index warm-up, model version pinning, a canary check, and a real rollback plan.
Zero Trust for a RAG Pipeline Means No Service Gets a Free Pass
A decision framework for applying zero trust to a production RAG pipeline: verifying every service and user call, not just the ones at the edge.
Where RAG Latency Actually Goes, and How to Budget It
Break a RAG request into its four latency stages, find out which one is actually slow, and set a budget for each before you start tuning blindly.
What to Check First in a RAG Pipeline Security Audit
A practical order of operations for auditing a production RAG pipeline: data exposure, prompt injection, access control, logging, and vendor risk.
Mapping SOC 2 Controls to a RAG Pipeline's Real Components
SOC 2 auditors ask about access, change management, and vendors in the abstract. Here's what each control actually maps to in a RAG pipeline.
How Vector Search Throughput Degrades as Your Index Grows
Throughput doesn't fall off gradually as a vector index grows. Here's why it degrades in steps, and how sharding, replicas, and quantization each help.