The Data Pipeline Bug That Only Shows Up After a Retry
A data pipeline that isn't idempotent works perfectly until the exact moment it needs to retry, at which point it quietly double-counts a batch of records, double-charges a batch of invoices, or duplicates rows in a table nobody notices until a monthly report looks wrong. The failure is specifically invisible in normal operation, since a pipeline that never fails never exposes the bug.
This walks through the idempotency key pattern that makes a pipeline safe to rerun from any point, and the places teams most often forget to apply it.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Why 'just add a retry' makes the problem worse, not better
The instinctive fix for a flaky pipeline step is to wrap it in a retry loop, and that instinct is exactly backwards if the step itself isn't idempotent. A retry on a step that inserts rows, sends a webhook, or charges a payment will, on a partial failure, redo work that partially succeeded the first time, producing duplicate rows, duplicate webhook deliveries, or a duplicate charge. Adding retries to a non-idempotent pipeline doesn't add resilience, it adds a new, harder-to-reproduce bug that only appears under the exact failure conditions retries are meant to handle.
The idempotency key pattern, worked through an example
Say your pipeline processes a batch of orders from an upstream source once an hour. Attach a stable idempotency key to each unit of work, in this case the order ID plus the processing window, and before writing any result, check whether that key has already been recorded as processed. If a run fails halfway through a batch of 10,000 orders and restarts, it reprocesses the same batch, but every order already recorded gets skipped rather than reinserted, and only the genuinely unprocessed remainder does real work. The key insight is that the check-and-record step must be atomic with the write itself, not a separate step that can itself fail out of order.
Where teams forget to apply this
- Webhook delivery: a retried webhook send after a timeout can double-deliver even if the first attempt actually succeeded on the receiving end; include an idempotency key the receiver can deduplicate on.
- Downstream API calls that mutate state: a payment charge, an inventory decrement, or an email send are the highest-cost places to get this wrong, since the damage is customer-visible.
- Aggregation and rollup jobs: a job that recomputes a daily total from raw events is naturally idempotent if it fully recomputes from source; one that increments a running counter is not, and a retry will double-count.
- Multi-step pipelines where only the last step has a key: if step three has an idempotency key but step one doesn't, a partial failure between one and three still causes step one's side effects to duplicate on retry.
Making the idempotency store itself reliable
The table or store tracking which keys have already been processed needs its own durability guarantees, since if it's lost or reset, every previously processed key looks unprocessed again and the entire safety mechanism silently stops working. Set a clear retention policy, since keeping every key forever isn't usually necessary, but expiring keys too aggressively can reopen the duplication window for a legitimately delayed retry. Back this store with the same database that has your strongest durability guarantees, not a cache that can be evicted under memory pressure.
Testing idempotency deliberately, not hoping it holds
The only reliable way to know a pipeline is actually idempotent is to deliberately kill it mid-run in a test environment and rerun it, then check for duplicates in the output. Most teams never do this and only discover a gap in production, during the exact retry event the fix was supposed to protect against. Build this kill-and-rerun test into your CI pipeline for any job that writes to a durable store, and rerun it whenever the pipeline's write logic changes.
A common mistake is testing idempotency only on the happy path of a clean rerun. For example, a team kills a job between two steps, reruns it, sees no duplicate rows in the main table, and declares the pipeline safe. But the same run also sent webhooks and updated a summary counter, and both fired twice. The fix is to widen the assertion: after the rerun, compare every output the job touches, including tables, external calls, and aggregates, against a single clean run. Record the calls made to any mocked downstream service, and fail the test if any idempotency key produces more than one side effect. That turns the kill-and-rerun test from a spot check into a guarantee about the whole write path.
What Good Looks Like
A reliable pipeline attaches a stable idempotency key to every unit of work with an external or durable side effect, stores processed keys durably, and is deliberately tested by killing and rerunning it mid-batch.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
Does adding an idempotency key slow down the pipeline meaningfully?
The check-and-record step adds a small amount of latency per unit of work, usually negligible compared to the processing itself. The bigger cost is the engineering time to implement it correctly, which is worth weighing against the cost of a duplicate-charge incident.
Is a database unique constraint enough to guarantee idempotency on its own?
It helps prevent duplicate rows specifically, but it doesn't protect a downstream side effect like a webhook send or an external API call that happens before the database write. You still need an explicit idempotency key for any step with an external, non-database side effect.
How long should we retain idempotency keys before expiring them?
Long enough to cover your realistic retry window, including a delayed manual rerun after an incident, which is often longer than teams initially assume. A month is a reasonable starting point for most pipelines; extend it for anything with a slower, less automated retry path.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Making Your Data Pipeline Safe to Rerun
A nightly ETL job fails halfway through, someone reruns it, and revenue gets double counted. A worked example of building a pipeline safe to replay.
Making a Data Ingestion Pipeline Safe to Retry Without Duplicating Records
How to design idempotency keys and deduplication so a retried or replayed ingestion job never double counts or double writes a record.
Why Your Data Pipeline Needs to Survive Being Run Twice
How to design data ingestion jobs so a retry, a replay, or a duplicate message never produces duplicate rows, with concrete patterns for common failure points.
Making a Data Pipeline Safe to Replay
A worked example of tracing a batch through a pipeline to find every place a retry could duplicate it, and the idempotency key pattern that fixes it.
Making an Ingestion Pipeline Retry-Safe: A Walkthrough With Idempotency Keys
A worked example of a duplicate-row bug in a webhook ingestion pipeline, and how idempotency keys with an upsert actually fix it, versus fixes that don't.
Why Your Ingestion Pipeline Needs to Survive Being Rerun
A guide to idempotent data pipelines, why retries and reruns are inevitable, and the specific patterns that keep a rerun from duplicating or corrupting data.