Designing a Data Pipeline That Survives Being Run Twice
Every pipeline gets retried eventually, whether that's an orchestrator re-running a failed step, a webhook provider resending an event it thinks might not have arrived, or an engineer manually replaying a batch after a fix. If running the same input twice produces a different result than running it once, that's not an edge case. It's a bug waiting for the right retry to expose it.
Idempotency is the property that fixes this: the same input, processed any number of times, leaves the system in the same state it would be in after processing it once.
Why running a pipeline twice breaks things
The most common failure is double counting: an aggregation step that adds a batch's revenue to a running total will double that revenue if the batch runs again after a retry. The second most common is a side effect that shouldn't repeat, like sending a notification or charging a card, triggered again because the step that was supposed to record it happening failed before it could write that record.
Both failures share a root cause: the system has no reliable way to tell 'I already did this' from 'I haven't done this yet,' so a retry can't distinguish a genuine gap from a duplicate.
Idempotency keys: dedup at the write, not the read
An idempotency key is a unique identifier attached to a unit of work, generated once at the source, carried through every retry of that same unit. The receiving system checks whether it has already recorded that key before doing anything else. If it has, it returns the previous result instead of redoing the work.
The key has to be generated before the first attempt, not regenerated on each retry, or every retry looks like a brand new unit of work to the receiving system, which defeats the entire point.
Upserts versus append-only, and when each wins
For state that represents a current value, like a customer's subscription status, an upsert keyed on a stable identifier is naturally idempotent: writing the same value twice leaves the same result both times. For an event log, where you need to know that something happened and when, append-only with a unique constraint on the idempotency key gets you the same protection without collapsing history you actually want to keep.
Picking the wrong one is a common mistake: append-only logic on data that's really a current-state field produces duplicate rows on retry, and upsert logic on an event log erases the fact that something happened more than once when it genuinely should have.
A worked example: a webhook that fires twice
Say a payment provider sends the same 'payment succeeded' webhook twice because your endpoint was slow to respond the first time and it assumed the delivery failed. Without an idempotency check, that fires your order fulfillment logic twice. With one, keyed on the provider's own event id, the second delivery is recognized as a duplicate and safely ignored before it ever reaches the fulfillment code.
This pattern generalizes past payments to almost any inbound webhook: store the provider's event id the moment you receive it, check for that id before processing, and the exact-once guarantee falls out of that one check. The check itself has to happen inside the same transaction as recording the event id, or a second webhook arriving in the narrow window before that write completes can still slip through.
Testing idempotency on purpose
Don't wait for a real retry to find out whether a pipeline is actually idempotent. Take a real batch or a real webhook payload and deliberately replay it in a test environment, then check that the resulting state is identical to running it once. If it isn't, that gap is exactly what a production retry will eventually expose, usually at a worse time than a scheduled test.
Make this a standing part of your test suite, not a one-time exercise. Every new pipeline that has a real side effect should ship with a rerun test alongside its normal tests, so idempotency is checked automatically every time the pipeline changes, not just the day it was first built.
A deliberate replay test can follow these steps:
- Take a real batch or a real webhook payload instead of a synthetic example.
- Run it once in a test environment and record the resulting state.
- Replay the same input in the same environment, ideally more than once.
- Compare the resulting state with the first run and investigate any difference.
- Treat any gap as the same failure a production retry will eventually expose, usually at a worse time.
What Good Looks Like
Good idempotency design means you can replay any real batch or event in a test environment and get exactly the same resulting state as running it once.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
What's the difference between idempotency and deduplication?
Deduplication usually means filtering out duplicates after the fact, once they've already been recorded. Idempotency prevents the duplicate from ever taking effect in the first place, by checking a unique key before any side effect happens, which is the safer and generally preferred approach.
Where should the idempotency key come from?
Ideally from the source system itself, like a webhook provider's own event id, generated once and carried through every retry. If you're generating the key yourself, generate it once at the start of the operation and reuse it on every retry attempt, never regenerate it per attempt.
Do we need idempotency for every pipeline?
Any pipeline that could plausibly be retried, which in practice is nearly all of them, benefits from it. The exceptions are pipelines with no meaningful side effects at all, where running twice genuinely produces no different outcome than running once.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
A Rollout Checklist for Swapping Models in Production
A rollout checklist for swapping AI models in production: evaluation gates, canary and shadow traffic, and fast rollback paths.
Why Inference Latency Creeps Up After You Ship
Where AI model-serving latency actually hides: tokenization, queueing, batching windows, and network hops, plus a worked example fix.
Zero Trust for Machines Calling Your Model Endpoints
Why internal services calling your inference endpoints still need identity checks, and where CrowdStrike-style posture checks and Tenable-style scanning fit.
What a Real Security Audit of Model Serving Should Cover
A practical checklist for auditing AI model serving and inference: endpoint access, weight security, prompt logging, and patch timelines.
What SOC 2 Actually Expects From a Model-Serving Team
What SOC 2 expects from a team serving AI models: how change, access, patch, and vendor controls apply, and the evidence to have ready.
Data Residency Questions to Settle Before Picking an Inference Region
The questions to answer before you choose where inference runs: where prompts are processed, where logs live, and what to confirm with regulated customers.