Model Context Protocol & Agentic ArchitecturePlaybook4 min readUpdated September 2026

Making Data Pipeline Retries Safe: A Walkthrough

A data pipeline is safe to retry only when it is idempotent, meaning a repeated write or event produces the same result as the first one. Without that, a network timeout, a crashed worker, or a webhook that fires twice will quietly double-count revenue or duplicate customer records, and the bug surfaces weeks later as a number that won't reconcile.

This walks through making a pipeline safe to retry: picking a key that actually identifies a duplicate, handling a batch that fails partway through, and testing that the guarantee holds instead of just assuming it does.

The Failure That Makes Idempotency Non-Negotiable

Every pipeline eventually gets retried, whether you designed for it or not: a deploy interrupts a running job, a downstream API times out and your retry logic fires, a message queue redelivers a message it already delivered because the consumer didn't acknowledge it in time. None of these are exotic edge cases. They're the normal operating conditions of a distributed system running continuously.

If reprocessing the same input produces a different result than processing it once, the pipeline isn't safe to retry, which means every one of those ordinary failure modes above is also a data corruption risk. Idempotency isn't an optimization you add later. It's the property that makes retrying safe at all, and without it, your actual reliability strategy is hoping nothing ever needs a retry.

How do you pick a dedup key that prevents duplicates?

The dedup key needs to identify the same logical event across retries, not just the same row. A key built from an auto-incrementing database ID assigned during processing doesn't work, since a retry generates a new ID for what's logically the same event. A key built from the source system's own event ID, a webhook's event identifier, an upstream system's transaction ID, is what actually survives a retry, since the source sends the same ID both times.

If the upstream system doesn't provide a stable ID, derive one deterministically from the event's own content, a hash of the fields that make it unique, so the same logical event always hashes to the same key even without an explicit ID to rely on. Whatever the source, check the key against real duplicate events before trusting it. A key that looks unique in a sample of clean data can still collide once you see the actual variety of malformed or edge-case events your source produces.

How do you handle partial batch failures without reprocessing everything?

When a large batch fails partway through, reprocessing the whole batch from the start is the easy answer and the wrong one if any of the earlier records already wrote successfully and aren't themselves idempotent on the retry. Track progress at the record level, not just the batch level, so a retry resumes from the actual failure point or safely skips records already confirmed written.

The cleanest pattern writes each record's dedup key to a tracking table in the same transaction as the record itself, so a retry can check that table first and skip anything already committed, rather than trusting an external progress counter that might be stale if the job crashed before updating it.

Idempotency at the Write, Not Just at the Read

Deduplicating on read, filtering duplicates out of a query after they've already landed, is a workaround, not a fix. It leaves the duplicate data sitting in storage, adds a filtering step every downstream consumer has to remember to apply, and doesn't help if a downstream system already consumed the duplicate before your filter ran. The dedup check needs to happen at write time: an upsert keyed on the dedup key, or a conditional insert that fails silently, and is caught, rather than creating a duplicate row.

For side effects beyond the database, sending an email, calling a third-party API, charging a card, idempotency has to extend to that call too, usually by checking a log of dedup keys that already triggered the side effect before triggering it again. A database-level upsert alone doesn't stop a duplicate email from going out if the email send happens before the database write.

Testing That Your Pipeline Is Actually Idempotent

Don't assume idempotency works because the design looks right on paper. Test it directly:

  • Replay the exact same batch of input twice in a row and confirm the resulting data is identical, not just that no error was thrown.
  • Kill the pipeline worker mid-batch, on purpose, in a staging environment, and confirm the retry picks up cleanly without duplicating anything already written.
  • Send a duplicate webhook or event deliberately and confirm downstream side effects, emails, API calls, notifications, don't fire twice.
  • Check the dedup key logic against real production data samples, not just clean test fixtures, since malformed or unusual source events are exactly where a dedup key assumption tends to break.

A pipeline that passes these checks is idempotent in practice, not just in the design document. Revisit the tests whenever the source system's event format changes, since a new field or a changed ID format can quietly invalidate a dedup key that used to work.

Executive Capability Standard

What Good Looks Like

A good idempotent pipeline means replaying the same input twice produces identical output, verified by an actual replay test, not assumed from the design.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Trace what happens today if your pipeline's most critical job gets retried mid-batch: does it duplicate data, or does nothing change.
2. Do Manually:Add a dedup key derived from a stable source identifier to your highest-risk pipeline first, and manually replay a batch to confirm it holds.
3. Delegate:Have whoever owns each pipeline document its dedup key and confirm it survives a partial batch failure, not just a clean full retry.
4. Automate:Build partial-batch tracking into your pipeline framework so a failure at any point resumes from the right place automatically instead of reprocessing everything.
5. Buy:Consider a managed data integration platform with idempotent delivery built in if you're building many similar ingestion pipelines from scratch and reinventing this guarantee each time.

How to Get Started

Frequently Asked Questions

Is deduplicating on read good enough if we're careful about it?

It's a workaround, not a fix. It leaves duplicate data in storage, requires every downstream consumer to remember to filter, and doesn't help once a duplicate has already been consumed by a system that read before your filter applied. Deduplicate at write time instead.

What if the upstream system doesn't give us a stable event ID?

Derive a dedup key deterministically from the event's own content, a hash of the fields that make it unique, so the same logical event always produces the same key even without an explicit ID. Test it against real duplicate and edge-case events before trusting it in production.

Do side effects like sending emails need to be idempotent too, or just the database writes?

Both. A database upsert prevents a duplicate row, but it doesn't stop a duplicate email or API call if that side effect happens before the database write. Check a log of already-triggered side effects, keyed the same way as your data dedup key, before firing one again.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides