Distributed Systems & Enterprise ResiliencePlaybook3 min readUpdated September 2026

Making an Ingestion Pipeline Retry-Safe: A Walkthrough With Idempotency Keys

A webhook sends an event, your ingestion pipeline times out before acknowledging it, and the sender retries. Nothing about that is unusual; it's how most webhook systems are designed to behave. The problem shows up downstream: your pipeline had actually finished processing the first attempt, so the retry creates a second row for the same event, and now a report built on that table is double-counting.

The fix isn't catching duplicates after they land. It's making the write itself incapable of producing one in the first place.

The bug, concretely

The consumer received the event, wrote a row, and was in the middle of sending its acknowledgment when the connection dropped. From the sender's point of view, the request failed, so it retried on schedule. From the pipeline's point of view, it had already done the work; it just never got to say so. The result is one real event represented as two rows, and nothing in the pipeline itself noticed anything went wrong, because both writes succeeded.

The fix that doesn't actually work: deduping after the fact

A tempting patch is a nightly job that looks for rows with matching content and removes the duplicates. It's fragile in the exact way that matters: it depends on defining 'matching' correctly for every event type, it runs after the duplicate has already been counted in anything that reads the table during the day, and it silently breaks the moment two genuinely different events happen to look similar under whatever matching rule was chosen. It treats the symptom, not the cause, and it does so on a delay.

The fix that works: an idempotency key and an upsert

Every event needs a stable identifier that's the same across retries: an event ID assigned by the sender, or if the sender doesn't provide one, a deterministic hash of the payload's immutable fields. The write itself changes from a plain insert to an upsert keyed on that identifier (ON CONFLICT DO NOTHING, or DO UPDATE if a retry might carry corrected data). The second attempt at writing the same event either does nothing or safely overwrites with the same result, instead of creating a second row. The database, not a nightly cleanup job, is what enforces uniqueness, and it enforces it at write time, not after the fact.

To make the write itself retry-safe, follow these steps:

  1. Give every event a stable identifier that stays the same across retries, using the sender's event ID when one exists.
  2. If the sender provides none, derive a deterministic hash from the payload's immutable fields, excluding anything that changes per attempt.
  3. Change the write from a plain insert to an upsert keyed on that identifier, so a retry can't create a second row.
  4. Choose DO NOTHING when retries carry identical data, and DO UPDATE when a retry may carry corrected data that should win.

Where idempotency keys quietly break down

Two things undermine this pattern even after it's implemented. First, a key that's generated at send time rather than derived from the payload's own content isn't actually stable across retries if the sender regenerates it on each attempt; check the sender's documentation for whether the ID it provides is guaranteed constant across retries of the same logical event, not just unique per request. Second, a check-then-insert pattern (look up whether the key exists, then insert if not) has a race condition under concurrent processing that a single upsert statement doesn't, because two workers can both pass the check before either one writes. The upsert needs to be one atomic database operation, not two separate steps in application code.

Testing it, since it's rarely automated

The actual test is simple and almost never written down: take a real batch, replay it twice through the pipeline in a staging environment, and assert the row count after the second run equals the row count after the first. If it doesn't, the pipeline isn't idempotent yet regardless of what the code looks like. This is worth adding as a standing integration test, not a one-time manual check, because a future change to the write path can silently reintroduce the plain-insert bug the upsert was meant to prevent.

Idempotency further downstream, past the first write

The first write isn't the only place duplication can sneak back in. A downstream job that aggregates the ingestion table into a daily rollup, or a consumer that reads from the table and pushes to a third system, needs its own idempotency story if it can also be retried or rerun, since a clean upstream table doesn't automatically protect anything built on top of it. Treat every stage that writes as a stage that needs a stable key and a safe re-run behavior, not just the first one in the chain.

Executive Capability Standard

What Good Looks Like

A retry-safe pipeline uses a stable, payload-derived idempotency key and an atomic upsert at write time, verified by a standing test that replays a real batch twice and checks the row count doesn't change.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Audit your ingestion pipelines for plain inserts on data that could plausibly be retried by an upstream sender.
2. Do Manually:Manually replay a real batch twice against a staging copy of the highest-risk pipeline and confirm whether duplicates appear.
3. Delegate:Have an engineer add stable idempotency keys and convert the affected writes to upserts, starting with the pipeline that failed the manual replay test.
4. Automate:Add a standing integration test that replays a batch twice and asserts the row count doesn't change, so a future change can't silently reintroduce the bug.
5. Buy:Bring in data engineering advisory if idempotency issues are already showing up as data quality problems across multiple pipelines and need a systematic fix.

How to Get Started

Frequently Asked Questions

What if the upstream sender doesn't provide a stable event ID?

Derive one yourself with a deterministic hash of the payload's immutable fields, excluding anything that changes between retries of the same logical event, such as a timestamp added at send time. The key needs to be identical across retries of the same event, not just unique per request, or it won't catch the duplicate.

Is ON CONFLICT DO NOTHING always the right choice over DO UPDATE?

Use DO NOTHING when a retry always carries identical data, which is the common case. Use DO UPDATE when a retry might legitimately carry corrected data for the same event, in which case you want the second write to win, not be silently dropped.

Can we add idempotency at the application layer instead of the database?

You can, with a check-then-insert pattern, but it introduces a race condition under concurrent processing that a database-level upsert doesn't have, since two workers can both pass the check before either writes. A single atomic upsert statement is the safer default unless there's a specific reason the database can't enforce it.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides