Developer Productivity & Platform EngineeringPlaybook3 min readUpdated September 2026

Making Your Data Pipeline Safe to Rerun

A nightly pipeline fails halfway through. Someone reruns it, and by morning half of yesterday's revenue rows are counted twice in the warehouse, because nothing about the pipeline was designed to be replayed safely.

At-least-once delivery, a message redelivered, a batch retried after a partial failure, is the normal case for most pipelines, not an edge case. Idempotency isn't optional if you ever plan to rerun a job, which you will.

Why 'Just Rerun It' Is the Failure Mode, Not the Fix

Retries, redelivered messages, and reruns after a partial failure happen constantly in any real pipeline, whether you planned for them or not. If your writes aren't idempotent, every one of those normal events is a chance to duplicate or corrupt data. Treating idempotency as a nice-to-have instead of a correctness requirement is how a routine retry turns into a data quality incident.

The Core Pattern: A Stable Key and an Upsert

Give every record a deterministic key derived from the source data itself, a source record id combined with an event type or version, rather than relying on an auto-generated id assigned at write time. Write with an upsert against that key instead of a plain insert, so replaying the same batch twice produces the same end state instead of two copies of every row.

Where This Gets Harder Than It Looks

Raw ingestion is the easy part. Aggregation steps are where idempotency quietly breaks: a naive "add this batch's total to the running total" step double counts on rerun even if the raw rows underneath it are perfectly deduplicated, because the increment itself isn't idempotent. The fix is recomputing the aggregate from the already-idempotent raw table on every run, rather than incrementing a running number.

A Worked Example

Consider an order-events pipeline. Raw events land in a staging table keyed by order id, event type, and event timestamp, written with an upsert, so a redelivered event overwrites itself instead of duplicating. A daily revenue rollup is then recomputed from that staging table on every run, rather than incremented from the previous day's total. If the job fails halfway through and gets rerun, the staging table ends up in the same state either way, and the rollup recomputed from it produces the identical number, not a doubled one.

A Checklist Before You Trust a Pipeline to Rerun Safely

  • Every write step uses a deterministic key derived from source data, not an auto-generated one assigned at write time.
  • Every aggregate is recomputed from idempotent raw data on each run, not incremented from a previous total.
  • A partial failure halfway through a batch doesn't leave the destination table in a state a rerun can't reconcile cleanly.
  • You've actually tested a rerun after a deliberate partial failure, not just assumed the design will hold up.

Handling Late-Arriving and Out-of-Order Data

A record that arrives a day late, common when an upstream system retries its own failed delivery, is a special case of the same idempotency problem: your pipeline needs to accept it into the correct place without disturbing everything computed since. An upsert keyed on the record's own stable identity handles this the same way it handles a straightforward rerun, since the pipeline doesn't actually need to know the record arrived late, only that it needs to land in the same place it would have on time.

Aggregates that were already computed for the affected period still need to be recomputed once the late record lands, which is exactly why the recompute-from-raw-data pattern matters more than an incremental one: a recompute naturally picks up the late arrival, while an incremental running total has already moved on and has no mechanism to go back and correct itself.

Testing Idempotency Like You'd Test Anything Else

Most teams test a pipeline's happy path and stop, so a rerun scenario is never actually exercised until it happens for real, usually during an incident. Add a test that runs the pipeline once, kills it partway through a second run of the same input, reruns it again, and asserts the final output matches a single clean run byte for byte, not just approximately.

Run that same test after any schema change to the destination table, not only when the pipeline logic itself changes. A new column or a changed constraint can quietly break the upsert key's uniqueness guarantee even when the pipeline code hasn't been touched at all.

Executive Capability Standard

What Good Looks Like

Good pipeline idempotency means a rerun after a partial failure produces exactly the same result as if the job had run once cleanly, with no manual cleanup step required afterward.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Audit your highest-value pipeline's write steps and check whether any of them would duplicate data if the job ran twice.
2. Do Manually:Manually convert your riskiest insert-only write step to a keyed upsert as a first fix before touching anything else.
3. Delegate:Assign one engineer to own a rerun-safety checklist that every new pipeline has to pass before it ships.
4. Automate:Add an automated test that deliberately reruns each pipeline after a simulated partial failure and checks the output matches a single clean run.
5. Buy:Bring in a data engineering specialist to review your highest-value pipelines for idempotency before a migration or a major schema change.

How to Get Started

Frequently Asked Questions

Is exactly-once delivery from the message queue the real fix?

No, because exactly-once delivery is hard to guarantee end to end across real systems. Even where a queue offers it, a bug in your own processing logic can still duplicate a write. Building idempotent writes at the destination is a more reliable safeguard than trusting delivery semantics from every upstream component.

How do you make an aggregation step idempotent when the raw ingestion already is?

Recompute the aggregate from the idempotent raw table on every run instead of incrementing a running total. Incrementing is what breaks on a rerun, because the same batch gets added twice even if the underlying rows were correctly deduplicated. A full recompute over deduplicated raw data always lands on the same answer.

What if the source system doesn't provide a natural stable key?

Construct one from a combination of fields that are stable across redeliveries, such as a source id plus an event timestamp or a content hash of the record itself. The goal is a key that produces the same value every time the same logical event is processed, even if it's delivered more than once.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides