Incident Management & On-Call Operations4 min readUpdated September 2026

Pipelines Fail Quietly. Detection Is the Real Gap.

A pipeline can fail at two in the morning and nobody notices until a client opens a dashboard at nine and finds yesterday's numbers still showing. Nothing paged anyone, because from the pipeline's point of view, nothing crashed, it just stopped producing fresh data.

For a BI or data engineering consultancy, that detection gap matters more than the choice between incident.io and PagerDuty, which only helps once something is actually generating an alert worth escalating.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

What actually needs to trigger an alert here

A job that fails outright is easy to catch. A job that succeeds but produces incomplete or stale data is the harder case, and it is the one that actually damages client trust, since the dashboard looks fine right up until someone notices the numbers have not moved.

Build a freshness or row-count check as its own alert source, separate from job success or failure, before worrying about which paging tool receives it. That check is the part most consultancies skip, and it is also the part that actually catches the failures clients notice first.

Alerts worth building for a data pipeline:

  • A freshness check that fires when a dashboard's underlying data hasn't updated when expected, even though the job technically succeeded.
  • A row-count check that catches jobs that finish but produce incomplete data.
  • Suppression rules for jobs known to run long, so month-end reconciliations don't train the team to ignore pages.
  • Routing that pages one on-call analyst first and escalates only if that person doesn't acknowledge.

Suppression matters as much as detection

Some jobs legitimately run long, a large monthly reconciliation, a backfill, a client's unusually large data drop. If every long-running job pages someone, the team starts ignoring pages, which defeats the point.

Whichever tool you use, invest time in suppression rules for known-long jobs specifically, so the alerts that do fire are ones worth waking up for. A team that has learned to ignore its own alerts is in a worse position than a team with no alerts at all, because at least the second team knows it has a gap.

Route a warehouse failure to one analyst, not the whole team

A single warehouse or pipeline failure rarely needs five people paged at once. incident.io's flexible routing inside Slack and PagerDuty's escalation policies can both be configured to page one on-call analyst first, with escalation to others only if that person does not acknowledge.

That is a meaningfully different default from a group alert that trains everyone to assume someone else will handle it, which is exactly what a broadcast-style alert tends to produce over time.

A worked example: the Monday reconciliation that runs long every month

Say your client's month-end reconciliation job reliably takes three times as long as a normal daily run, every single month, without failing. A naive alert threshold set for the daily job will page someone every month-end for a job that is actually fine, training the on-call analyst to dismiss the alert without checking.

A suppression rule scoped specifically to that job's known schedule fixes this permanently, and it is a better fix than raising the general threshold, which would just let a real daily failure run longer before anyone notices.

Client-facing freshness matters more than internal recovery speed here

Engineering teams often measure themselves on how quickly they recover from a failed deployment. For a data consultancy, the client-facing metric that actually matters is closer to data freshness: how long a dashboard sat stale before someone noticed and fixed it, which depends entirely on whether the detection layer described above exists at all.

A fast incident response process built on top of no detection layer still leaves a client staring at yesterday's numbers all morning.

incident.io's channel workflow fits well once an alert exists

Once a real alert fires, incident.io's Slack-native channel creation and timeline logging work well for a data team's response, especially for documenting exactly which upstream source or transformation step caused the failure.

That detail is easy to lose if it only lives in someone's memory by the time the postmortem gets written, and it is often the single most useful line in the postmortem for preventing the same failure next month.

What the postmortem needs to say about the upstream source, not just your pipeline

When a client's dashboard goes stale, the root cause is often not your own pipeline code at all, it is a schema change in an upstream source system that nobody told you about, a vendor API that silently changed its response format, or a source database that ran a migration over a weekend. Writing a postmortem that only documents what your pipeline did, without naming the upstream trigger clearly, leaves the client with an incomplete picture of whether the next failure is something you can prevent or something that depends on a system outside your control.

Say a client's CRM vendor renames a field during a routine update, and your extraction job either fails outright or, worse, silently maps the wrong column into a downstream metric. The postmortem's most useful line for preventing a repeat is not a description of your retry logic, it is a clear statement that the trigger was an upstream schema change and a note on whether you now have a check that would catch a similar change before it reaches a client dashboard.

That distinction also matters for the client relationship directly. A client who understands that a specific vendor's behavior caused the outage trusts your team more than one who is left assuming your own code was simply unreliable, and it gives you a concrete answer when the same client eventually asks whether this could happen again.

Executive Capability Standard

What Good Looks Like

A strong data consultancy detects a stale or incomplete pipeline before a client notices it on a dashboard, routes the resulting alert to one on-call analyst rather than the whole team, and documents which upstream source or transformation step caused each failure.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Audit your top client dashboards and identify which ones have no freshness or row-count check behind them, only job success monitoring.
2. Do Manually:Have someone manually spot-check key dashboards each morning before a client would see them, as a stopgap while proper alerting gets built.
3. Delegate:Assign a data reliability owner per client engagement, responsible for freshness checks and suppression rules on that client's pipelines specifically.
4. Automate:Route freshness and data-quality alerts into incident.io or PagerDuty, with suppression rules for known long-running jobs, so real alerts reach one on-call analyst.
5. Buy:Build a standard data observability layer you attach to every new client pipeline at delivery, so alerting is part of the initial build rather than added after the first stale-dashboard incident.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

AWS

If your pipelines run on AWS-hosted infrastructure, its native health checks and multi-Availability Zone patterns can catch some infrastructure-level failures before they ever reach the data-quality layer.

Visit AWS→

Frequently Asked Questions

Why doesn't a failed job always trigger an alert for a data pipeline?

Because many pipeline problems are not failures in the technical sense: the job runs, exits successfully, but produces stale, incomplete, or wrong data. Catching that requires a separate freshness or data-quality check, not just monitoring whether the job itself crashed.

How do we stop alert fatigue from jobs that legitimately run long?

Build explicit suppression rules for known long-running jobs, based on their normal duration, rather than a blanket alert threshold applied to every job equally. Without that, the team learns to ignore alerts, which defeats the purpose of having them.

Who should get paged when a data warehouse job fails?

One on-call analyst first, with escalation to others only if that person doesn't acknowledge. A single warehouse or pipeline failure rarely needs five people paged at once, and both incident.io's routing in Slack and PagerDuty's escalation policies can be configured that way.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides