Developer Productivity & Platform EngineeringPlaybook3 min readUpdated September 2026

Why Your SLA Dashboard Says Green While Customers Are Down

A customer emails to say the app has been down for twenty minutes. Your dashboard says every service is healthy. This gap, between what your monitoring measures and what a customer actually experiences, is the most common reason SLA enforcement fails, and it usually isn't a monitoring outage, it's a monitoring design problem.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Why do health checks miss real outages?

Most health checks confirm a process is running and can respond to a ping. That's not the same as confirming the feature a customer relies on actually works end to end. A service can return a healthy check while the specific database query behind your login page times out for every real request.

Build at least one synthetic check per critical user flow, actually logging in, actually loading the dashboard, on the same schedule as your infrastructure health checks. The difference between "the process is up" and "the feature works" is exactly the gap where customers notice problems before your monitoring does.

Your Uptime Budget Should Set Your Alert Thresholds

If you've committed to an uptime target, your alerting should be calibrated against the actual downtime budget that target allows, not an arbitrary error rate someone picked at launch1. A tighter commitment needs faster detection and a lower tolerance for sustained errors before paging someone.

Work backward from the budget: if your annual allowance is small, your alert has to fire within minutes of a real outage starting, not after an hour of degraded performance quietly burns through a chunk of the year's allowance in one incident.

For example, suppose your monitoring alerts only when the error rate stays high for a full hour. If a bad release breaks sign-in at the start of that window, a large slice of the year's downtime allowance can disappear before anyone is paged. Working backward from the budget shows that the alert needs to fire within minutes, which usually means a short evaluation window on a synthetic sign-in check rather than a long window on aggregate errors. Then test it: break the check on purpose in a staging environment and time how long the page takes to arrive.

How do aggregate metrics hide localized outages?

An SLA dashboard built on aggregate success rate across all traffic can stay green while one customer, one region, or one feature is completely broken, simply because the volume of everything else drowns out the failure in the average.

Break your SLA metrics down by the dimensions that matter to your actual customers: by region, by plan tier, by the specific critical flow. A single aggregate number is convenient for a status page and nearly useless for catching the outage a customer is calling you about right now.

Alert Fatigue Is Why Real Alerts Get Ignored

A team that gets paged for every minor blip stops trusting alerts within a few weeks, and the real outage alert arrives in the same noisy channel as everything else. This is often the actual root cause when a team says "the alert fired, we just didn't act on it fast enough."

Tune thresholds so a page means something is genuinely wrong, and route lower-severity signals to a dashboard someone checks, not a channel that wakes someone up. Fewer, more trustworthy alerts beat more, ignorable ones every time.

Close the Loop With a Real Postmortem

Every time the dashboard said green while a customer was actually down, that's a specific, fixable gap in what you measure. Track these separately from ordinary incidents, because they're telling you something different: not that something broke, but that your detection has a blind spot.

Fix the specific check, not just the incident. A team that patches the immediate outage without asking why monitoring missed it will find the same blind spot again, usually at a worse time.

Close the gap between the dashboard and the customer with these checks:

  • Add a synthetic check for each critical user flow, such as signing in and loading the dashboard, on the same schedule as infrastructure health checks.
  • Set alert thresholds from the downtime budget your uptime commitment allows, so a page fires within minutes of a real outage.
  • Break SLA metrics down by region, plan tier, and critical flow so one broken segment cannot hide inside the aggregate.
  • Reserve paging for issues that are customer-facing and time-sensitive, and send lower-severity signals to a dashboard.
  • Log every case where customers noticed before monitoring did, and fix the specific check that missed it.

Involve Whoever Writes the Status Page Update

The person updating a public status page during an incident often has better visibility into what customers are actually reporting than the dashboard does, because support tickets and social mentions arrive faster than some metrics catch up, and that lag is exactly the gap this whole problem lives in.

Feed that channel back into your monitoring design. If support consistently hears about an outage before the dashboard reflects it, that's a specific, addressable gap between what customers experience and what you measure, worth fixing directly rather than accepting it as simply how things are, quarter after quarter, incident after incident, year after year.

Executive Capability Standard

What Good Looks Like

Working SLA enforcement alerts on what customers actually experience, calibrated against your real uptime budget, broken down by region and flow rather than a single aggregate number.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Review your last few real incidents and measure the gap between when monitoring detected them and when a customer noticed.
2. Do Manually:Manually add a synthetic check for your one or two most critical user flows and watch it for a week before automating alerts on it.
3. Delegate:Assign an on-call engineer ownership of alert tuning, with authority to mute or adjust noisy alerts without asking permission each time.
4. Automate:Automate alert thresholds tied to your actual uptime budget instead of a static error rate picked once at launch.
5. Buy:Bring in a monitoring or observability specialist if your team has never built synthetic monitoring before a launch that needs it.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Frequently Asked Questions

How many synthetic checks do we actually need for a small product?

Start with the two or three flows that would be a genuine emergency if broken, usually sign-in and whatever the core action of your product is. You don't need synthetic coverage of every page, just the handful where an outage would actually hurt customers and revenue.

Should every alert page someone immediately?

No. Reserve paging for issues that are customer-facing and time-sensitive. Route everything else to a dashboard or a lower-urgency channel someone checks during business hours, so the alerts that do page carry real weight and get acted on quickly.

What's the fastest way to find our current monitoring blind spots?

Look back at your last two or three real incidents and ask, for each one, how long it took for monitoring to detect it versus how long it took a customer to notice. Any gap there points directly at where your synthetic checks or alert thresholds need work.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides