Why Your SLA Dashboard Says Green While Customers Are Down
A customer emails to say the app has been down for twenty minutes. Your dashboard says every service is healthy. This gap, between what your monitoring measures and what a customer actually experiences, is the most common reason SLA enforcement fails, and it usually isn't a monitoring outage, it's a monitoring design problem.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Why do health checks miss real outages?
Most health checks confirm a process is running and can respond to a ping. That's not the same as confirming the feature a customer relies on actually works end to end. A service can return a healthy check while the specific database query behind your login page times out for every real request.
Build at least one synthetic check per critical user flow, actually logging in, actually loading the dashboard, on the same schedule as your infrastructure health checks. The difference between "the process is up" and "the feature works" is exactly the gap where customers notice problems before your monitoring does.
Your Uptime Budget Should Set Your Alert Thresholds
If you've committed to an uptime target, your alerting should be calibrated against the actual downtime budget that target allows, not an arbitrary error rate someone picked at launch1. A tighter commitment needs faster detection and a lower tolerance for sustained errors before paging someone.
Work backward from the budget: if your annual allowance is small, your alert has to fire within minutes of a real outage starting, not after an hour of degraded performance quietly burns through a chunk of the year's allowance in one incident.
For example, suppose your monitoring alerts only when the error rate stays high for a full hour. If a bad release breaks sign-in at the start of that window, a large slice of the year's downtime allowance can disappear before anyone is paged. Working backward from the budget shows that the alert needs to fire within minutes, which usually means a short evaluation window on a synthetic sign-in check rather than a long window on aggregate errors. Then test it: break the check on purpose in a staging environment and time how long the page takes to arrive.
How do aggregate metrics hide localized outages?
An SLA dashboard built on aggregate success rate across all traffic can stay green while one customer, one region, or one feature is completely broken, simply because the volume of everything else drowns out the failure in the average.
Break your SLA metrics down by the dimensions that matter to your actual customers: by region, by plan tier, by the specific critical flow. A single aggregate number is convenient for a status page and nearly useless for catching the outage a customer is calling you about right now.
Alert Fatigue Is Why Real Alerts Get Ignored
A team that gets paged for every minor blip stops trusting alerts within a few weeks, and the real outage alert arrives in the same noisy channel as everything else. This is often the actual root cause when a team says "the alert fired, we just didn't act on it fast enough."
Tune thresholds so a page means something is genuinely wrong, and route lower-severity signals to a dashboard someone checks, not a channel that wakes someone up. Fewer, more trustworthy alerts beat more, ignorable ones every time.
Close the Loop With a Real Postmortem
Every time the dashboard said green while a customer was actually down, that's a specific, fixable gap in what you measure. Track these separately from ordinary incidents, because they're telling you something different: not that something broke, but that your detection has a blind spot.
Fix the specific check, not just the incident. A team that patches the immediate outage without asking why monitoring missed it will find the same blind spot again, usually at a worse time.
Close the gap between the dashboard and the customer with these checks:
- Add a synthetic check for each critical user flow, such as signing in and loading the dashboard, on the same schedule as infrastructure health checks.
- Set alert thresholds from the downtime budget your uptime commitment allows, so a page fires within minutes of a real outage.
- Break SLA metrics down by region, plan tier, and critical flow so one broken segment cannot hide inside the aggregate.
- Reserve paging for issues that are customer-facing and time-sensitive, and send lower-severity signals to a dashboard.
- Log every case where customers noticed before monitoring did, and fix the specific check that missed it.
Involve Whoever Writes the Status Page Update
The person updating a public status page during an incident often has better visibility into what customers are actually reporting than the dashboard does, because support tickets and social mentions arrive faster than some metrics catch up, and that lag is exactly the gap this whole problem lives in.
Feed that channel back into your monitoring design. If support consistently hears about an outage before the dashboard reflects it, that's a specific, addressable gap between what customers experience and what you measure, worth fixing directly rather than accepting it as simply how things are, quarter after quarter, incident after incident, year after year.
What Good Looks Like
Working SLA enforcement alerts on what customers actually experience, calibrated against your real uptime budget, broken down by region and flow rather than a single aggregate number.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Track postmortem follow-ups from missed alerts as tasks in ClickUp so the fix doesn't get lost after the incident review meeting.
Document your on-call escalation runbook in Trainual so alert response doesn't depend on one person's memory of what each alert means.
Frequently Asked Questions
How many synthetic checks do we actually need for a small product?
Start with the two or three flows that would be a genuine emergency if broken, usually sign-in and whatever the core action of your product is. You don't need synthetic coverage of every page, just the handful where an outage would actually hurt customers and revenue.
Should every alert page someone immediately?
No. Reserve paging for issues that are customer-facing and time-sensitive. Route everything else to a dashboard or a lower-urgency channel someone checks during business hours, so the alerts that do page carry real weight and get acted on quickly.
What's the fastest way to find our current monitoring blind spots?
Look back at your last two or three real incidents and ask, for each one, how long it took for monitoring to detect it versus how long it took a customer to notice. Any gap there points directly at where your synthetic checks or alert thresholds need work.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
Why Your SLA Dashboard Doesn't Know You Breached an SLA
An uptime dashboard is not SLA monitoring. Here is how to define a breach precisely enough to detect it automatically, before a customer emails about it.
The Common Mistakes That Make Automated SLA Alerts Untrustworthy
The specific mistakes that turn automated SLA breach detection into noise nobody responds to, and what to fix in each one before adding more alerts.
Why Your SLA Alerts Keep Missing Real Breaches
Why polling-based SLA monitoring breaks down on a real-time pipeline at scale, and how to detect breaches from the event stream itself instead.
Why Your SLA Alerts Stop Firing Once You Scale
Why the alerting setup that caught every SLA breach with five services quietly stops working at fifty, and what to fix before a customer finds the gap first.
Catching SLA Breaches Before Your Customers Do
How to build automated SLA breach detection that catches an availability or latency problem before a customer has to report it to you first.
Why Your SLA Monitoring Keeps Missing Real Breaches
Why synthetic uptime checks miss real SLA breaches, how to build monitoring that matches the contract you actually signed, and what to do once one is confirmed.