Catching SLA Breaches Before Your Customers Do
If a customer is the one who tells you about an SLA breach, your monitoring already failed at its main job. Automated SLA enforcement means your own systems know you're in breach, or about to be, before support gets the email. Here's how to build that instead of relying on dashboards someone has to remember to check.
How do you turn an SLA into something monitoring can measure?
A contractual SLA often uses language like available or responsive that doesn't map cleanly to a single metric. Before you can automate detection, translate the SLA into a specific, measurable definition: which endpoints count, what response counts as a failure, and over what rolling window availability is calculated. If your legal SLA language and your monitoring definition disagree, even slightly, you can be technically breaching the contract while your dashboards show green, or the reverse.
Measure from outside your own infrastructure
Monitoring that runs inside your own network can miss the exact failure mode that matters most to a customer: your API being reachable from your own servers but not from the outside world, due to a DNS issue, a CDN misconfiguration, or a network path problem between you and your customers. Run synthetic checks from external locations that approximate where your actual customers connect from, not just internal health checks, so what you measure matches what your SLA actually promises.
Where should your internal alert threshold sit?
If your SLA promises a monthly availability figure, don't set your internal alert to fire only once you've already crossed it. Set an earlier internal warning threshold with enough runway to intervene before the contractual line is crossed, based on your current rate of downtime accumulating through the month. This turns SLA management from a monthly report you dread into an early warning system that gives you a chance to actually prevent the breach, not just document it after the fact.
For example, suppose your contract measures availability over a month, and a bad afternoon early in the month has already used up much of your allowance. A single alert at the contractual line would tell you nothing until it was too late. An earlier internal warning, based on how quickly downtime is accumulating, gives the on-call owner time to stop the damage, slow risky deploys, or warn the customer. The common mistake is treating the contract number as the alert number. Treat it as the finish line, and put your warning far enough before it that someone can still act.
Downtime budgets are smaller than most teams assume
It's easy to underestimate how little downtime a high availability target actually allows. Google's widely used SRE reference table puts the allowed downtime for a 99.99% availability target at 52.6 minutes a year1. Once you calculate your own target's actual budget in minutes rather than as a percentage, it becomes obvious why automated detection matters: a single unremediated incident can consume a meaningful share of an entire year's allowance in one sitting.
Automate the escalation, not just the detection
Detecting a breach in progress only helps if it reaches someone who can act on it immediately, at any hour, not just during business hours when someone happens to be watching a dashboard. Wire your SLA threshold alerts into the same on-call escalation path as your other production incidents, with a clear owner and a defined response expectation, rather than routing them to a channel that might not get checked until the next morning.
The steps for automated breach detection:
- Translate the contractual SLA into a measurable definition: which endpoints count, what counts as a failure, and the rolling window.
- Run synthetic checks from external locations that approximate where your customers connect from.
- Set an internal warning threshold below the contractual line, with enough runway to intervene.
- Route threshold alerts into the same on-call escalation path as your other production incidents.
- Notify affected customers yourself, and log near-misses alongside confirmed breaches.
Report your own breaches before the customer notices
When a breach does happen despite the guardrails above, proactively notifying the affected customer with what happened and what you're doing about it, before they file a ticket asking why your API was slow, changes the entire tone of that conversation. It also builds the kind of track record that makes SLA conversations at renewal time far less adversarial, since the customer already trusts that you catch problems yourself rather than waiting to be told.
Keep a running log of near-misses, not just confirmed breaches
An alert that fired and got resolved before crossing the contractual line is still useful information: it tells you where your margin is thinnest and which part of your system tends to be the first to degrade. Keep a lightweight record of these near-misses alongside actual breaches, and review both together on a regular cadence, since a pattern of frequent near-misses in the same area is often an early signal of a capacity or architecture problem worth fixing before it becomes a real breach.
What Good Looks Like
Good SLA enforcement means your own systems detect and escalate an approaching breach automatically, with enough lead time to act, rather than relying on a customer complaint or a manually checked dashboard.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How much earlier should our internal alert threshold be than our contractual SLA?
Enough to give your team real time to intervene before the contractual line is crossed, which depends on how quickly your typical incidents get resolved. A team with a fast incident response might set the threshold closer to the line than a team where remediation typically takes longer.
Should SLA monitoring run from inside our own infrastructure or externally?
Externally, from locations that approximate where your real customers connect from. Internal monitoring can look healthy while customers experience a real outage caused by a DNS, CDN, or network path issue between your infrastructure and the outside world.
Is it worth proactively telling customers about a breach they haven't noticed yet?
Generally yes, since it demonstrates you're monitoring the SLA seriously and gives you control over the framing instead of reacting defensively to a complaint. This matters more, not less, for breaches a customer would have eventually noticed anyway.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
Continuous Device Verification for a Zero-Trust API
How continuous device and identity verification actually works in a zero-trust architecture, and where to draw the line for a small engineering team.
Why Your SLA Dashboard Says Green While Customers Are Down
Why automated SLA monitoring so often shows green during a real outage, and how to build alerting that actually reflects what customers experience.
Synthetic Monitoring: Catching Outages Before Customers Do
How to design synthetic transaction probes that catch a real outage instead of false alarms, and where they can't replace real user monitoring.
Rolling Out Zero Trust in Production Without a Broad Outage
A checklist for rolling out stricter API authentication and authorization in production, and the pitfalls that turn a rollout into an incident.
The Common Mistakes That Make Automated SLA Alerts Untrustworthy
The specific mistakes that turn automated SLA breach detection into noise nobody responds to, and what to fix in each one before adding more alerts.
Why Your SLA Dashboard Doesn't Know You Breached an SLA
An uptime dashboard is not SLA monitoring. Here is how to define a breach precisely enough to detect it automatically, before a customer emails about it.