Building Synthetic Checks That Catch an Outage Before Your Customers Do
Synthetic monitoring catches outages early when a few probes exercise your most important transactions, run from where customers actually are, and page only after repeated failures. Built carelessly, the same checks become noise that trains the team to ignore alerts, while dashboards stay green and support tickets pile up.
Pick the transactions that matter, not every endpoint
Three to five flows that represent real revenue or core value, such as login, checkout, or the specific API call your customers' businesses depend on, are worth a dedicated synthetic probe each. A probe for every endpoint in your system produces a wall of checks that nobody can triage during an incident, and most of them will never catch anything a customer would actually notice. Start from 'what would a customer notice is broken' and work backward to the transaction that exercises it end to end, not from a list of routes in your API.
A good synthetic check also exercises the real dependency chain, not a shortcut through it. A login probe that only hits an internal health endpoint proves the process is running, not that a customer can actually authenticate; a probe that submits real credentials through the same path a browser would use, against a dedicated test account, is the version that actually catches the outage a customer would hit.
Run from where your customers actually are
A probe running from inside your own VPC monitors the happiest possible path: your own network, your own region, none of the DNS resolution or CDN routing a real customer's request goes through. If you serve customers across multiple regions, run probes from multiple external vantage points, ideally the regions and networks your actual traffic comes from. An outage that's specific to one region's routing or one CDN edge location is invisible to a single internal probe and very visible to the customers actually affected by it.
Tune cadence and alert thresholds deliberately
Run probes often enough to catch a brief outage before a customer files a ticket about it, but require two or three consecutive failures before paging, not a single blip; a single failed check is at least as likely to be a transient network hiccup between the probe and your service as a real outage. If your target is three nines of uptime, you're budgeting well under nine hours of downtime a year for the whole system1. A synthetic check that only notices an outage after a customer emails support has already spent a meaningful slice of that budget on detection delay alone, before anyone's even started fixing anything.
Tie a failure to a runbook, not just a page
Each synthetic check should point whoever gets paged to the specific dashboard for that transaction and the first two or three things worth checking, not a generic 'investigate' notification. A probe that fires at 3 a.m. with no context costs the on-call engineer the first ten minutes just figuring out where to look, which is exactly the time an outage is costing customers the most. The runbook doesn't need to be exhaustive; it needs to shortcut the first, most common diagnosis step.
What synthetic monitoring won't tell you
A synthetic check tells you a transaction is broken. It doesn't tell you why, which service in the chain is actually failing, or whether the failure is isolated to a subset of real users the probe's specific path doesn't represent. Pair it with real user monitoring and distributed tracing, or the team ends up with a clear 'it's down' signal and no path from there to root cause, which just moves the delay from detection to diagnosis instead of removing it. Comparing observability platforms is a reasonable next step once synthetic checks are catching outages reliably and the gap becomes diagnosis speed.
It's also worth remembering a synthetic check only tests the one path it was written to test. A checkout probe that always uses the same test card, the same shipping address, and the same product doesn't catch a bug specific to a different payment method or a different region's tax calculation. Rotate the probe's inputs occasionally, or add a second probe for the next most common variant, rather than treating one green check as proof the whole flow works for every customer.
A workable synthetic monitoring setup covers these points:
- Probe a handful of flows that represent revenue or core value, such as login or checkout, end to end with a dedicated test account.
- Run probes from several external vantage points that match where your real traffic comes from, not only from inside your own network.
- Require a few consecutive failures before paging, so a single network blip doesn't wake anyone.
- Link each check to a dashboard and the first things to inspect, so the on-call engineer isn't starting from nothing.
- Pair the probes with real user monitoring and distributed tracing, since a probe says something is broken but not why.
What Good Looks Like
Good synthetic monitoring covers the handful of transactions that represent real customer value, runs from where customers actually are, and pages on a sustained failure with a runbook attached, not on a single blip with no context.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How many synthetic checks should we actually run?
Start with the three to five transactions that represent your core value or revenue, not every endpoint you have. It's easier to add a check for a flow that later turns out to matter than to triage a wall of checks nobody trusts during an actual incident.
Why require multiple consecutive failures before paging?
A single failed probe is often a transient network issue between the probe and your service rather than a real outage. Requiring two or three consecutive failures filters most of that noise out while still catching a real outage within roughly the interval you're running the probes at.
Is synthetic monitoring a replacement for real user monitoring?
No. Synthetic checks tell you a specific transaction from a specific vantage point is broken; they don't capture what your actual, varied user base is experiencing or help you diagnose why. Use synthetic checks for fast detection and real user monitoring plus tracing for understanding scope and root cause.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
Datadog vs New Relic vs Dynatrace: Cloud Observability Platforms Compared
Compare Datadog, New Relic, and Dynatrace for cloud observability: log ingestion costs, distributed tracing, APM overhead, and MTTR compression.
Stress Testing Without Taking Down the System You're Trying to Protect
How to run stress tests aggressive enough to find real breaking points without risking the production system or the customers depending on it.
Synthetic Monitoring: Testing the Paths Users Take
A green uptime dashboard can hide a broken checkout for hours. How to pick the handful of flows worth simulating and alert on them well.
Why Your SLA Monitoring Keeps Missing Real Breaches
Why synthetic uptime checks miss real SLA breaches, how to build monitoring that matches the contract you actually signed, and what to do once one is confirmed.
A Production Deployment Checklist That Actually Catches Problems
A stage-by-stage deployment checklist for distributed systems, covering rollback readiness, dependency ordering, and the checks teams skip under pressure.
Your Uptime Monitor Looks Fine. Your Customers Disagree
A checklist for building synthetic monitoring that catches what a basic uptime check misses, and the common mistakes that leave it blind to real outages.