Distributed Systems & Enterprise ResiliencePlaybook3 min readUpdated September 2026

Building Synthetic Checks That Catch an Outage Before Your Customers Do

Synthetic monitoring catches outages early when a few probes exercise your most important transactions, run from where customers actually are, and page only after repeated failures. Built carelessly, the same checks become noise that trains the team to ignore alerts, while dashboards stay green and support tickets pile up.

Pick the transactions that matter, not every endpoint

Three to five flows that represent real revenue or core value, such as login, checkout, or the specific API call your customers' businesses depend on, are worth a dedicated synthetic probe each. A probe for every endpoint in your system produces a wall of checks that nobody can triage during an incident, and most of them will never catch anything a customer would actually notice. Start from 'what would a customer notice is broken' and work backward to the transaction that exercises it end to end, not from a list of routes in your API.

A good synthetic check also exercises the real dependency chain, not a shortcut through it. A login probe that only hits an internal health endpoint proves the process is running, not that a customer can actually authenticate; a probe that submits real credentials through the same path a browser would use, against a dedicated test account, is the version that actually catches the outage a customer would hit.

Run from where your customers actually are

A probe running from inside your own VPC monitors the happiest possible path: your own network, your own region, none of the DNS resolution or CDN routing a real customer's request goes through. If you serve customers across multiple regions, run probes from multiple external vantage points, ideally the regions and networks your actual traffic comes from. An outage that's specific to one region's routing or one CDN edge location is invisible to a single internal probe and very visible to the customers actually affected by it.

Tune cadence and alert thresholds deliberately

Run probes often enough to catch a brief outage before a customer files a ticket about it, but require two or three consecutive failures before paging, not a single blip; a single failed check is at least as likely to be a transient network hiccup between the probe and your service as a real outage. If your target is three nines of uptime, you're budgeting well under nine hours of downtime a year for the whole system1. A synthetic check that only notices an outage after a customer emails support has already spent a meaningful slice of that budget on detection delay alone, before anyone's even started fixing anything.

Tie a failure to a runbook, not just a page

Each synthetic check should point whoever gets paged to the specific dashboard for that transaction and the first two or three things worth checking, not a generic 'investigate' notification. A probe that fires at 3 a.m. with no context costs the on-call engineer the first ten minutes just figuring out where to look, which is exactly the time an outage is costing customers the most. The runbook doesn't need to be exhaustive; it needs to shortcut the first, most common diagnosis step.

What synthetic monitoring won't tell you

A synthetic check tells you a transaction is broken. It doesn't tell you why, which service in the chain is actually failing, or whether the failure is isolated to a subset of real users the probe's specific path doesn't represent. Pair it with real user monitoring and distributed tracing, or the team ends up with a clear 'it's down' signal and no path from there to root cause, which just moves the delay from detection to diagnosis instead of removing it. Comparing observability platforms is a reasonable next step once synthetic checks are catching outages reliably and the gap becomes diagnosis speed.

It's also worth remembering a synthetic check only tests the one path it was written to test. A checkout probe that always uses the same test card, the same shipping address, and the same product doesn't catch a bug specific to a different payment method or a different region's tax calculation. Rotate the probe's inputs occasionally, or add a second probe for the next most common variant, rather than treating one green check as proof the whole flow works for every customer.

A workable synthetic monitoring setup covers these points:

  • Probe a handful of flows that represent revenue or core value, such as login or checkout, end to end with a dedicated test account.
  • Run probes from several external vantage points that match where your real traffic comes from, not only from inside your own network.
  • Require a few consecutive failures before paging, so a single network blip doesn't wake anyone.
  • Link each check to a dashboard and the first things to inspect, so the on-call engineer isn't starting from nothing.
  • Pair the probes with real user monitoring and distributed tracing, since a probe says something is broken but not why.
Executive Capability Standard

What Good Looks Like

Good synthetic monitoring covers the handful of transactions that represent real customer value, runs from where customers actually are, and pages on a sustained failure with a runbook attached, not on a single blip with no context.

Building The Capability (5-Stage Skill Ladder)

1. Learn:List the three to five transactions that represent your core customer value or revenue, and identify the regions or networks your real traffic comes from.
2. Do Manually:Manually run each of those transactions from an external vantage point periodically and compare the experience with what your internal dashboards say.
3. Delegate:Have an engineer set up automated synthetic checks for those transactions with multi-region coverage and tuned alert thresholds.
4. Automate:Wire each check's failure alert to a specific runbook and dashboard link, so on-call has a starting point instead of a bare notification.
5. Buy:Bring in observability or SRE advisory if you're standing up synthetic monitoring for the first time across a multi-region system and want the coverage right from the start.

How to Get Started

Frequently Asked Questions

How many synthetic checks should we actually run?

Start with the three to five transactions that represent your core value or revenue, not every endpoint you have. It's easier to add a check for a flow that later turns out to matter than to triage a wall of checks nobody trusts during an actual incident.

Why require multiple consecutive failures before paging?

A single failed probe is often a transient network issue between the probe and your service rather than a real outage. Requiring two or three consecutive failures filters most of that noise out while still catching a real outage within roughly the interval you're running the probes at.

Is synthetic monitoring a replacement for real user monitoring?

No. Synthetic checks tell you a specific transaction from a specific vantage point is broken; they don't capture what your actual, varied user base is experiencing or help you diagnose why. Use synthetic checks for fast detection and real user monitoring plus tracing for understanding scope and root cause.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides