Synthetic Monitoring: Catching Outages Before Customers Do
Real user monitoring tells you something broke once a real customer hits it. Synthetic monitoring is how you find out before that happens: a scripted probe runs the same critical path a customer would, on a schedule, from outside your infrastructure, and alerts the moment it fails.
The hard part isn't running the probe. It's designing one specific enough to catch a real problem without paging your on-call team over noise.
What to probe: critical paths, not every endpoint
A synthetic probe on every API endpoint sounds thorough and produces mostly noise. Start with the handful of paths where a failure directly costs revenue or trust: login, checkout, the core action your product exists to perform.
Script each probe to actually complete the transaction, not just hit a health check endpoint, since a basic health route returning success tells you almost nothing about whether a customer can actually log in or complete a purchase.
Protecting your uptime budget, not just watching for total outages
If your SLA promises three nines of availability, you're working with an annual error budget of 8.76 hours of downtime, full stop, whether an outage shows up as a customer support ticket or not1.
Synthetic probes are how you find out you're burning that budget in real time, from a partial degradation that never fully takes the site down but still fails a meaningful share of checkout attempts. A probe running every few minutes catches that kind of partial failure long before a slow trickle of support tickets does.
Avoiding false alarms
- Run probes from more than one region and require two consecutive failures before paging, so a transient network blip between the probe location and your infrastructure doesn't wake anyone up for nothing.
- Keep probe credentials and test data separate from production, and flag the probe's actions, a test purchase, a test signup, clearly so they don't pollute your real analytics or trigger real fraud checks.
- Alert on the specific step that failed, not just that the checkout probe failed, so the on-call engineer isn't starting from zero when the page fires overnight.
Where synthetic monitoring can't substitute for real user monitoring
A synthetic probe runs the same script from the same handful of locations every time, which means it will never catch a failure specific to a particular browser, device, ISP, or geography that your actual user base includes. It also can't tell you about slowness that's still technically working, like a checkout that completes but takes noticeably longer than usual for real users on mobile networks.
Run both: synthetic probes for fast detection of hard failures on critical paths, real user monitoring for the broader picture of what your actual customers are experiencing.
Where to run probes from
A probe running from the same cloud region as your infrastructure catches an application-level failure but misses anything specific to how the wider internet reaches you: a DNS resolution problem, a CDN misconfiguration, a network path issue affecting one part of the world. Run probes from at least two or three geographically distinct locations, ideally including wherever a meaningful share of your actual customers are, so a regional problem doesn't hide behind a probe that happens to have a clean path to your servers.
A worked example: a partial payment failure
Say a payment provider integration starts silently failing for one specific card network while every other transaction type keeps working. Total transaction volume barely dips, since most traffic uses other card networks, so an alert on overall error rate or total checkout volume never fires. A scripted probe that walks the full checkout flow with a specific test card on that network would catch it within minutes; a probe that only confirms the checkout page loads would not.
This is the general shape of the failures synthetic monitoring exists for: not the outage that takes everything down, but the narrower failure that's easy to miss in an aggregate metric and easy to catch with a scripted probe built around one specific real path.
What Good Looks Like
A good synthetic monitoring setup scripts complete transactions on your handful of truly critical paths, alerts on the specific step that failed, and requires confirmation from more than one location before paging anyone.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How often should a synthetic probe run?
For a critical path like checkout, every few minutes is typical, since the goal is catching a failure fast enough that it's resolved before it shows up as a wave of support tickets. Less critical paths can run less frequently to keep monitoring costs and noise down.
Why does our synthetic monitoring keep paging us for nothing?
The most common cause is alerting on a single failure from a single probe location, which catches ordinary transient network blips between the probe and your infrastructure. Requiring two consecutive failures, ideally from more than one region, cuts most false alarms without meaningfully slowing real detection.
Do we still need real user monitoring if we have synthetic probes?
Yes. Synthetic probes run the same script from a handful of fixed locations, so they can't catch a failure specific to a particular device, browser, or region your real users are on. Use synthetic monitoring for fast detection of hard failures and real user monitoring for the fuller picture.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
Continuous Device Verification for a Zero-Trust API
How continuous device and identity verification actually works in a zero-trust architecture, and where to draw the line for a small engineering team.
Load Testing an Authenticated API Without Setting Off Your Own Defenses
Four safeguards for load testing a zero trust API so the test doesn't trip rate limits, skew results with one shared identity, or miss the real bottleneck.
Synthetic Monitoring: Testing the Paths Users Take
A green uptime dashboard can hide a broken checkout for hours. How to pick the handful of flows worth simulating and alert on them well.
Catching SLA Breaches Before Your Customers Do
How to build automated SLA breach detection that catches an availability or latency problem before a customer has to report it to you first.
Your Uptime Monitor Looks Fine. Your Customers Disagree
A checklist for building synthetic monitoring that catches what a basic uptime check misses, and the common mistakes that leave it blind to real outages.
A Checklist for Synthetic Monitoring That Actually Catches Outages Early
A practical checklist for setting up synthetic transaction probes that catch real customer-facing failures, plus the common pitfalls that make them useless.