A Checklist for Synthetic Monitoring That Actually Catches Outages Early
Synthetic monitoring means running a scripted, fake transaction against your product on a schedule, logging in, adding an item to a cart, submitting a form, and alerting when it fails. Done well, it catches the failures that matter to customers before a customer reports them. Done poorly, it's a source of false alarms that gets muted within a month and provides no coverage at all.
The difference is almost entirely in the setup. Here's the checklist and the pitfalls that undermine it.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Pick Transactions That Reflect Real Revenue or Real Risk
Start with the handful of flows that would actually hurt if they broke: signup, login, checkout, whatever the core action of your product is. Resist the urge to synthetically monitor every page. A probe on your marketing homepage tells you less than a probe on your checkout flow, but it's tempting to add anyway because it's easy. Every additional probe is something someone has to maintain when the UI changes, so each one should earn its place.
Run Probes From Where Your Customers Actually Are
A probe running from the same cloud region as your servers will miss latency and routing problems that a customer three regions away experiences every day. If you have customers in multiple geographies, run probes from multiple locations, not just your own infrastructure's home region. This is also the only reliable way to catch a DNS or CDN misconfiguration that's invisible from inside your own network but very visible to an actual customer.
Set Thresholds Based on What's Actually Abnormal, Not a Round Number
A common mistake is picking an alert threshold that sounds reasonable, three seconds, five seconds, without checking what your probe's normal response time actually looks like. Pull two weeks of probe history first, and set the threshold a meaningful distance above your actual p95, not an arbitrary round number. A threshold set too tight fires constantly and gets ignored; a threshold set too loose misses the slow degradation that precedes a full outage.
The Checklist Before You Call It Done
Before trusting a synthetic monitoring setup, confirm each of these:
- Every probe covers a transaction, not just a page load, since a page that loads but a form that silently fails is the failure mode that hurts customers most.
- Probes run from at least two geographic locations if you have customers outside your primary server region.
- Alert thresholds are set from real historical data, not a guess.
- Alerts route to whoever is actually on call, tested with a real page, not assumed to work.
- Each probe has an owner who updates its script when the underlying flow changes, so it doesn't start failing for reasons that have nothing to do with an actual outage.
Where Teams Let Synthetic Monitoring Rot
The most common failure isn't a missing probe, it's a stale one. A UI change breaks a login probe's selector, the probe starts failing constantly for a reason that has nothing to do with an outage, someone mutes the alert to stop the noise, and it stays muted for months. Give every probe a named owner, review the list quarterly, and treat a probe that's been failing for more than a day as either a real incident or a maintenance task, never as something to just silence and forget.
Connecting Probes to Your Actual Uptime Target
How aggressively you invest in synthetic coverage should track your own uptime target, not an arbitrary sense of thoroughness. A team promising 99.9 percent availability has a downtime budget measured in hours a year, not days1. A probe that only checks infrequently could let an outage burn through a meaningful chunk of that budget before anyone notices it. If your target is looser, checking less often might genuinely be enough, and spending engineering time tightening probe frequency further wouldn't buy you anything real. Decide the frequency from the budget, not the other way around.
A Worked Example: Catching a Silent Checkout Failure
Say a payment provider's SDK ships an update that silently swallows a specific card-decline error instead of surfacing it, so checkout appears to load fine but a subset of real transactions fail with no visible error on the page. A page-load probe would never catch this, since the page renders normally. A transaction probe that actually completes a test purchase with a known test card, and checks for a specific success confirmation rather than just an HTTP 200, would fail immediately. This is the concrete difference between monitoring that a page loads and monitoring that a transaction works, and it's usually the second one that matches what a customer actually experiences.
What Good Looks Like
Good synthetic monitoring means every core revenue or signup flow has a scripted probe with a data-driven alert threshold, a named owner, and alerts that have actually been tested to reach the right person.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
How often should synthetic probes actually run?
For a critical revenue flow like checkout, every one to five minutes is typical, since that's tight enough to catch a partial outage quickly without generating excessive cost. For lower-priority flows, every fifteen to thirty minutes is usually enough. Match the frequency to how much damage a delay in detection would actually cause.
Do synthetic probes replace real user monitoring?
No, they answer different questions. Synthetic probes tell you a specific scripted path works, on a schedule, from a controlled environment. Real user monitoring tells you what actual customers are experiencing across every device, browser, and network condition. You want both: synthetic for early, consistent detection, real user monitoring for the full picture.
What should we do when a probe fails but we can't reproduce the issue manually?
Treat it as real until proven otherwise rather than dismissing it. Check the probe's raw response and screenshot if your tool captures one, and check for regional or timing patterns, since an intermittent failure that a probe catches consistently is often a real edge case a manual check happens to miss.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
Synthetic Monitoring: Testing the Paths Users Take
A green uptime dashboard can hide a broken checkout for hours. How to pick the handful of flows worth simulating and alert on them well.
Stress-Testing a System Without Taking Down Real Traffic
How to run a synthetic load test that finds where a system actually breaks, without accidentally taking down production traffic in the process.
Your Uptime Monitor Looks Fine. Your Customers Disagree
A checklist for building synthetic monitoring that catches what a basic uptime check misses, and the common mistakes that leave it blind to real outages.
Synthetic Monitoring That Watches What Customers Actually Do
A checklist for synthetic transaction monitoring: which journeys to probe, how to avoid alert fatigue, and where synthetic checks miss what real users hit.
What Synthetic Monitoring Catches That Your Alerts Don't
How synthetic transaction probes catch outages that server metrics and error-rate alerts miss, and how to set them up without drowning in false alarms.
Synthetic Monitoring: Catching Outages Before Customers Do
How to design synthetic transaction probes that catch a real outage instead of false alarms, and where they can't replace real user monitoring.