Cloud FinOps & Infrastructure ScalingPlaybook3 min readUpdated September 2026

A Checklist for Synthetic Monitoring That Actually Catches Outages Early

Synthetic monitoring means running a scripted, fake transaction against your product on a schedule, logging in, adding an item to a cart, submitting a form, and alerting when it fails. Done well, it catches the failures that matter to customers before a customer reports them. Done poorly, it's a source of false alarms that gets muted within a month and provides no coverage at all.

The difference is almost entirely in the setup. Here's the checklist and the pitfalls that undermine it.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Pick Transactions That Reflect Real Revenue or Real Risk

Start with the handful of flows that would actually hurt if they broke: signup, login, checkout, whatever the core action of your product is. Resist the urge to synthetically monitor every page. A probe on your marketing homepage tells you less than a probe on your checkout flow, but it's tempting to add anyway because it's easy. Every additional probe is something someone has to maintain when the UI changes, so each one should earn its place.

Run Probes From Where Your Customers Actually Are

A probe running from the same cloud region as your servers will miss latency and routing problems that a customer three regions away experiences every day. If you have customers in multiple geographies, run probes from multiple locations, not just your own infrastructure's home region. This is also the only reliable way to catch a DNS or CDN misconfiguration that's invisible from inside your own network but very visible to an actual customer.

Set Thresholds Based on What's Actually Abnormal, Not a Round Number

A common mistake is picking an alert threshold that sounds reasonable, three seconds, five seconds, without checking what your probe's normal response time actually looks like. Pull two weeks of probe history first, and set the threshold a meaningful distance above your actual p95, not an arbitrary round number. A threshold set too tight fires constantly and gets ignored; a threshold set too loose misses the slow degradation that precedes a full outage.

The Checklist Before You Call It Done

Before trusting a synthetic monitoring setup, confirm each of these:

  • Every probe covers a transaction, not just a page load, since a page that loads but a form that silently fails is the failure mode that hurts customers most.
  • Probes run from at least two geographic locations if you have customers outside your primary server region.
  • Alert thresholds are set from real historical data, not a guess.
  • Alerts route to whoever is actually on call, tested with a real page, not assumed to work.
  • Each probe has an owner who updates its script when the underlying flow changes, so it doesn't start failing for reasons that have nothing to do with an actual outage.

Where Teams Let Synthetic Monitoring Rot

The most common failure isn't a missing probe, it's a stale one. A UI change breaks a login probe's selector, the probe starts failing constantly for a reason that has nothing to do with an outage, someone mutes the alert to stop the noise, and it stays muted for months. Give every probe a named owner, review the list quarterly, and treat a probe that's been failing for more than a day as either a real incident or a maintenance task, never as something to just silence and forget.

Connecting Probes to Your Actual Uptime Target

How aggressively you invest in synthetic coverage should track your own uptime target, not an arbitrary sense of thoroughness. A team promising 99.9 percent availability has a downtime budget measured in hours a year, not days1. A probe that only checks infrequently could let an outage burn through a meaningful chunk of that budget before anyone notices it. If your target is looser, checking less often might genuinely be enough, and spending engineering time tightening probe frequency further wouldn't buy you anything real. Decide the frequency from the budget, not the other way around.

A Worked Example: Catching a Silent Checkout Failure

Say a payment provider's SDK ships an update that silently swallows a specific card-decline error instead of surfacing it, so checkout appears to load fine but a subset of real transactions fail with no visible error on the page. A page-load probe would never catch this, since the page renders normally. A transaction probe that actually completes a test purchase with a known test card, and checks for a specific success confirmation rather than just an HTTP 200, would fail immediately. This is the concrete difference between monitoring that a page loads and monitoring that a transaction works, and it's usually the second one that matches what a customer actually experiences.

Executive Capability Standard

What Good Looks Like

Good synthetic monitoring means every core revenue or signup flow has a scripted probe with a data-driven alert threshold, a named owner, and alerts that have actually been tested to reach the right person.

Building The Capability (5-Stage Skill Ladder)

1. Learn:List your product's core transactions and check which ones currently have no synthetic monitoring coverage at all.
2. Do Manually:Manually walk through your top uncovered flow on a schedule for a couple of weeks to understand its normal behavior before scripting a probe.
3. Delegate:Assign a named owner to each synthetic probe who's responsible for updating its script when the underlying flow changes.
4. Automate:Set up scripted synthetic probes from multiple regions with alert thresholds derived from real historical response times.
5. Buy:Use a dedicated synthetic monitoring or observability platform once you have more than a handful of critical flows to cover, rather than maintaining custom scripts yourself.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

ClickUp

When a probe does catch something, logging the incident and the follow-up fix in a shared tracker like ClickUp keeps a record of which flows have actually broken before, which is useful the next time you're deciding what deserves a probe.

Visit ClickUp→

Frequently Asked Questions

How often should synthetic probes actually run?

For a critical revenue flow like checkout, every one to five minutes is typical, since that's tight enough to catch a partial outage quickly without generating excessive cost. For lower-priority flows, every fifteen to thirty minutes is usually enough. Match the frequency to how much damage a delay in detection would actually cause.

Do synthetic probes replace real user monitoring?

No, they answer different questions. Synthetic probes tell you a specific scripted path works, on a schedule, from a controlled environment. Real user monitoring tells you what actual customers are experiencing across every device, browser, and network condition. You want both: synthetic for early, consistent detection, real user monitoring for the full picture.

What should we do when a probe fails but we can't reproduce the issue manually?

Treat it as real until proven otherwise rather than dismissing it. Check the probe's raw response and screenshot if your tool captures one, and check for regional or timing patterns, since an intermittent failure that a probe catches consistently is often a real edge case a manual check happens to miss.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides