Synthetic Monitoring That Watches What Customers Actually Do
A health-check endpoint returning 200 tells you the process is up. It doesn't tell you whether a customer can actually complete checkout, log in, or upload a file, because those flows touch a database, a third-party payment processor, and session state that a bare health check never exercises.
This is a checklist for building synthetic probes around what actually matters, and the pitfalls that turn a good monitoring setup into a source of ignored alerts.
Probe the Journeys That Actually Make Money
List your top three or four revenue-critical paths before writing a single probe script: signup, checkout, the core action your product exists to perform. A probe that logs in and clicks around a dashboard is easy to write and tells you almost nothing compared to one that completes an actual checkout flow, including the third-party payment step most teams are tempted to stub out of the probe because it's harder to automate. The harder-to-build probe is usually the more valuable one.
Frequency and Geographic Spread Should Match Your User Base
A probe running every five minutes from a single region catches less than one running every minute from the regions your actual traffic comes from, especially for latency regressions that only show up from certain geographies due to CDN routing or a third-party dependency with regional outages. Match probe frequency to how quickly you need to know something broke, a checkout flow probably warrants tighter intervals than an internal admin page, and match geographic spread to where your customers actually are, not where your infrastructure happens to sit.
Alert Fatigue From Flaky Probes Kills the Whole Program
A probe that occasionally times out because of its own network path, not a real outage, and pages someone every time, trains the on-call rotation to ignore synthetic alerts within a month. Require consecutive failures from more than one probe location before paging, and separate "this probe is degraded" from "the service is down" as distinct alert severities. A synthetic monitoring program that's been muted by its own team has negative value: it looks like coverage exists when it doesn't.
Synthetic Monitoring and Real User Monitoring Answer Different Questions
Synthetic probes tell you a specific, scripted path works from a specific location on a schedule; real user monitoring tells you what's actually happening across every real session, every browser, every device, all the time. Synthetic catches an outage before a human notices; RUM tells you which real-world edge case, a specific browser version, a slow connection, is actually hurting users right now. Run both. Synthetic without RUM misses the long tail of real conditions; RUM without synthetic means you only find out about an outage from angry customers instead of a proactive alert.
Keeping Probes From Rotting as the UI Changes
A probe built against specific selectors or a specific flow breaks the moment product ships a redesign, and if nobody owns probe maintenance, it either starts failing constantly on false positives or gets quietly disabled and never re-enabled. Give one engineer explicit ownership of probe health as part of the release process for the flows they cover, and treat a probe failure after a deploy as a signal to check first, not dismiss, since it's just as likely to be a real regression as a stale selector.
A solid synthetic monitoring setup covers these points:
- Three or four revenue-critical journeys, such as signup, checkout, and your product's core action, instead of a bare health endpoint.
- Probe frequency and locations that match where your real traffic comes from.
- Alerts that fire only after consecutive failures from more than one probe location.
- One named engineer who owns probe maintenance as the UI changes.
- Real user monitoring alongside probes, since a probe only exercises the exact path its script follows.
What Probes Miss That You Still Need to Cover
A synthetic check exercises the exact path its script follows and nothing else, so it won't catch a bug that only appears for a specific account state, a specific plan tier, or a specific combination of feature flags outside its scripted path. Treat probes as coverage for your most critical, common paths, not a substitute for broader end-to-end test coverage or the kind of exploratory testing that catches edge cases a fixed script never will. A team that believes synthetic monitoring means the product is fully covered is usually the one surprised by a bug that lived in an untested corner for months.
Choosing Probe Locations to Match Real Failure Patterns
A third-party dependency, a CDN, a DNS provider, an upstream API, can fail regionally without your own infrastructure being at fault at all, and a probe running from only one location can't distinguish "our service is down" from "this specific region can't reach us." Running probes from a handful of geographically distinct locations, ideally overlapping with where your actual customer base sits, turns an ambiguous alert into an actionable one: a failure isolated to one probe location points at network or regional issues, while a failure across every location points squarely at your own service.
What Good Looks Like
Synthetic checks should exercise the same critical path a paying customer takes, not just a health endpoint that always returns 200.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How many synthetic probes is enough for a small engineering team?
Three to five, covering your top revenue-critical journeys, is a reasonable starting point. More than that with a small team usually means maintenance burden outpaces the signal you get, especially once the UI starts changing faster than probes get updated.
Should synthetic probes run against production or a staging environment?
Production, for the ones meant to catch real customer-facing outages. Staging probes are useful for pre-deploy smoke testing, but they don't tell you whether the actual production environment, with its real third-party integrations and real data volume, is healthy right now.
What's a reasonable threshold before a probe failure pages someone?
Two to three consecutive failures from at least two different probe locations, rather than a single failed check. This filters out transient network blips in the probe's own path while still catching a genuine outage within a few minutes.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Synthetic Monitoring: Testing the Paths Users Take
A green uptime dashboard can hide a broken checkout for hours. How to pick the handful of flows worth simulating and alert on them well.
Why Synthetic Load Tests Miss the Failures That Actually Happen
The specific ways a synthetic load test differs from a real traffic spike, and what to build into the test so it catches what actually breaks.
Building Synthetic Probes That Catch an Outage Before Customers Do
How to design synthetic transaction probes that actually catch real failures, instead of monitoring theater that stays green while customers see errors.
Your Uptime Monitor Looks Fine. Your Customers Disagree
A checklist for building synthetic monitoring that catches what a basic uptime check misses, and the common mistakes that leave it blind to real outages.
A Checklist for Synthetic Monitoring That Actually Catches Outages Early
A practical checklist for setting up synthetic transaction probes that catch real customer-facing failures, plus the common pitfalls that make them useless.
What Synthetic Monitoring Catches That Your Alerts Don't
How synthetic transaction probes catch outages that server metrics and error-rate alerts miss, and how to set them up without drowning in false alarms.