Engineering Leadership & Technical HiringPlaybook3 min readUpdated September 2026

Synthetic Monitoring That Watches What Customers Actually Do

A health-check endpoint returning 200 tells you the process is up. It doesn't tell you whether a customer can actually complete checkout, log in, or upload a file, because those flows touch a database, a third-party payment processor, and session state that a bare health check never exercises.

This is a checklist for building synthetic probes around what actually matters, and the pitfalls that turn a good monitoring setup into a source of ignored alerts.

Probe the Journeys That Actually Make Money

List your top three or four revenue-critical paths before writing a single probe script: signup, checkout, the core action your product exists to perform. A probe that logs in and clicks around a dashboard is easy to write and tells you almost nothing compared to one that completes an actual checkout flow, including the third-party payment step most teams are tempted to stub out of the probe because it's harder to automate. The harder-to-build probe is usually the more valuable one.

Frequency and Geographic Spread Should Match Your User Base

A probe running every five minutes from a single region catches less than one running every minute from the regions your actual traffic comes from, especially for latency regressions that only show up from certain geographies due to CDN routing or a third-party dependency with regional outages. Match probe frequency to how quickly you need to know something broke, a checkout flow probably warrants tighter intervals than an internal admin page, and match geographic spread to where your customers actually are, not where your infrastructure happens to sit.

Alert Fatigue From Flaky Probes Kills the Whole Program

A probe that occasionally times out because of its own network path, not a real outage, and pages someone every time, trains the on-call rotation to ignore synthetic alerts within a month. Require consecutive failures from more than one probe location before paging, and separate "this probe is degraded" from "the service is down" as distinct alert severities. A synthetic monitoring program that's been muted by its own team has negative value: it looks like coverage exists when it doesn't.

Synthetic Monitoring and Real User Monitoring Answer Different Questions

Synthetic probes tell you a specific, scripted path works from a specific location on a schedule; real user monitoring tells you what's actually happening across every real session, every browser, every device, all the time. Synthetic catches an outage before a human notices; RUM tells you which real-world edge case, a specific browser version, a slow connection, is actually hurting users right now. Run both. Synthetic without RUM misses the long tail of real conditions; RUM without synthetic means you only find out about an outage from angry customers instead of a proactive alert.

Keeping Probes From Rotting as the UI Changes

A probe built against specific selectors or a specific flow breaks the moment product ships a redesign, and if nobody owns probe maintenance, it either starts failing constantly on false positives or gets quietly disabled and never re-enabled. Give one engineer explicit ownership of probe health as part of the release process for the flows they cover, and treat a probe failure after a deploy as a signal to check first, not dismiss, since it's just as likely to be a real regression as a stale selector.

A solid synthetic monitoring setup covers these points:

  • Three or four revenue-critical journeys, such as signup, checkout, and your product's core action, instead of a bare health endpoint.
  • Probe frequency and locations that match where your real traffic comes from.
  • Alerts that fire only after consecutive failures from more than one probe location.
  • One named engineer who owns probe maintenance as the UI changes.
  • Real user monitoring alongside probes, since a probe only exercises the exact path its script follows.

What Probes Miss That You Still Need to Cover

A synthetic check exercises the exact path its script follows and nothing else, so it won't catch a bug that only appears for a specific account state, a specific plan tier, or a specific combination of feature flags outside its scripted path. Treat probes as coverage for your most critical, common paths, not a substitute for broader end-to-end test coverage or the kind of exploratory testing that catches edge cases a fixed script never will. A team that believes synthetic monitoring means the product is fully covered is usually the one surprised by a bug that lived in an untested corner for months.

Choosing Probe Locations to Match Real Failure Patterns

A third-party dependency, a CDN, a DNS provider, an upstream API, can fail regionally without your own infrastructure being at fault at all, and a probe running from only one location can't distinguish "our service is down" from "this specific region can't reach us." Running probes from a handful of geographically distinct locations, ideally overlapping with where your actual customer base sits, turns an ambiguous alert into an actionable one: a failure isolated to one probe location points at network or regional issues, while a failure across every location points squarely at your own service.

Executive Capability Standard

What Good Looks Like

Synthetic checks should exercise the same critical path a paying customer takes, not just a health endpoint that always returns 200.

Building The Capability (5-Stage Skill Ladder)

1. Learn:List your top three revenue-critical user journeys before writing a single probe.
2. Do Manually:Run through the checkout or signup flow yourself from a few regions before automating it as a probe.
3. Delegate:Give one engineer explicit ownership of probe maintenance so scripts don't rot when the UI changes.
4. Automate:Schedule probes from multiple regions and alert on consecutive failures, not a single blip.
5. Buy:Consider a managed synthetic monitoring platform once you're maintaining probe infrastructure across more regions than makes sense in-house.

How to Get Started

Frequently Asked Questions

How many synthetic probes is enough for a small engineering team?

Three to five, covering your top revenue-critical journeys, is a reasonable starting point. More than that with a small team usually means maintenance burden outpaces the signal you get, especially once the UI starts changing faster than probes get updated.

Should synthetic probes run against production or a staging environment?

Production, for the ones meant to catch real customer-facing outages. Staging probes are useful for pre-deploy smoke testing, but they don't tell you whether the actual production environment, with its real third-party integrations and real data volume, is healthy right now.

What's a reasonable threshold before a probe failure pages someone?

Two to three consecutive failures from at least two different probe locations, rather than a single failed check. This filters out transient network blips in the probe's own path while still catching a genuine outage within a few minutes.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides