Developer Productivity & Platform EngineeringPlaybook3 min readUpdated September 2026

Synthetic Monitoring: Testing the Paths Users Take

An uptime dashboard reads 99.98 percent up for the month while support tickets say checkout has been broken for six hours, because the health check only pings a root path that returns 200 regardless of whether checkout actually works.

Synthetic monitoring closes that gap by simulating the flows customers actually take, not just checking that a server responds. Getting value from it means picking the right handful of flows and alerting on them without crying wolf.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Why a Green Health Check Can Hide a Broken Product

A shallow health check verifies that a process is running and can answer a request. It says nothing about whether the database connection that checkout depends on is healthy, whether a third-party payment call is timing out, or whether a recent deploy broke the coupon code path while leaving the root route untouched.

A synthetic probe instead walks through an actual sequence: load the page, add an item to the cart, apply a coupon, submit a test payment. If any step in that sequence fails, the probe fails, even though the server's root path would still return a perfectly healthy 200.

Picking the Handful of Flows Worth Simulating

You don't need to simulate every page. Pick the flows where a failure directly costs revenue or trust: signup, checkout, password reset, and whatever single action is the core of your product. Running probes against every minor page adds cost and noise without adding much protection, and a flood of low-value alerts is how a team stops trusting the whole monitoring setup.

For example, a subscription product might list five candidate flows: signup, checkout, password reset, a settings page, and a help center article. Only the first three touch revenue or account access, so a sensible first version probes those and the core action of the product, and leaves the rest to real user monitoring. A common mistake is adding a new probe after every incident until nobody remembers why each one exists. The fix is a short review: for each probe, write down what it costs you when that flow fails, and retire any probe whose honest answer is not much.

Where Synthetic Probes Overlap With Real User Monitoring, and Where They Don't

Synthetic probes run from fixed locations on a schedule, which means they catch an outage before a real customer hits it, often within minutes. What they can miss is anything specific to a particular browser, region, or account state that only shows up in genuine traffic, an edge case a scripted probe never exercises.

Real user monitoring catches those edge cases but only after a real customer has already been affected. Running both together, rather than picking one, covers each other's blind spot.

Setting Alert Thresholds That Don't Cry Wolf

  • Alert on two or three consecutive failures, not a single blip, since a probe run from one location can hit a transient network hiccup that has nothing to do with your service.
  • Treat a slow response and a broken response as separate alerts, since the fix and the urgency for each are usually different.
  • Page a human only for flows tied directly to revenue or account access; route everything else to a ticket queue instead.

The Cost You're Actually Trading Off

Running synthetic probes too frequently, across too many regions, without ever pruning the list, quietly adds up in monitoring cost for very little extra protection. Weigh that against the cost of an undetected outage: say a checkout outage costs your business two thousand dollars an hour in abandoned carts. A probe that catches that outage in five minutes instead of six hours can pay for a full year of monitoring in a single incident.

What to Do With the Data Once You Have It

A probe that just fires an alert and gets forgotten only helps in the moment of the outage. Keep a record of every failure and its cause, and you start to see patterns: the same third-party payment step timing out every Monday morning, or checkout failures clustering right after a specific kind of deploy.

That history also makes incident review faster. Instead of reconstructing what happened from logs after the fact, you already have a timestamped record of exactly when the flow started failing and which step broke, which turns a vague postmortem into a specific one with a real fix attached.

Review the failure log on a regular cadence, not just after an incident. A flow that's been silently flaky for weeks without ever crossing the alert threshold is exactly the kind of problem a periodic review catches before it becomes the outage that does cross it.

Executive Capability Standard

What Good Looks Like

Good synthetic monitoring means the flows that make you money, signup, checkout, the core action of the product, get caught within minutes of breaking, while low-value pages don't generate alert noise nobody reads.

Building The Capability (5-Stage Skill Ladder)

1. Learn:List the handful of user flows where a failure would directly cost revenue or trust, and check whether any of them are currently monitored at all.
2. Do Manually:Manually walk through your top two flows on a schedule for a week to see how often they actually break, before investing in automated probes.
3. Delegate:Assign one engineer to own the probe list and alert thresholds so tuning doesn't happen ad hoc during an incident.
4. Automate:Automate synthetic checks for your revenue-critical flows on a schedule short enough to catch an outage in minutes, not hours.
5. Buy:Bring in a reliability specialist to set up region coverage and alert routing if outages are currently found by customers before your team.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

ClickUp

A task tool like ClickUp can route a confirmed probe failure straight into a ticket for the owning team, separate from the pages that go to an on-call human.

Visit ClickUp→

Frequently Asked Questions

How is synthetic monitoring different from basic uptime monitoring?

Uptime monitoring checks that a server responds, usually to a single lightweight endpoint. Synthetic monitoring simulates an actual user flow, like adding an item to a cart and checking out, so it can catch a broken feature even when the server itself is technically up and responding to every request.

How many flows should I actually be monitoring this way?

Start with the handful where a failure directly costs revenue or account access: signup, checkout, password reset, and your product's core action. Adding more flows than that mostly adds noise and cost without meaningfully improving coverage, since most other pages fail loudly enough to show up in error logs anyway.

How fast should a synthetic probe alert fire after a failure?

Wait for two or three consecutive failures before paging anyone, since a single failed run can just be a transient network blip from the probe's own location. Once you've confirmed a real failure, page immediately for revenue-critical flows and route lower-priority flows to a ticket instead.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides