Building Synthetic Monitoring That Catches Real Outages
A homepage check that returns a 200 tells you almost nothing about whether a customer can actually log in, add something to a cart, or complete checkout. Most outages that matter happen one layer deeper than a simple health check ever looks.
Synthetic monitoring closes that gap by scripting the actual steps a user takes and running them on a schedule from outside your network, the same way a real customer would hit the service.
Test the user journey, not the homepage
Pick the two or three flows that would be a genuine incident if they broke: signing in, completing a purchase, sending the first message in a new conversation. Script each one as a sequence of real requests or browser actions, not a single ping to the root URL.
A synthetic probe that fires every minute from outside your network is what actually lets you measure whether you're living inside your downtime budget or already over it1, because it sees the same failures a customer would, including ones your internal health checks never touch.
Where to run probes from, and how many
Run probes from at least two geographic regions, ideally ones close to where your actual customers are, so a regional network problem doesn't get mistaken for your service being down everywhere. Running from a single location is a common gap: it catches your outages but misses a CDN or DNS provider's regional problem entirely.
Frequency matters too. A check every five minutes might mean a real outage runs for five minutes before anyone's paged. For a flow that generates revenue every minute it's broken, a tighter interval is worth the extra probe traffic. There's a cost side to this as well: every probe run is itself a request against your production system, so a very tight interval across many regions and many flows can add up to a meaningful, if small, slice of your own traffic.
Alert thresholds that don't page you for nothing
A single failed probe run is often a blip: a transient network hiccup between the probe location and your service, not a real outage. Alerting on the first failure trains your on-call rotation to distrust the alert. Requiring two or three consecutive failures, or failures from more than one region at once, before paging someone cuts noise without meaningfully slowing your detection of a real incident.
Pitfalls that make synthetic monitoring lie to you
A probe that logs in with the same test account every run can mask a real problem if that account happens to have cached or special-cased data that regular customers don't. A probe that never gets updated when the actual user flow changes ends up testing a page that no longer matters, while the page customers actually use goes unchecked.
A less obvious pitfall: a probe account that never expires, never changes its password, and never gets rate limited can quietly bypass the exact protections you'd want tested. If your real customers hit a captcha or a fraud check under some conditions, make sure your probe account isn't permanently exempt from the flow it's supposed to be verifying.
Treat the probe scripts as part of the product, reviewed and updated whenever the flow they test changes, not as a one-time setup task you can forget about once it's green.
What to check before trusting a green dashboard
Confirm the probe is actually exercising the full flow, including the steps most likely to fail, like payment processing or a third-party integration, rather than stopping short at an easy checkpoint. And run a deliberate failure test now and then, breaking the flow on purpose in a lower environment, to confirm the alert actually fires the way you expect before you need it to.
It's worth reviewing who actually receives the alert, too. A page that routes to a Slack channel nobody watches on weekends is functionally the same as no alert at all, and that gap only shows up during an incident, which is the worst possible time to discover it.
Before you trust a green dashboard, confirm the following:
- The probe exercises the whole user journey, including the steps most likely to fail, such as payment processing or a third-party integration.
- Probes run from at least two regions so a regional network or DNS problem is not mistaken for a global outage.
- The test account does not have cached or special-cased data that regular customers lack.
- The probe script has been updated to match the current user flow, not a page that no longer matters.
- A deliberate failure in a lower environment has shown that the alert fires the way you expect.
What Good Looks Like
Good synthetic monitoring means the scripted flows match what customers actually do today, and a deliberate failure test confirms the alert path really works.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How is synthetic monitoring different from uptime monitoring?
Basic uptime monitoring usually checks that a single endpoint responds. Synthetic monitoring scripts a full user journey, like logging in and completing a purchase, so it catches failures deeper in the stack that a simple health check would never see.
How many consecutive failures should trigger a page?
Two or three consecutive failures, ideally confirmed from more than one probe location, is a reasonable default. A single failed run is often a transient network blip rather than a real outage, and paging on every blip trains the team to ignore the alert.
Do we need probes in multiple regions?
Yes, if your customers are spread across regions. A single probe location can miss a regional network or DNS problem that's genuinely affecting a slice of your customers while looking completely fine to a probe sitting somewhere else.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
How to Load-Test a Model Endpoint Without Faking the Results
How to load-test and stress-test an AI model-serving endpoint with realistic traffic, and what to watch beyond pass or fail.
Why Automated SLA Alerts on Inference Break at Scale
Why latency and uptime alerts on a model serving endpoint stop working as traffic grows, and how to set thresholds and route the alerts that matter.
A Rollout Checklist for Swapping Models in Production
A rollout checklist for swapping AI models in production: evaluation gates, canary and shadow traffic, and fast rollback paths.
Why Inference Latency Creeps Up After You Ship
Where AI model-serving latency actually hides: tokenization, queueing, batching windows, and network hops, plus a worked example fix.
What Synthetic Monitoring Catches That Real Traffic Misses
A checklist for setting up synthetic transaction probes that catch real failures early, plus the common pitfalls that make teams stop trusting them.
Zero Trust for Machines Calling Your Model Endpoints
Why internal services calling your inference endpoints still need identity checks, and where CrowdStrike-style posture checks and Tenable-style scanning fit.