What Synthetic Monitoring Catches That Real Traffic Misses
Real user monitoring tells you something broke after a user already hit it. Synthetic monitoring, a scripted transaction that runs on a schedule from outside your infrastructure, login, checkout, a critical API call, tells you before that, as long as it's set up to actually mimic what a user does instead of just pinging a health check endpoint.
The gap between having synthetic monitoring and having synthetic monitoring that actually catches things comes down to which transactions you probe, how you set thresholds, and where the probes run from. This works through the setup decisions in order, plus where teams most often end up with a wall of alerts nobody reads.
What Synthetic Probes Are For
A health check endpoint tells you the process is running. A synthetic probe tells you the thing a user actually cares about still works: can someone log in, can a payment complete, does a search return results. Those are different failure classes. A service can return a healthy status on its own health check while the login flow it depends on is broken three services downstream, and only a synthetic probe walking the actual flow catches that gap.
Don't replace real user monitoring with synthetic probes; they answer different questions. Real user monitoring tells you what's actually happening to your real traffic right now, including edge cases you'd never think to script. Synthetic probes tell you whether the core flows work at all, on a predictable schedule, even during low-traffic hours when real user monitoring might not catch a break for a while.
Picking Which Transactions to Probe
Probe the handful of flows where a failure costs real money or trust in the first few minutes, not every possible user path. Login, checkout or the core conversion action, and any API call a paying customer depends on directly are the usual short list. Adding probes for every secondary page mostly adds noise and alert fatigue without adding much protection, since a failure there is rarely as urgent.
For each probe, script the actual multi-step flow, not just a single request. A login probe that only checks the login page loads misses the case where the page loads fine but the authentication service behind it is down. Walk the whole flow the way a real user would, including the step that confirms it actually worked.
Setting Thresholds That Don't Cry Wolf
A probe that pages someone on a single failed run trains the team to ignore pages, since transient network blips happen even when nothing's actually wrong. Require two or three consecutive failures, or failures from more than one probe location at once, before it escalates to a page. Reserve the single-failure alert for a lower-urgency channel someone checks during business hours.
Set latency thresholds from your own historical data, not a round number that sounds reasonable. Pull a week of real probe run times, and set the warning threshold above your normal variance rather than an arbitrary guess, so a threshold set too tight doesn't start paging on ordinary variance.
Where Probes Should Run From
Run probes from outside your own infrastructure and ideally from more than one geographic region, since a probe running inside your own network can stay green while an external DNS or CDN issue makes the service unreachable to real users everywhere else. This is the single most common gap in a synthetic monitoring setup: probes that only prove the service works from a vantage point no real user actually has.
If your users are concentrated in a specific region, weight your probe locations toward where your traffic actually comes from rather than spreading evenly across every available region. A tool comparison like Datadog vs. New Relic vs. Dynatrace is worth reading before committing to one platform, since probe location coverage varies meaningfully between vendors.
Common Pitfalls That Kill Trust in Probes
Watch for these, since each one is a reason teams end up muting the alerts instead of fixing the underlying flow:
- Probing a login flow with a real test account that occasionally gets rate-limited or flagged by the same fraud detection meant to protect real users, causing false failures.
- Letting probe scripts go stale after a UI change, so the probe fails not because the flow broke but because it can't find a button that moved.
- Alerting the whole team on every probe failure instead of routing to whoever actually owns that flow, which trains everyone to assume someone else will look at it.
- Never reviewing probe coverage after launch, so new critical flows ship with no synthetic coverage at all while old, now-irrelevant flows keep getting probed.
Review probe coverage on the same cadence you review incident postmortems. A probe that's been silently failing or silently useless for months is a gap you won't find any other way.
What Good Looks Like
Good synthetic monitoring means the handful of flows that matter most, login, checkout, core paid API calls, are probed as full multi-step transactions from outside your network, with thresholds tuned to your own historical variance.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How is synthetic monitoring different from real user monitoring?
Real user monitoring shows what's actually happening to real traffic right now, including cases you'd never think to script. Synthetic monitoring runs a scripted transaction on a fixed schedule from outside your infrastructure, so it catches core flows breaking even during low-traffic hours when real user monitoring might not notice for a while.
How many consecutive failures should trigger a page?
Two or three consecutive failures, or failures from more than one probe location at once, is a common starting point. A single failure is usually a transient network blip rather than a real outage, and paging on every single one trains the team to ignore the pages entirely.
Why do probes need to run from outside our own network?
A probe running inside your infrastructure can stay green during a DNS, CDN, or edge routing failure that makes the service unreachable to every real user outside your network. Running probes externally, ideally from more than one region, is what actually proves the service works the way a real user experiences it.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Datadog vs New Relic vs Dynatrace: Cloud Observability Platforms Compared
Compare Datadog, New Relic, and Dynatrace for cloud observability: log ingestion costs, distributed tracing, APM overhead, and MTTR compression.
Load Testing an Agent System Before It Meets Real Traffic
Answers to the practical questions CTOs have about load testing agentic systems, from what to simulate to how much traffic is actually enough.
Synthetic Monitoring: Testing the Paths Users Take
A green uptime dashboard can hide a broken checkout for hours. How to pick the handful of flows worth simulating and alert on them well.
Why Automated SLA Alerts Keep Missing Real Breaches
Why single-threshold SLA alerts miss real breaches, and how matching the contract's window, error budget and failure modes catches them before customers do.
Rolling Out Agentic Workflows Without Breaking Production
A practical rollout checklist for shipping an AI agent to production, from a shadow-mode test run through the guardrails that catch it if it misbehaves.
Building Synthetic Monitoring That Catches Real Outages
How to set up synthetic monitoring that tests the journeys customers actually take, without drowning your on-call rotation in false alarms.