Your Uptime Monitor Looks Fine. Your Customers Disagree
A simple uptime check that pings the homepage every minute will stay green while checkout is silently broken, a specific payment method fails every time, or search returns empty results for every query. Synthetic monitoring closes that gap by scripting the actual user journeys that make the product work, not just confirming the server responds.
This is a checklist for building synthetic probes that catch what a homepage ping misses, along with the mistakes that quietly leave a monitoring setup blind to a real outage.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Why does a homepage ping miss real outages?
A basic uptime check confirms the server returns a 200 status code, which tells you almost nothing about whether the product actually works. The database can be timing out on every write, a third-party payment processor can be down, or a recent deploy can have broken the checkout flow specifically, and a homepage ping stays green through all of it. Synthetic monitoring means scripting an actual transaction, logging in, adding an item, completing checkout, and confirming the result at every step, not just the final status code.
The probes worth building first
- Authentication: a scripted login with a dedicated test account, run from outside your infrastructure, catches an auth provider outage or a broken session cookie config.
- The core transaction: whatever action makes your product money, a checkout, a booking, a form submission, scripted end to end with assertions on the actual result, not just a 200 response.
- Every payment method separately: a payment gateway can fail for one card network or region while others work fine, so checking only one method leaves the others unmonitored.
- Search or the primary data query: confirming a known query returns a known result catches an index that silently stopped updating.
Start with whichever of these, if it broke silently for four hours, would cost the most revenue or trust, and build outward from there.
Running probes from outside your own infrastructure
A probe that runs from inside your own cloud region will miss a DNS issue, a CDN misconfiguration, or a regional network problem that only affects customers in a specific geography. Run synthetic checks from multiple external locations, ideally matching where your actual customer base is concentrated, so a probe from a region with few customers doesn't mask a real outage affecting the region with most of them. This also catches the class of incident where your infrastructure looks perfectly healthy from the inside while being unreachable from the outside.
How do you stop synthetic alerts from turning into noise?
A single failed check from one location is often a transient network blip, not a real outage, and paging someone for every single failure trains the team to ignore the alert within a month. Require two or three consecutive failures, or failures from multiple independent locations, before paging, and route a single-location failure to a lower-priority channel for investigation rather than an immediate page. The goal is a monitor the team trusts enough to act on immediately, not one that cries wolf so often it gets muted.
Tying probes to your error budget, not just uptime
Synthetic probe results are most useful when tied to the same availability target the rest of the organization already tracks. If the team has committed to three-nines availability, the allowed downtime budget is well under half a day a year1, and every synthetic probe failure that crosses the alert threshold should count against that budget in the same dashboard leadership already looks at. This turns synthetic monitoring from an engineering-only tool into something that connects directly to a commitment the business has made to customers, especially once it sits in the same observability platform the rest of the team already checks daily.
A worked example: the payment method nobody was watching
Say a product accepts three card networks, and a routing change at the payment processor quietly breaks authorization for just one of them on a Friday afternoon. A homepage ping stays green the entire weekend, since the server is healthy and two of the three card networks still work fine. Without a probe scripted against each payment method separately, the first signal the team gets is a spike in support tickets from customers whose cards keep failing, sometimes days later, well after the trust damage is done. A synthetic probe that checks all three networks on a short interval would have caught this within minutes of the routing change, not days after it.
What Good Looks Like
Mature synthetic monitoring scripts the actual revenue-generating transaction end to end from multiple external locations, requires consecutive failures before paging, and ties results to the same availability budget the business has committed to.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
How often should a synthetic probe run?
For a core transaction like checkout, every one to five minutes is typical, since that determines how quickly you detect an outage. Less critical journeys can run every 15 to 30 minutes. Balance detection speed against the cost of running the probe at scale.
Do we need synthetic monitoring if we already have detailed application logging?
Logging tells you what happened after a customer hits a problem, or requires you to notice it in the logs. Synthetic monitoring proactively exercises the same path a customer would take, often catching a failure before any customer reports it or generates an error log entry at all.
Should synthetic probes use real customer accounts or dedicated test accounts?
Dedicated test accounts, isolated from real customer data and billing, run on a schedule you control and won't be affected by rate limits, fraud checks, or data changes a real account might trigger. Never point automated probes at a production customer account.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
Datadog vs New Relic vs Dynatrace: Cloud Observability Platforms Compared
Compare Datadog, New Relic, and Dynatrace for cloud observability: log ingestion costs, distributed tracing, APM overhead, and MTTR compression.
Synthetic Monitoring: Testing the Paths Users Take
A green uptime dashboard can hide a broken checkout for hours. How to pick the handful of flows worth simulating and alert on them well.
A Checklist for Synthetic Monitoring That Actually Catches Outages Early
A practical checklist for setting up synthetic transaction probes that catch real customer-facing failures, plus the common pitfalls that make them useless.
Synthetic Monitoring That Watches What Customers Actually Do
A checklist for synthetic transaction monitoring: which journeys to probe, how to avoid alert fatigue, and where synthetic checks miss what real users hit.
What Synthetic Monitoring Catches That Your Alerts Don't
How synthetic transaction probes catch outages that server metrics and error-rate alerts miss, and how to set them up without drowning in false alarms.
Synthetic Monitoring: Catching Outages Before Customers Do
How to design synthetic transaction probes that catch a real outage instead of false alarms, and where they can't replace real user monitoring.