Uptime Monitoring Checklist: What to Watch and How to Alert
A good uptime monitoring checklist covers the public pages and APIs, the critical user flows, the certificate and DNS layers, the background jobs and the third-party services you depend on. Each check should run from outside your network and alert a person who can act.
The goal isn't a wall of green dots. It's hearing about an outage before your customers tell you, and not being paged for false alarms. Work through the sections below and tick each one.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
What should you monitor?
Cover each layer a customer's request passes through:
- Public website and marketing pages: a simple availability check on the home page.
- Application and login: the app loads and a real sign-in works.
- API health: a lightweight endpoint that checks its own dependencies, such as the database and cache, and reports them honestly.
- Critical flows: signup, checkout or whatever earns your revenue, tested end to end with a synthetic script.
- DNS and TLS: resolution works and certificates aren't about to expire. Certificate expiry is a preventable outage that still catches teams.
- Background jobs and scheduled tasks: heartbeat checks that alert when an expected ping doesn't arrive.
- Third-party dependencies: payment, email and authentication providers that your app cannot work without.
If you can only build three things this week, choose the login flow, the API health check and certificate expiry.
How should you design each check?
Cheap checks give false confidence, so design them to fail when customers would notice:
- Verify content, not just a 200 status. An error page or a login wall can return success.
- Run from more than one region, and require two or more locations to fail before alerting, to filter out a bad network path.
- Confirm a failure with a quick retry before paging.
- Set a timeout that matches user tolerance, so a check that takes 30 seconds counts as down.
- Don't hit expensive or state-changing endpoints often. Use a dedicated health route and a test account.
- Choose an interval that fits the impact: a minute for the checkout, five minutes for the marketing site.
A short interval multiplied by many locations and many checks can turn into a large monitoring bill, so match tightness to importance.
Who gets alerted, and how?
An alert is only useful if it reaches someone who can fix the problem. Decide these things in advance:
- A named on-call person, with a backup, and a clear escalation time if the first person doesn't acknowledge.
- Channels: phone or push for outages, chat for warnings, email for reports. Don't send everything to one noisy channel.
- Severity: page for customer-facing outages, and ticket for slow degradation that can wait until morning.
- Runbook links inside the alert so the first responder knows where to look.
Track how many alerts are false. If people start muting the channel, fix or delete the noisy monitors immediately, because an ignored alert system is worse than none.
How much downtime does a target actually allow?
It helps to know what a number means before you promise it. An availability target of 99% allows about 3.65 days of downtime per year, while 99.9% allows about 8.76 hours1.
Say you run a single database instance with no failover. A four-nines promise would be fragile, because one restart could use much of the year's allowance. Match your target to your architecture, and write down what you promise customers only after you've measured what you deliver.
What should you do about communication and upkeep?
When something breaks, customers want an acknowledgment and an estimate. Keep a status page, hosted separately from your own infrastructure so it stays up during an outage, and publish incidents promptly with plain updates. Tell customers what you know, what you're doing and when you'll update next.
Maintain the system itself:
- Test your monitors monthly by intentionally breaking a non-production check or pausing a heartbeat.
- Add a monitor for every incident that customers reported first.
- Silence checks during planned maintenance, and remove the silence afterward.
- Review the list quarterly and delete monitors for retired services.
A tool such as Datadog can run synthetic checks alongside your metrics, and the platform comparison covers other choices. The production deployment checklist covers what to watch right after releases.
What Good Looks Like
Critical flows, APIs, certificates, DNS and scheduled jobs are checked from outside your network, with alerts that reach a named responder and a status page for customers.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
How often should uptime checks run?
Every minute or so for revenue-critical flows and every few minutes for lower-risk pages. Shorter intervals find outages sooner but cost more and can trigger false alarms. Match the interval to how quickly you'd need to respond.
Why check from multiple locations?
A single location can fail because of its own network problem. Requiring failures from two or more regions before alerting removes most false alarms and shows whether an outage is regional.
Is a 200 status code enough for a health check?
No. A broken app can return 200 with an error page. Check for expected content, and have the health endpoint verify critical dependencies such as the database, without making it expensive to call.
Do we need a status page?
For any product with paying customers, yes. Host it away from your own infrastructure so it works during outages, and post updates early. It reduces support load and builds trust.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
Datadog vs New Relic vs Dynatrace: Cloud Observability Platforms Compared
Compare Datadog, New Relic, and Dynatrace for cloud observability: log ingestion costs, distributed tracing, APM overhead, and MTTR compression.
A Production Deployment Checklist: Before, During and After
A deployment checklist covering pre-release checks, safe database migrations, feature flags, post-deploy verification and rollback.
Your Uptime Monitor Looks Fine. Your Customers Disagree
A checklist for building synthetic monitoring that catches what a basic uptime check misses, and the common mistakes that leave it blind to real outages.
New Developer Onboarding: A Checklist for the First 30 Days
A phased checklist for onboarding a new developer: access before day one, a first merged change in week one, and how to measure it worked.
A Checklist for Synthetic Monitoring That Actually Catches Outages Early
A practical checklist for setting up synthetic transaction probes that catch real customer-facing failures, plus the common pitfalls that make them useless.
Synthetic Monitoring: Testing the Paths Users Take
A green uptime dashboard can hide a broken checkout for hours. How to pick the handful of flows worth simulating and alert on them well.