What Your Alerts Are Actually Telling You
An observability stack can generate thousands of data points a minute and still fail the one test that matters: when something breaks, does the right person find out fast enough to fix it before a customer notices. Most teams do not have too little monitoring, they have monitoring nobody trusts, because half the alerts turned out to be noise the first month and everyone stopped reading them.
This is a short audit for checking what your current setup actually catches, what it misses, and where to spend the next hour of engineering time.
Start by counting alerts nobody acted on last month
Pull your alert history and sort it into three piles: alerts someone acted on, alerts someone acknowledged and ignored, and alerts nobody saw at all. If the second and third piles outnumber the first, the problem is not coverage, it is signal quality, and adding more monitoring will only make it worse.
Delete or rework any alert that fired more than a handful of times last month without leading to a real fix. A threshold set too sensitively trains your team to ignore your monitoring, which is more dangerous than not having it. A single alert that pages someone overnight for a condition that resolves itself within a couple of minutes does more damage to trust in your monitoring than a genuine outage does, because it teaches the on-call engineer that the phone buzzing does not necessarily mean something is actually wrong.
Audit last month's alerts with these steps:
- Pull your alert history and sort every alert into acted on, acknowledged and ignored, or never seen.
- Compare the piles: if ignored and unseen alerts outnumber acted-on ones, treat it as a signal quality problem, not a coverage gap.
- Delete or rework any alert that fired more than a handful of times without leading to a real fix.
- Re-check the alerts that page someone overnight, and remove any whose condition resolves itself within a couple of minutes.
Check whether you can answer 'is it us or them' in under five minutes
When a customer reports something slow or broken, the first question is always whether the problem is inside your system or in a dependency you do not control. If answering that takes longer than checking a dashboard and a trace, your telemetry has a gap, usually around third-party API calls or database query performance, that is worth closing before adding any new alert category.
A comparison of observability platforms is a reasonable place to start if your current tooling cannot answer that question cleanly, since the gap is usually a tooling limitation, not a process one.
Trace error budgets back to something you can actually see
An availability target only means something if you can measure against it. A 99.95% target allows a little over four hours of downtime a year1, and that number is only useful if your monitoring can tell you, at any moment, how much of that budget you have already spent.
If nobody can answer that question without doing math by hand, your dashboards are showing raw uptime instead of budget burn, and burn is the number that actually drives a decision about whether to slow down releases.
The gap that shows up most often: the deploy that broke it
The fastest way to diagnose most incidents is correlating a metric spike against a deploy timeline, and a surprising number of setups do not make that overlay easy to see. If your team spends the first ten minutes of every incident manually checking what shipped recently, that overlay is worth building before anything else on this list.
Once it exists, most regressions get caught within minutes of deploy instead of hours after a customer notices, simply because the correlation is visible instead of something someone has to remember to check.
What to add once the basics are trustworthy
Once error rate, latency, and deploy correlation are solid, the next useful addition is usually an external synthetic check that pings your product the same way a customer would, from outside your own infrastructure. That catches the failure mode internal monitoring tends to miss entirely: every internal system reporting healthy while a DNS misconfiguration or an expired certificate keeps real users locked out.
Route alerts by actual severity instead of sending everything to the same channel. A payment failure and a slow overnight batch job deserve different response times, and treating them the same is exactly what trains a team to start ignoring alerts in the first place.
For example, suppose every internal dashboard shows healthy services while customers cannot load the login page because a certificate expired. An external synthetic check that requests the login page from outside your infrastructure would have caught it. Route that check straight to the on-call engineer, and send a slow overnight batch job to a channel that is reviewed the next morning. The common mistake is adding the synthetic check but sending it to the same noisy channel as everything else, where it gets ignored just like the earlier alerts. Give each new check an owner and a severity when you create it.
What Good Looks Like
Good here means anyone on the team can answer whether an incident is inside your system or a dependency's within five minutes, using a dashboard, not by asking around in chat.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How many alerts is too many?
There is no fixed number, but if your on-call engineer cannot recall what most of last week's alerts were about, you have too many low-value ones drowning out the ones that matter. Fewer, trusted alerts beat comprehensive coverage nobody reads.
Should every service have its own dashboard?
Only if someone actually looks at it. A dashboard built once and never opened again is worse than no dashboard, since it creates a false sense that the service is being watched. Build dashboards around the questions someone asks during an incident, not one per service by default.
What should we monitor first if we are starting from nothing?
Start with error rate, latency, and traffic on your most customer-facing endpoints, plus deploy markers overlaid on all three. That combination answers most incident questions on its own and gives you a foundation to add depth to later, rather than trying to instrument everything at once.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
Datadog vs New Relic vs Dynatrace: Cloud Observability Platforms Compared
Compare Datadog, New Relic, and Dynatrace for cloud observability: log ingestion costs, distributed tracing, APM overhead, and MTTR compression.
What to Actually Monitor Before You Buy an Observability Tool
Answers to the questions engineering teams actually ask before setting up monitoring: what to track, how many alerts is too many, and when to add tracing.
Where Production Deployment Budgets Quietly Leak
The recurring places engineering teams overspend on production deployment architecture, and a practical order for fixing them without a full rebuild.
What "Zero Trust" Actually Means for Device Verification
Zero trust device verification means a device is trusted continuously, based on its current state, not once at login. Here is what that actually requires.
A 30-Minute Check for Blind Spots in Your Observability Setup
A quick, practical checklist CTOs can run in 30 minutes to find the gaps in telemetry and alerting that usually surface during an incident instead.
How to Benchmark Your System Before It Has to Scale
A practical runbook for benchmarking throughput and capacity before you actually need the headroom, so scaling decisions are based on data, not guesses.