A Practical Checklist for Observability That Gets Used
Most teams don't have too little observability. They have a pile of dashboards and alerts that no one trusts enough to act on during an actual incident, which is functionally the same as having none.
Here's a checklist for building the kind of observability that gets used, not just installed.
Instrument for the questions you'll actually ask during an incident
Before adding another dashboard, list the two or three questions you always ask when something breaks: is this affecting all users or a subset, is it a specific endpoint or everything, did a deploy just go out. Build your instrumentation to answer those questions directly, rather than collecting every metric a tool offers by default.
A dashboard built around a real incident question gets opened during the next incident. A generic one, built because a tool suggested it, usually doesn't.
Set alert thresholds off your own baseline, not a generic default
A default threshold, copied from documentation or a template, either fires constantly on normal variation or stays silent through a real problem, because it was never set against your actual traffic. Pull a few weeks of your own data first, and set thresholds relative to that baseline.
This matters most for anything with a strong daily or weekly pattern. An alert tuned for your quiet Tuesday morning will misfire every single Monday if you didn't account for the cycle.
The three signals worth paying for: logs, metrics, and traces
Logs tell you what happened in detail but are slow to search across a whole system. Metrics tell you that something changed, quickly, but not why. Traces connect a single request across every service it touched, which is what actually lets you find where in a distributed system something went wrong.
Most teams over-invest in logs early because they're the easiest to add, then discover during a real incident that tracing a request across services is what they actually needed and never built.
Pitfalls: dashboards nobody opens and alerts nobody trusts
Watch for these patterns, since each one quietly erodes trust in the whole system:
- An alert that's fired incorrectly enough times that people mute it instead of investigating
- A dashboard built once and never updated as the system it monitors changed
- Alerting on a symptom (high CPU) instead of user impact (failed requests, slow responses)
- No clear owner for a given alert, so it pages a rotation that doesn't know what to do with it
An observability setup nobody trusts is worse than none, because it creates a false sense that someone would notice if something broke.
A short checklist before you call observability done
You're in reasonable shape once you can say yes to all of these: you can trace a single slow or failed request across every service it touched, your alert thresholds are based on your own traffic patterns, every alert has a clear owner who knows what action it means, and you've actually tested that a real problem would trigger the alert you think it would. If any of those is a no, that's where to focus next, not on adding another dashboard. See how observability platforms differ once you're choosing a specific one.
How this changes as you move from one service to many
A single service, running on one or two hosts, can usually get by on decent logging and a handful of metrics, because there's nowhere else for a problem to hide. Once a request has to travel through several services to get answered, the same logging and metrics stop being enough, because a metric can tell you that something is slow without telling you which of the several services along the way is responsible.
That transition is the point to add tracing, not before, since it's genuinely more work to set up and maintain than logs or metrics alone. Teams that add it too early spend effort on infrastructure their current system doesn't need yet, and teams that wait too long spend a painful incident or two trying to reconstruct a request's path by hand from scattered logs across services.
What Good Looks Like
Good here means an on-call engineer can trace a real incident to its cause using your own tools, and every alert that fires has an owner who trusts it enough to act on it.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How many alerts is too many for a small engineering team?
There's no fixed number, but if your team routinely mutes or ignores alerts, you have too many or the wrong ones. A smaller set of alerts tied directly to user-facing impact, that people actually trust and act on, beats a large set that gets tuned out.
Do we need distributed tracing if we only have a few services?
Probably not yet. Tracing earns its cost once a single request routinely touches enough services that logs and metrics alone can't tell you where time or errors are coming from. A smaller system can usually get by on good logging and metrics for longer.
Should on-call engineers be the ones who build the alerts?
They should at least review them closely. An alert built by someone who's never carried the pager for that system tends to fire on the wrong signal or miss the one that actually matters, because they haven't felt the cost of a bad alert firsthand.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Datadog vs New Relic vs Dynatrace: Cloud Observability Platforms Compared
Compare Datadog, New Relic, and Dynatrace for cloud observability: log ingestion costs, distributed tracing, APM overhead, and MTTR compression.
What to Actually Monitor Before You Buy an Observability Tool
Answers to the questions engineering teams actually ask before setting up monitoring: what to track, how many alerts is too many, and when to add tracing.
How to Ship a Risky Change Without a 2am Rollback
A concrete walkthrough of how to plan a risky production deployment: how to split it, what to watch, and when to decide the rollback trigger.
Build or Buy for Verifying Every Device That Connects?
How to split device identity from device posture checking, what building either one in house actually costs, and where a platform earns its keep instead.
A 30-Minute Check for Blind Spots in Your Observability Setup
A quick, practical checklist CTOs can run in 30 minutes to find the gaps in telemetry and alerting that usually surface during an incident instead.
Diagnosing Slow Requests Before You Blame the Database
A step-by-step way to find out whether a slowdown is the network, the app, or the database, before you add caching or upgrade infrastructure to fix it.