Cloud FinOps & Infrastructure ScalingPlaybook3 min readUpdated September 2026

A Practical Checklist for Observability That Gets Used

Most teams don't have too little observability. They have a pile of dashboards and alerts that no one trusts enough to act on during an actual incident, which is functionally the same as having none.

Here's a checklist for building the kind of observability that gets used, not just installed.

Instrument for the questions you'll actually ask during an incident

Before adding another dashboard, list the two or three questions you always ask when something breaks: is this affecting all users or a subset, is it a specific endpoint or everything, did a deploy just go out. Build your instrumentation to answer those questions directly, rather than collecting every metric a tool offers by default.

A dashboard built around a real incident question gets opened during the next incident. A generic one, built because a tool suggested it, usually doesn't.

Set alert thresholds off your own baseline, not a generic default

A default threshold, copied from documentation or a template, either fires constantly on normal variation or stays silent through a real problem, because it was never set against your actual traffic. Pull a few weeks of your own data first, and set thresholds relative to that baseline.

This matters most for anything with a strong daily or weekly pattern. An alert tuned for your quiet Tuesday morning will misfire every single Monday if you didn't account for the cycle.

The three signals worth paying for: logs, metrics, and traces

Logs tell you what happened in detail but are slow to search across a whole system. Metrics tell you that something changed, quickly, but not why. Traces connect a single request across every service it touched, which is what actually lets you find where in a distributed system something went wrong.

Most teams over-invest in logs early because they're the easiest to add, then discover during a real incident that tracing a request across services is what they actually needed and never built.

Pitfalls: dashboards nobody opens and alerts nobody trusts

Watch for these patterns, since each one quietly erodes trust in the whole system:

  • An alert that's fired incorrectly enough times that people mute it instead of investigating
  • A dashboard built once and never updated as the system it monitors changed
  • Alerting on a symptom (high CPU) instead of user impact (failed requests, slow responses)
  • No clear owner for a given alert, so it pages a rotation that doesn't know what to do with it

An observability setup nobody trusts is worse than none, because it creates a false sense that someone would notice if something broke.

A short checklist before you call observability done

You're in reasonable shape once you can say yes to all of these: you can trace a single slow or failed request across every service it touched, your alert thresholds are based on your own traffic patterns, every alert has a clear owner who knows what action it means, and you've actually tested that a real problem would trigger the alert you think it would. If any of those is a no, that's where to focus next, not on adding another dashboard. See how observability platforms differ once you're choosing a specific one.

How this changes as you move from one service to many

A single service, running on one or two hosts, can usually get by on decent logging and a handful of metrics, because there's nowhere else for a problem to hide. Once a request has to travel through several services to get answered, the same logging and metrics stop being enough, because a metric can tell you that something is slow without telling you which of the several services along the way is responsible.

That transition is the point to add tracing, not before, since it's genuinely more work to set up and maintain than logs or metrics alone. Teams that add it too early spend effort on infrastructure their current system doesn't need yet, and teams that wait too long spend a painful incident or two trying to reconstruct a request's path by hand from scattered logs across services.

Executive Capability Standard

What Good Looks Like

Good here means an on-call engineer can trace a real incident to its cause using your own tools, and every alert that fires has an owner who trusts it enough to act on it.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Learn to read your own traffic baseline well enough to know what a normal Tuesday looks like versus a real anomaly.
2. Do Manually:Build one dashboard and one alert around a real incident question you've actually needed answered, and test that it fires correctly.
3. Delegate:Give the on-call rotation ownership of tuning their own alerts, since they feel the cost of a bad one first.
4. Automate:Set alert thresholds programmatically off rolling baselines instead of static numbers, so they adjust as traffic patterns shift.
5. Buy:Bring in a specialist to set up distributed tracing once your system has grown past what logs and metrics alone can explain.

How to Get Started

Frequently Asked Questions

How many alerts is too many for a small engineering team?

There's no fixed number, but if your team routinely mutes or ignores alerts, you have too many or the wrong ones. A smaller set of alerts tied directly to user-facing impact, that people actually trust and act on, beats a large set that gets tuned out.

Do we need distributed tracing if we only have a few services?

Probably not yet. Tracing earns its cost once a single request routinely touches enough services that logs and metrics alone can't tell you where time or errors are coming from. A smaller system can usually get by on good logging and metrics for longer.

Should on-call engineers be the ones who build the alerts?

They should at least review them closely. An alert built by someone who's never carried the pager for that system tends to fire on the wrong signal or miss the one that actually matters, because they haven't felt the cost of a bad alert firsthand.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides