Developer Productivity & Platform EngineeringPlaybook3 min readUpdated September 2026

What to Actually Monitor Before You Buy an Observability Tool

Teams usually approach observability backward: they buy a tool, turn on every default dashboard it ships with, and end up with hundreds of metrics and no clearer sense of when something is actually wrong. The better starting point is a short list of questions, answered honestly, before any tool gets chosen.

The questions below are the ones Taj, MeetMyCTO's AI CTO, hears most often from teams setting up monitoring for the first time or migrating off a tool they've outgrown. Answering them first usually changes which tool makes sense, and sometimes removes the need for a new tool at all.

What should we actually be alerted on?

Alert on symptoms your users would notice, not on every internal metric that moves. A good starting set for most services is: error rate above a threshold, latency above your budget for a sustained window, and a health check failing for more than a minute or two. Everything else (CPU usage, queue depth, cache hit rate) is useful for diagnosis once you're alerted, but paging someone on CPU usage alone tends to produce alerts nobody trusts, and alerts nobody trusts get muted.

A useful rule when you're deciding whether something deserves a page versus a dashboard panel: would a user notice this on their own within the next ten minutes if nobody fixed it? If yes, page. If the honest answer is "only if it gets much worse," it belongs on a dashboard for the next business-hours review instead.

A starting set of user-facing alerts for most services:

  • Error rate above a threshold you have agreed on for that service.
  • Latency above your budget for a sustained window, not a single slow request.
  • A health check failing for more than a minute or two.
  • Keep CPU usage, queue depth and cache hit rate on dashboards for diagnosis instead of paging anyone on them alone.

How do we know if we have too many alerts?

If your on-call engineer regularly acknowledges an alert without taking action, that alert is noise and should be deleted or turned into a dashboard panel instead. A useful gut check: pull your last month of alerts and count how many led to an actual code or config change versus how many were acknowledged and ignored. Teams are often surprised that fewer than half their alerts have ever driven a real fix.

Alert fatigue also shows up as a delay pattern, not just a volume pattern: if the average time to acknowledge an alert has crept up over the last few months, that's often a sign of fatigue setting in before anyone consciously notices it. Track that number alongside your alert count, since it usually moves first.

When do we need distributed tracing instead of just logs and metrics?

Logs and metrics tell you that something is slow or failing; tracing tells you which hop in a multi-service request is responsible. If your architecture is a single service talking to one database, metrics and structured logs are enough. Once a single user request routes through three or more services, tracing pays for itself the first time you spend an afternoon guessing which service is the actual bottleneck instead of seeing it directly in a trace.

Adopting tracing doesn't require instrumenting your whole codebase at once. Start with your single highest-traffic or highest-complaint request path, add spans at each service boundary along that one path, and expand from there. Teams that try to instrument everything on day one usually stall out before finishing any single path well enough to be useful during an incident, and which tracing platform you pick matters far less than which path you chose to instrument first.

What's a reasonable uptime target to design around?

Pick a target based on what your product actually needs, not what sounds impressive. A 99.9% target allows about 8.76 hours of downtime a year; a 99.99% target shrinks that to roughly 53 minutes1. Most early-stage products are better served by the looser target: chasing 99.99% availability means multi-region failover, automated runbooks and on-call staffing that a ten-person engineering team usually isn't ready to operate, and the operational overhead can cost more than the downtime it prevents1.

Write the target down as an error budget, not just an aspiration: if you're targeting 99.9%, you have roughly 43 minutes of downtime to spend each month before you've broken the commitment. Spend it deliberately, a risky migration, a provider maintenance window, rather than discovering after the fact that an unplanned outage used it up.

What should the on-call dashboard actually show?

One screen, not a wall of panels: current error rate and latency against your budget for each critical service, the status of your last deployment, and a link to recent alerts. Anything an on-call engineer has to dig for during an incident should either be on that one screen or one click away. If you find your on-call rotation building custom queries during an active incident, that's a sign your default dashboard is missing something they need every time.

After your next incident, ask whoever was on call what they had to look up manually that wasn't already visible. Add that one thing to the dashboard and stop there; a dashboard that tries to answer every possible question up front usually ends up answering none of them quickly.

Executive Capability Standard

What Good Looks Like

Good observability means alerts that map to what users would notice, a dashboard that answers most on-call questions without a custom query, and tracing once request paths span more than a couple of services.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Pull your last month of alerts and tag each one as led-to-a-fix or acknowledged-and-ignored to see how much noise you actually have.
2. Do Manually:Set up symptom-based alerts (error rate, latency, health checks) by hand for your two or three most important services.
3. Delegate:Have a platform engineer own building the single on-call dashboard and pruning alerts that never lead to action.
4. Automate:Wire distributed tracing into your request path so on-call engineers can see the bottleneck directly instead of reconstructing it from logs.
5. Buy:Bring in an SRE consultant for a short engagement to set error budgets and an on-call rotation if you're staffing on-call for the first time.

How to Get Started

Frequently Asked Questions

Do we need application performance monitoring and infrastructure monitoring as separate tools?

Not necessarily. Most modern observability platforms cover both in one product now. What matters more than the tool count is whether the data from each layer is correlated, so an infrastructure alert and the application trace that explains it show up together.

How many services justify setting up a full observability stack versus just logging?

Once you have more than two or three services calling each other, or once a single incident has taken more than an hour to diagnose because no one could see the full request path, it's time. Below that, structured logs with good search are usually enough.

Should error budgets be set per team or centrally?

Set the target centrally (what uptime the product needs) but let each team own the budget for their own services, since only the team running a service can make the tradeoffs between reliability work and feature work day to day.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides