Distributed Systems & Enterprise ResiliencePlaybook3 min readUpdated September 2026

What to Instrument First in a Distributed System

Observability isn't logs, metrics, and traces sitting in three separate tools. It's the ability to answer one question fast during an incident: what broke, and why, without paging four teams to find out whose service is actually at fault.

Most distributed systems collect plenty of data and still fail that test, because the data isn't connected. An engineer ends up tabbing between five dashboards and three log tools trying to reconstruct what a single request actually did, which is slow exactly when speed matters most.

Here's what to fix first, and what to leave for later.

Why start with correlation instead of data volume?

Before adding another dashboard, make sure a single request can be traced across every service it touches. That means a request ID generated at the edge and propagated through every downstream call, logged consistently so you can pull every line related to one failing request in one query instead of grepping five services by hand.

Teams with plenty of logs but no correlation ID still end up guessing which service log to check first during an incident, which is the exact problem observability is supposed to solve.

Propagating the ID is the hard part in practice, not generating it. Every internal client, queue message, and background job needs to carry it forward, which usually means adding it to a shared library once rather than hoping every team remembers on their own.

Why alert on what users would notice, not internal noise?

Say CPU on one box runs hot: that isn't an incident if nothing downstream is affected. An error rate climbing on a customer-facing endpoint is. Page on symptoms tied to user impact, an SLO burning down faster than expected, elevated error rates, requests timing out, and route internal resource signals to dashboards instead of pages.

Tie your alert thresholds to the availability target you've actually committed to. A service promising 99.9 percent has roughly 8.76 hours of downtime to spend across the whole year1. An alert that doesn't fire until you've already burned a meaningful chunk of that budget in one incident is set too late.

A worked example: an intermittent error that only shows up under load

Say customers occasionally report a failed checkout, but it never reproduces when an engineer tries it manually. Without correlation, this stays a mystery. With a request ID threaded through logs, the pattern becomes visible: the failures cluster around moments when a downstream inventory service's connection pool is exhausted, which only happens under concurrent load an engineer testing alone can't recreate.

The fix here isn't better error messages, it's a metric on connection pool saturation that would have shown the exhaustion happening well before customers noticed anything.

The dashboards nobody looks at

  • A dashboard built for one incident and never updated as the system changed
  • Metrics with no owner, so nobody notices when they stop reporting data at all
  • Alerts tuned so loose they never fire, or so tight the on-call engineer mutes them out of habit
  • Logs with no consistent structure, so search means guessing at field names service by service

A dashboard that isn't tied to an action someone takes when it changes isn't observability, it's wallpaper.

Structured logs beat clever log messages

A log line that reads well in a terminal is usually hard to query at scale. Structured logging, consistent field names for request ID, service name, status, and duration across every service, turns your logs into something you can actually filter and aggregate instead of just read.

This matters more as the number of services grows. Two services with inconsistent log formats is an annoyance. Twenty services with inconsistent formats means every incident starts with reverse-engineering that day's logging conventions before you can even start debugging.

Link every page to a runbook, not tribal knowledge

An alert that pages someone with no next step attached turns every incident into a puzzle the on-call engineer has to solve from scratch, usually at the worst possible hour. Attach a short runbook to every alert: what this usually means, where to look first, and who to escalate to if it's outside your area.

The runbook doesn't need to be exhaustive. A few lines pointing at the right dashboard and the service's known failure modes turns a 2am page from a guessing game into a checklist, which matters most for whoever is newest on the rotation.

Executive Capability Standard

What Good Looks Like

Good observability means any engineer can trace a failing request across every service it touched, alerts fire on user-visible symptoms tied to your actual SLOs, and logs share a consistent structure across the whole system.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Pick one recent incident and try to reconstruct the full request path through logs alone, noting every place the trail went cold.
2. Do Manually:Add a request ID to your highest-traffic endpoint and thread it through every service it calls by hand.
3. Delegate:Assign an engineer to own alert quality, reviewing which pages fired in the last month and cutting the ones that didn't need a human.
4. Automate:Standardize structured logging across every service and instrument distributed tracing so correlation happens automatically, not by convention.
5. Buy:Bring in a fractional CTO or observability specialist once your incident response time stays slow despite having the data, since the gap is usually in how it's organized, not how much of it exists.

How to Get Started

Frequently Asked Questions

How many alerts should be paging someone at 3am?

As few as possible, and each one should require a human action that can't wait until morning. If an alert fires and the response is always "check again in the morning," it shouldn't be a page, it should be a dashboard entry reviewed during business hours.

Do we need distributed tracing if we only have five services?

Probably, once any single request touches more than two of them. The value of tracing scales with the number of hops a request takes, not the total service count, and five services calling each other in a chain already makes manual log correlation slow.

What's the difference between an SLO and an alert threshold?

An SLO is the target you've promised for something like availability or error rate. An alert threshold is set well before you'd breach that target, so you have time to react. Setting them to the same number means you only find out after you've already failed the commitment.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides