What to Instrument First in a Distributed System
Observability isn't logs, metrics, and traces sitting in three separate tools. It's the ability to answer one question fast during an incident: what broke, and why, without paging four teams to find out whose service is actually at fault.
Most distributed systems collect plenty of data and still fail that test, because the data isn't connected. An engineer ends up tabbing between five dashboards and three log tools trying to reconstruct what a single request actually did, which is slow exactly when speed matters most.
Here's what to fix first, and what to leave for later.
Why start with correlation instead of data volume?
Before adding another dashboard, make sure a single request can be traced across every service it touches. That means a request ID generated at the edge and propagated through every downstream call, logged consistently so you can pull every line related to one failing request in one query instead of grepping five services by hand.
Teams with plenty of logs but no correlation ID still end up guessing which service log to check first during an incident, which is the exact problem observability is supposed to solve.
Propagating the ID is the hard part in practice, not generating it. Every internal client, queue message, and background job needs to carry it forward, which usually means adding it to a shared library once rather than hoping every team remembers on their own.
Why alert on what users would notice, not internal noise?
Say CPU on one box runs hot: that isn't an incident if nothing downstream is affected. An error rate climbing on a customer-facing endpoint is. Page on symptoms tied to user impact, an SLO burning down faster than expected, elevated error rates, requests timing out, and route internal resource signals to dashboards instead of pages.
Tie your alert thresholds to the availability target you've actually committed to. A service promising 99.9 percent has roughly 8.76 hours of downtime to spend across the whole year1. An alert that doesn't fire until you've already burned a meaningful chunk of that budget in one incident is set too late.
A worked example: an intermittent error that only shows up under load
Say customers occasionally report a failed checkout, but it never reproduces when an engineer tries it manually. Without correlation, this stays a mystery. With a request ID threaded through logs, the pattern becomes visible: the failures cluster around moments when a downstream inventory service's connection pool is exhausted, which only happens under concurrent load an engineer testing alone can't recreate.
The fix here isn't better error messages, it's a metric on connection pool saturation that would have shown the exhaustion happening well before customers noticed anything.
The dashboards nobody looks at
- A dashboard built for one incident and never updated as the system changed
- Metrics with no owner, so nobody notices when they stop reporting data at all
- Alerts tuned so loose they never fire, or so tight the on-call engineer mutes them out of habit
- Logs with no consistent structure, so search means guessing at field names service by service
A dashboard that isn't tied to an action someone takes when it changes isn't observability, it's wallpaper.
Structured logs beat clever log messages
A log line that reads well in a terminal is usually hard to query at scale. Structured logging, consistent field names for request ID, service name, status, and duration across every service, turns your logs into something you can actually filter and aggregate instead of just read.
This matters more as the number of services grows. Two services with inconsistent log formats is an annoyance. Twenty services with inconsistent formats means every incident starts with reverse-engineering that day's logging conventions before you can even start debugging.
Link every page to a runbook, not tribal knowledge
An alert that pages someone with no next step attached turns every incident into a puzzle the on-call engineer has to solve from scratch, usually at the worst possible hour. Attach a short runbook to every alert: what this usually means, where to look first, and who to escalate to if it's outside your area.
The runbook doesn't need to be exhaustive. A few lines pointing at the right dashboard and the service's known failure modes turns a 2am page from a guessing game into a checklist, which matters most for whoever is newest on the rotation.
What Good Looks Like
Good observability means any engineer can trace a failing request across every service it touched, alerts fire on user-visible symptoms tied to your actual SLOs, and logs share a consistent structure across the whole system.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How many alerts should be paging someone at 3am?
As few as possible, and each one should require a human action that can't wait until morning. If an alert fires and the response is always "check again in the morning," it shouldn't be a page, it should be a dashboard entry reviewed during business hours.
Do we need distributed tracing if we only have five services?
Probably, once any single request touches more than two of them. The value of tracing scales with the number of hops a request takes, not the total service count, and five services calling each other in a chain already makes manual log correlation slow.
What's the difference between an SLO and an alert threshold?
An SLO is the target you've promised for something like availability or error rate. An alert threshold is set well before you'd breach that target, so you have time to react. Setting them to the same number means you only find out after you've already failed the commitment.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
A Production Deployment Checklist That Actually Catches Problems
A stage-by-stage deployment checklist for distributed systems, covering rollback readiness, dependency ordering, and the checks teams skip under pressure.
Verifying Devices Before They Touch Production, Not After
How to build device verification into a zero-trust rollout, what actually counts as a trust signal, and where teams stop checking too early.
Finding the Real Source of Latency in a Distributed System
A decision guide for narrowing down whether a slow request is a network problem, a database problem, a queue problem, or your own code.
Cache Invalidation Is Still the Hard Part
A practical guide to choosing a caching layer and, more importantly, keeping it from serving stale or wrong data across a distributed system.
Why Your Redis Lock Let Two Jobs Run at Once (and How to Fix It)
A walkthrough of a real double-charge bug caused by a Redis lock's TTL expiring mid-job, and the fencing-token pattern that actually fixes it.
Load Testing Numbers That Don't Match What Users Actually Feel
Why a clean throughput benchmark often fails to predict real-world scaling behavior, and how to build one around your real traffic mix and first bottleneck.