Distributed Systems & Enterprise ResiliencePlaybook3 min readUpdated September 2026

A Worksheet for Deciding What to Instrument With OpenTelemetry First

Rolling out OpenTelemetry across an entire service fleet at once is how a lot of tracing initiatives stall: months of instrumentation work before anyone gets a usable trace, and the services most likely to cause the next incident often aren't the ones that got instrumented first, because the easiest services to instrument tend to be the smallest and least central ones.

A short worksheet fixes the ordering problem. It doesn't need a project management tool, just three columns and an honest scoring pass.

How do you decide what to instrument with OpenTelemetry first?

List your top four or five user-facing flows down the side (checkout, login, the core API call customers depend on) and, for each, the services it actually touches. That list alone is often the first useful output: many teams discover a flow touches more services than anyone remembered, because ownership drifted service by service over time and nobody had recently mapped the whole path end to end.

Score for blast radius, not for ease of instrumentation

For each service on the list, score three things: how often it's named in a recent incident or postmortem, how many other services call it (a rough proxy for blast radius if it fails), and whether it already emits any structured logs you could correlate against a trace once one exists. Instrument the highest-scoring services first, even if their SDK support is less mature or the integration is more work, because those are the services where a trace actually shortens your next incident, not the ones that were easiest to wire up over an afternoon.

Propagate context across every hop or the trace is a stub

A span with no propagated context from the caller isn't part of a distributed trace, it's an isolated log entry that happens to use tracing vocabulary. Every hop, including message queues, background job processors, and any internal HTTP call between services, needs to pass the trace context forward. It's tempting to instrument the easy, synchronous request path first and leave the asynchronous hops for later, but an incident that involves a queue is exactly the kind of incident a trace is supposed to help with, and a broken chain there is where teams usually discover the gap, mid-incident, instead of ahead of time.

How should you sample traces to keep errors and slow requests?

Sampling every request in a high-traffic service gets expensive fast, in both storage cost and query performance against the tracing backend. Keep all traces that include an error or exceed a latency threshold, and sample the rest at a much lower rate; that combination preserves exactly the traces you'll actually want during an investigation while controlling cost for the traffic that isn't interesting. A flat sampling rate across all traffic, chosen mainly to hit a cost target, tends to throw away the rare, expensive traces that would have mattered most.

The mistake: instrumenting the easy services first

Services with generous, well-documented SDK support get instrumented quickly, and it feels like progress. If those services rarely appear in your incidents, the tracing rollout has produced coverage without producing much diagnostic value where it's actually needed. Revisit the worksheet's score column periodically rather than letting 'which services are left to do' default to whatever order is easiest, and be willing to reprioritize a harder integration ahead of an easier one if a new incident makes its blast-radius score go up. Comparing tracing and observability backends is worth doing once you know which services actually need to send data there.

Keeping the worksheet current instead of treating it as a one-time exercise

Service ownership and call patterns shift as a system grows, and a worksheet built once during the initial rollout goes stale the same way an architecture diagram does. Revisit it after any incident that involved a service not already near the top of the list, and after any quarter where a new service picked up meaningfully more traffic than it had when it was last scored. Treating the worksheet as a living document, reviewed on the same cadence as an on-call retro, keeps the instrumentation priority aligned with where the system's actual risk has moved.

Use the worksheet in this order:

  1. List your top user-facing flows down the side, and note the services each flow actually touches.
  2. Score each service on how often it appears in incidents, how many other services call it, and whether it already emits structured logs.
  3. Instrument the highest-scoring services first, rather than the ones that are easiest to set up.
  4. Propagate trace context across every hop, including message queues, background workers and internal HTTP calls.
  5. Keep all error and slow-request traces, sample the rest at a lower rate, and revisit the worksheet after any incident.
Executive Capability Standard

What Good Looks Like

Good tracing coverage instruments services in order of incident frequency and blast radius, propagates context across every hop including asynchronous ones, and samples deliberately instead of applying one flat rate to all traffic.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Build the three-column worksheet: top user flows, the services each one touches, and a blast-radius score per service based on incident history.
2. Do Manually:Manually trace one real incident's timeline against the worksheet to confirm the highest-scored services are actually the ones that would have helped.
3. Delegate:Have an engineer instrument the highest-scoring services first, including their asynchronous hops, not just their synchronous request paths.
4. Automate:Set up error-and-slow-request-biased sampling so the traces most likely to matter during an incident aren't the ones getting sampled out.
5. Buy:Bring in observability advisory if the worksheet reveals a flow touching more services than your current tracing backend can affordably cover at full fidelity.

How to Get Started

Frequently Asked Questions

Do we need full instrumentation before tracing is useful at all?

No. A trace that covers even the two or three highest-blast-radius services in a flow is useful immediately, since it shortens the time to identify which service in the chain is the actual problem. Full coverage is a long-term goal, not a prerequisite for getting value from the first services you instrument.

How do we handle context propagation through a message queue?

Most OpenTelemetry instrumentation libraries for common queue systems support injecting trace context into message headers or metadata and extracting it on the consumer side. Confirm your specific queue client's instrumentation actually does this, since it's a common gap even in otherwise well-instrumented systems.

What sampling rate should we start with for high-traffic services?

Keep all error and slow-request traces regardless of rate, then sample the remaining normal traffic low enough to control cost, often in the low single-digit percent for very high-traffic services. Adjust based on your actual trace storage cost and how often you find yourself wishing you had a trace that got sampled out.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides