A Worksheet for Deciding What to Instrument With OpenTelemetry First
Rolling out OpenTelemetry across an entire service fleet at once is how a lot of tracing initiatives stall: months of instrumentation work before anyone gets a usable trace, and the services most likely to cause the next incident often aren't the ones that got instrumented first, because the easiest services to instrument tend to be the smallest and least central ones.
A short worksheet fixes the ordering problem. It doesn't need a project management tool, just three columns and an honest scoring pass.
How do you decide what to instrument with OpenTelemetry first?
List your top four or five user-facing flows down the side (checkout, login, the core API call customers depend on) and, for each, the services it actually touches. That list alone is often the first useful output: many teams discover a flow touches more services than anyone remembered, because ownership drifted service by service over time and nobody had recently mapped the whole path end to end.
Score for blast radius, not for ease of instrumentation
For each service on the list, score three things: how often it's named in a recent incident or postmortem, how many other services call it (a rough proxy for blast radius if it fails), and whether it already emits any structured logs you could correlate against a trace once one exists. Instrument the highest-scoring services first, even if their SDK support is less mature or the integration is more work, because those are the services where a trace actually shortens your next incident, not the ones that were easiest to wire up over an afternoon.
Propagate context across every hop or the trace is a stub
A span with no propagated context from the caller isn't part of a distributed trace, it's an isolated log entry that happens to use tracing vocabulary. Every hop, including message queues, background job processors, and any internal HTTP call between services, needs to pass the trace context forward. It's tempting to instrument the easy, synchronous request path first and leave the asynchronous hops for later, but an incident that involves a queue is exactly the kind of incident a trace is supposed to help with, and a broken chain there is where teams usually discover the gap, mid-incident, instead of ahead of time.
How should you sample traces to keep errors and slow requests?
Sampling every request in a high-traffic service gets expensive fast, in both storage cost and query performance against the tracing backend. Keep all traces that include an error or exceed a latency threshold, and sample the rest at a much lower rate; that combination preserves exactly the traces you'll actually want during an investigation while controlling cost for the traffic that isn't interesting. A flat sampling rate across all traffic, chosen mainly to hit a cost target, tends to throw away the rare, expensive traces that would have mattered most.
The mistake: instrumenting the easy services first
Services with generous, well-documented SDK support get instrumented quickly, and it feels like progress. If those services rarely appear in your incidents, the tracing rollout has produced coverage without producing much diagnostic value where it's actually needed. Revisit the worksheet's score column periodically rather than letting 'which services are left to do' default to whatever order is easiest, and be willing to reprioritize a harder integration ahead of an easier one if a new incident makes its blast-radius score go up. Comparing tracing and observability backends is worth doing once you know which services actually need to send data there.
Keeping the worksheet current instead of treating it as a one-time exercise
Service ownership and call patterns shift as a system grows, and a worksheet built once during the initial rollout goes stale the same way an architecture diagram does. Revisit it after any incident that involved a service not already near the top of the list, and after any quarter where a new service picked up meaningfully more traffic than it had when it was last scored. Treating the worksheet as a living document, reviewed on the same cadence as an on-call retro, keeps the instrumentation priority aligned with where the system's actual risk has moved.
Use the worksheet in this order:
- List your top user-facing flows down the side, and note the services each flow actually touches.
- Score each service on how often it appears in incidents, how many other services call it, and whether it already emits structured logs.
- Instrument the highest-scoring services first, rather than the ones that are easiest to set up.
- Propagate trace context across every hop, including message queues, background workers and internal HTTP calls.
- Keep all error and slow-request traces, sample the rest at a lower rate, and revisit the worksheet after any incident.
What Good Looks Like
Good tracing coverage instruments services in order of incident frequency and blast radius, propagates context across every hop including asynchronous ones, and samples deliberately instead of applying one flat rate to all traffic.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Do we need full instrumentation before tracing is useful at all?
No. A trace that covers even the two or three highest-blast-radius services in a flow is useful immediately, since it shortens the time to identify which service in the chain is the actual problem. Full coverage is a long-term goal, not a prerequisite for getting value from the first services you instrument.
How do we handle context propagation through a message queue?
Most OpenTelemetry instrumentation libraries for common queue systems support injecting trace context into message headers or metadata and extracting it on the consumer side. Confirm your specific queue client's instrumentation actually does this, since it's a common gap even in otherwise well-instrumented systems.
What sampling rate should we start with for high-traffic services?
Keep all error and slow-request traces regardless of rate, then sample the remaining normal traffic low enough to control cost, often in the low single-digit percent for very high-traffic services. Adjust based on your actual trace storage cost and how often you find yourself wishing you had a trace that got sampled out.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Datadog vs New Relic vs Dynatrace: Cloud Observability Platforms Compared
Compare Datadog, New Relic, and Dynatrace for cloud observability: log ingestion costs, distributed tracing, APM overhead, and MTTR compression.
Why Your Redis Lock Let Two Jobs Run at Once (and How to Fix It)
A walkthrough of a real double-charge bug caused by a Redis lock's TTL expiring mid-job, and the fencing-token pattern that actually fixes it.
Cache Invalidation Is Still the Hard Part
A practical guide to choosing a caching layer and, more importantly, keeping it from serving stale or wrong data across a distributed system.
A Production Deployment Checklist That Actually Catches Problems
A stage-by-stage deployment checklist for distributed systems, covering rollback readiness, dependency ordering, and the checks teams skip under pressure.
Verifying Devices Before They Touch Production, Not After
How to build device verification into a zero-trust rollout, what actually counts as a trust signal, and where teams stop checking too early.
Finding the Real Source of Latency in a Distributed System
A decision guide for narrowing down whether a slow request is a network problem, a database problem, a queue problem, or your own code.