Setting Up Distributed Tracing Without Drowning in Spans
Distributed tracing follows a single request through every service it touches, and OpenTelemetry is the open standard behind most modern setups. Instrumenting services is easier than it used to be, but that does not make the data useful, and a poor rollout can bury the one important trace under noise.
Here's how to roll it out so it actually helps during an incident, not just in a demo.
Which request path should you instrument first?
Instrumenting every service at once produces a flood of trace data before anyone has built the habit of actually using it, and most of that early data goes unread. Pick the one request path that's caused the most confusing incidents, like a checkout flow that touches five services, and instrument that path end to end first. A team that can reliably use tracing to debug one real path develops the instinct to instrument the next one well, rather than treating tracing as a checkbox rolled out everywhere at once.
Propagate context correctly, or the trace breaks silently
A trace only stays connected across services if each one correctly passes along the trace context, usually through a header, to the next service it calls. Miss this in one service, often an async job queue or a message broker that isn't designed with tracing in mind by default, and the trace simply stops there, showing a gap with no error, which is confusing precisely because nothing visibly failed. Test context propagation explicitly across every hop in your critical path, including the async ones, not just the synchronous HTTP calls that are easy to get right.
Add meaningful attributes, not just span names
A trace made of spans named only after their function, with no other context, tells you where time was spent but not why. Attach attributes that answer real debugging questions, such as which customer or order the request belongs to, what branch of logic was taken, or the size of the payload being processed. The difference between a trace that's decorative and one that's genuinely useful during an incident is almost always in these attributes, not in the raw timing data. Agree on a small, consistent set of attribute names across services too, so an engineer searching traces doesn't have to remember that one team calls it order id and another calls the same field something else entirely.
How much distributed tracing data should you sample?
Storing a full trace for every single request gets expensive fast at any real scale, and most of that stored data is never looked at. Sample a smaller percentage of normal traffic while ensuring every request that errors or exceeds a latency threshold is always kept, so the trace data you're paying to store and the trace data you actually need during an incident are the same data, not two different sets. Revisit the sampling rule periodically too, since a threshold tuned for last year's traffic pattern can quietly stop capturing the requests that matter most as the product changes.
A common pitfall: tracing that nobody knows how to read during an incident
Rolling out tracing infrastructure without training the team to actually use it during an incident means it sits unused exactly when it would help most, replaced by the old habit of grepping through logs from five different services by hand. Walk the team through a real past incident using the new tracing tool, showing concretely how it would have cut the time to find the cause, so the tool becomes part of the actual incident response habit instead of infrastructure nobody remembers exists under pressure. Revisit that walkthrough every time a new engineer joins the on-call rotation, since a tool nobody was trained on might as well not be running at all during a live incident.
Roll tracing out in this order:
- Pick the one request path that has caused the most confusing incidents and instrument it end to end first.
- Confirm trace context propagates across every hop, including job queues and message brokers that skip tracing by default.
- Add attributes such as customer, order, logic branch, and payload size, not just span names.
- Sample normal traffic deliberately, but always keep traces for requests that error or exceed a latency threshold.
- Walk the team through a real past incident using the tracing tool, so they reach for it during the next one.
What Good Looks Like
A useful distributed tracing setup instruments the request paths that actually cause confusing incidents, propagates context correctly across every service including async hops, and samples deliberately so the data that's kept is the data an engineer would actually need during an investigation.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Do we need to instrument every service before tracing is useful?
No. Tracing is useful the moment a single critical request path is fully instrumented end to end, even if most of the codebase isn't yet. Waiting for full coverage before using it at all means delaying real value for a completeness goal that isn't actually required to start debugging with it.
What's the difference between logging and distributed tracing?
Logs capture discrete events within one service in isolation. A trace connects those events across every service a single request touched, in the order it touched them, which is exactly the view you need to answer where time went across a multi-service request, not just within one service's boundary.
How much of our traffic should we actually sample and store?
Keep a small sample of normal traffic and always keep every trace that errors or runs slower than your latency threshold. There is no universal percentage, but this approach preserves baseline visibility while retaining the traces you disproportionately need during an investigation.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Instrumenting Tracing Without Drowning in Spans
Turning on auto-instrumentation everywhere produces a bill bigger than the incident it was meant to debug. A rollout order that avoids that.
A Worksheet for Deciding What to Instrument With OpenTelemetry First
A simple worksheet for prioritizing which services get OpenTelemetry instrumentation first, based on incident history and blast radius, not ease of setup.
A Go-Live Checklist for Shipping a RAG Pipeline
The specific steps to check before a RAG pipeline goes live: index warm-up, model version pinning, a canary check, and a real rollback plan.
Setting Up OpenTelemetry So Traces Actually Connect
A worksheet for rolling out OpenTelemetry: context propagation across queues, sampling that keeps errors, and controlling cardinality before costs spike.
The OpenTelemetry Rollout Order That Keeps the Trace Bill Sane
A practical runbook for adopting OpenTelemetry tracing across a microservices stack: what to instrument first, sampling strategy, and cost control.
Three Places to Cache in a RAG Pipeline, and What Each Buys You
Embedding caches, chunk-set caches, and shared versus per-instance caching each solve a different RAG cost or latency problem. Here's how to pick.