Getting Real Value Out of OpenTelemetry Instead of Just Installing It
A lot of teams install OpenTelemetry, wire up the SDK across their services, and consider distributed tracing done. Then a real incident happens, a request that spans six services and takes eight seconds when it should take two hundred milliseconds, and the trace they pull up is a wall of spans with no clear story about where the time actually went.
Instrumentation and usefulness are two different milestones, and most of the value sits in the gap between them.
Why Installing the SDK Isn't the Hard Part
OpenTelemetry's auto-instrumentation libraries make it genuinely easy to get spans flowing from most common frameworks and libraries without writing custom code. That's real progress over the manual instrumentation era, but it also means teams stop there, with traces that show HTTP calls and database queries but nothing about what actually happened inside your own business logic. The trace shows you made a call to another service and it took three seconds; it doesn't show you why, unless you've added spans around the specific logic that matters.
Instrument the Decisions, Not Just the Calls
The traces that actually help during an incident have custom spans around the parts of your code where a decision gets made or where variable-cost work happens: a retry loop, a cache lookup that might miss, a batch operation whose size depends on the input. Auto-instrumentation gives you the network boundary for free; it can't know which internal function is your team's actual bottleneck. Pick the handful of services most often involved in past incidents and add targeted spans there first, rather than trying to instrument everything evenly.
Propagate Context Correctly, or the Trace Breaks at the Worst Moment
A distributed trace only stays connected across service boundaries if trace context gets propagated on every call, HTTP headers, queue message metadata, background job payloads. It's common for teams to get this right for synchronous HTTP calls and miss it for asynchronous work: a message published to a queue that doesn't carry the trace ID forward breaks the chain right at the point where debugging usually gets hardest, an async handoff. Audit every place your architecture crosses a synchronous-to-asynchronous boundary and confirm trace context survives it.
Sampling: The Setting Most Teams Get Wrong by Default
Tracing every single request at high fidelity gets expensive fast, so most setups sample, keeping only a percentage of traces. The problem is that a flat percentage sample, say ten percent, is likely to miss the rare, slow, or erroring requests that matter most for debugging, precisely because they're rare. Tail-based sampling, which decides whether to keep a trace after seeing how it completed, keeping all the slow or erroring ones and a smaller sample of the normal ones, gets you a much more useful dataset for the same storage cost. It's more complex to set up than head-based sampling, but it's usually worth the tradeoff once tracing volume becomes a real cost line.
Making Traces Useful During an Actual Incident, Not Just in a Postmortem
A trace that's only ever pulled up after the fact, during a postmortem, is missing half its value. Link your tracing tool directly from your alerting and dashboards, so when a latency alert fires, the on-call engineer is one click from a relevant trace instead of starting a manual search through a tracing UI under pressure. Teams that ship every day especially benefit from this, since a fast deploy cadence means more candidate changes to check against when something goes wrong, and a trace that's hard to find slows that check down exactly when speed matters1.
Checks that turn installed tracing into usable tracing:
- Add custom spans around retry loops, cache lookups that might miss, and batch operations whose size depends on the input.
- Propagate trace context on every hop, including queue message metadata and background job payloads, not only synchronous HTTP headers.
- Consider tail-based sampling so rare slow or erroring requests are kept, since a flat percentage sample tends to miss exactly those.
- Link alerts and dashboards directly to a relevant trace so the on-call engineer is one click away from it during an incident.
A Worked Example: Chasing an Eight-Second Request
Say a checkout request that should take two hundred milliseconds is taking eight seconds for a subset of users, and the trace shows five spans: the API gateway, an auth check, an inventory service call, a pricing service call, and the database write, with most of the time sitting inside the pricing service span and nothing more specific than that. Without a custom span inside the pricing service around its actual logic, discount lookup, tax calculation, currency conversion, you're left guessing which of those three did the damage. Add spans around each one, and the next occurrence points straight at the actual bottleneck instead of just the service boundary that contains it.
What Good Looks Like
Useful tracing means the services most often involved in past incidents have custom spans around their actual decision points, trace context survives every async boundary, and an on-call engineer can reach a relevant trace in one click from an alert.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Do we need to instrument every service before tracing is useful at all?
No. Partial coverage still helps, especially if you start with the services most often implicated in past incidents. A trace that goes dark at an uninstrumented service still tells you where to look next, even if it doesn't tell you the full story yet. Expand coverage service by service rather than waiting to instrument everything before getting value.
What's the difference between logging, metrics, and tracing?
Metrics tell you something is wrong in aggregate, like elevated latency. Logs tell you what happened at a specific point in one service. Tracing connects those points across every service a single request touched, which is what lets you find where in a multi-service chain the time or error actually occurred, rather than guessing from separate logs.
Is tail-based sampling worth the added complexity for a small team?
If your tracing storage costs are still small, head-based sampling with a reasonably generous percentage is fine to start with. Tail-based sampling becomes worth the setup effort once your trace volume grows enough that a flat percentage sample is genuinely missing the rare, important requests you'd want to investigate.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.
Related Guides
Instrumenting Tracing Without Drowning in Spans
Turning on auto-instrumentation everywhere produces a bill bigger than the incident it was meant to debug. A rollout order that avoids that.
A Worksheet for Deciding What to Instrument With OpenTelemetry First
A simple worksheet for prioritizing which services get OpenTelemetry instrumentation first, based on incident history and blast radius, not ease of setup.
The OpenTelemetry Rollout Order That Keeps the Trace Bill Sane
A practical runbook for adopting OpenTelemetry tracing across a microservices stack: what to instrument first, sampling strategy, and cost control.
Setting Up OpenTelemetry So Traces Actually Connect
A worksheet for rolling out OpenTelemetry: context propagation across queues, sampling that keeps errors, and controlling cardinality before costs spike.
Setting Up Distributed Tracing Without Drowning in Spans
A practical guide to rolling out OpenTelemetry distributed tracing: what to instrument first, and how to keep trace data useful instead of overwhelming.
Rolling Out OpenTelemetry Without Drowning in Spans
A practical rollout plan for OpenTelemetry distributed tracing, including sampling strategy, span naming, and the mistakes that make traces unusable.