Cloud FinOps & Infrastructure ScalingPlaybook3 min readUpdated September 2026

Getting Real Value Out of OpenTelemetry Instead of Just Installing It

A lot of teams install OpenTelemetry, wire up the SDK across their services, and consider distributed tracing done. Then a real incident happens, a request that spans six services and takes eight seconds when it should take two hundred milliseconds, and the trace they pull up is a wall of spans with no clear story about where the time actually went.

Instrumentation and usefulness are two different milestones, and most of the value sits in the gap between them.

Why Installing the SDK Isn't the Hard Part

OpenTelemetry's auto-instrumentation libraries make it genuinely easy to get spans flowing from most common frameworks and libraries without writing custom code. That's real progress over the manual instrumentation era, but it also means teams stop there, with traces that show HTTP calls and database queries but nothing about what actually happened inside your own business logic. The trace shows you made a call to another service and it took three seconds; it doesn't show you why, unless you've added spans around the specific logic that matters.

Instrument the Decisions, Not Just the Calls

The traces that actually help during an incident have custom spans around the parts of your code where a decision gets made or where variable-cost work happens: a retry loop, a cache lookup that might miss, a batch operation whose size depends on the input. Auto-instrumentation gives you the network boundary for free; it can't know which internal function is your team's actual bottleneck. Pick the handful of services most often involved in past incidents and add targeted spans there first, rather than trying to instrument everything evenly.

Propagate Context Correctly, or the Trace Breaks at the Worst Moment

A distributed trace only stays connected across service boundaries if trace context gets propagated on every call, HTTP headers, queue message metadata, background job payloads. It's common for teams to get this right for synchronous HTTP calls and miss it for asynchronous work: a message published to a queue that doesn't carry the trace ID forward breaks the chain right at the point where debugging usually gets hardest, an async handoff. Audit every place your architecture crosses a synchronous-to-asynchronous boundary and confirm trace context survives it.

Sampling: The Setting Most Teams Get Wrong by Default

Tracing every single request at high fidelity gets expensive fast, so most setups sample, keeping only a percentage of traces. The problem is that a flat percentage sample, say ten percent, is likely to miss the rare, slow, or erroring requests that matter most for debugging, precisely because they're rare. Tail-based sampling, which decides whether to keep a trace after seeing how it completed, keeping all the slow or erroring ones and a smaller sample of the normal ones, gets you a much more useful dataset for the same storage cost. It's more complex to set up than head-based sampling, but it's usually worth the tradeoff once tracing volume becomes a real cost line.

Making Traces Useful During an Actual Incident, Not Just in a Postmortem

A trace that's only ever pulled up after the fact, during a postmortem, is missing half its value. Link your tracing tool directly from your alerting and dashboards, so when a latency alert fires, the on-call engineer is one click from a relevant trace instead of starting a manual search through a tracing UI under pressure. Teams that ship every day especially benefit from this, since a fast deploy cadence means more candidate changes to check against when something goes wrong, and a trace that's hard to find slows that check down exactly when speed matters1.

Checks that turn installed tracing into usable tracing:

  • Add custom spans around retry loops, cache lookups that might miss, and batch operations whose size depends on the input.
  • Propagate trace context on every hop, including queue message metadata and background job payloads, not only synchronous HTTP headers.
  • Consider tail-based sampling so rare slow or erroring requests are kept, since a flat percentage sample tends to miss exactly those.
  • Link alerts and dashboards directly to a relevant trace so the on-call engineer is one click away from it during an incident.

A Worked Example: Chasing an Eight-Second Request

Say a checkout request that should take two hundred milliseconds is taking eight seconds for a subset of users, and the trace shows five spans: the API gateway, an auth check, an inventory service call, a pricing service call, and the database write, with most of the time sitting inside the pricing service span and nothing more specific than that. Without a custom span inside the pricing service around its actual logic, discount lookup, tax calculation, currency conversion, you're left guessing which of those three did the damage. Add spans around each one, and the next occurrence points straight at the actual bottleneck instead of just the service boundary that contains it.

Executive Capability Standard

What Good Looks Like

Useful tracing means the services most often involved in past incidents have custom spans around their actual decision points, trace context survives every async boundary, and an on-call engineer can reach a relevant trace in one click from an alert.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Review your last three incidents and check whether existing traces would have shown you where the time went, or just that a call happened.
2. Do Manually:Add custom spans manually to the one or two services most often implicated in past incidents.
3. Delegate:Assign a platform or backend engineer to audit trace context propagation across every async boundary in your architecture.
4. Automate:Link tracing directly into alerting and dashboards so an on-call engineer reaches a relevant trace automatically instead of searching manually.
5. Buy:Adopt a managed observability platform with built-in tail-based sampling if your trace volume has outgrown what a flat percentage sample can usefully cover.

How to Get Started

Frequently Asked Questions

Do we need to instrument every service before tracing is useful at all?

No. Partial coverage still helps, especially if you start with the services most often implicated in past incidents. A trace that goes dark at an uninstrumented service still tells you where to look next, even if it doesn't tell you the full story yet. Expand coverage service by service rather than waiting to instrument everything before getting value.

What's the difference between logging, metrics, and tracing?

Metrics tell you something is wrong in aggregate, like elevated latency. Logs tell you what happened at a specific point in one service. Tracing connects those points across every service a single request touched, which is what lets you find where in a multi-service chain the time or error actually occurred, rather than guessing from separate logs.

Is tail-based sampling worth the added complexity for a small team?

If your tracing storage costs are still small, head-based sampling with a reasonably generous percentage is fine to start with. Tail-based sampling becomes worth the setup effort once your trace volume grows enough that a flat percentage sample is genuinely missing the rare, important requests you'd want to investigate.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.

Related Guides