Rolling Out OpenTelemetry Without Drowning in Spans
Distributed tracing is supposed to answer one question fast: which of the dozen services a slow request touched is actually the one that's slow. OpenTelemetry gives you a vendor-neutral way to instrument for that, but the rollout is where most teams either succeed quickly or generate a mountain of spans nobody can navigate.
The technology is the easy part. Sampling strategy and span naming are what determine whether the traces you're generating are actually useful six months in.
Where should you start instrumenting with OpenTelemetry?
Instrument service boundaries first, every inbound and outbound call between services, before instrumenting inside a single service's internal logic. This gets you the request-level trace, which service called which and how long each hop took, which is the view that actually answers why a request was slow for most incidents.
Deep internal instrumentation inside a single service is worth adding later, specifically for the services that show up as the bottleneck in that request-level view, not everywhere up front.
How should you sample traces to keep tracing affordable?
Tracing every single request at high traffic volume gets expensive fast, both in storage and in the overhead added to every request. Head-based sampling, deciding whether to trace a request before it starts, is simple but risks dropping exactly the slow or erroring requests you most want visibility into.
Tail-based sampling, deciding after the request completes based on whether it was slow or errored, keeps the traces that actually matter at the cost of buffering every request's spans until a sampling decision can be made. Most production setups land on tail-based sampling for anything above a modest traffic volume, specifically because it biases toward keeping the traces you'd actually want to look at.
Span naming conventions that don't collapse into noise
A span named after a dynamic value, like an individual user ID or order ID baked into the span name itself, makes every trace look unique to your tracing backend and destroys any ability to aggregate or search across similar requests.
Name spans after the operation, not the instance, a create-order call, not a specific order's path. Put the dynamic values in span attributes instead, where they're still searchable but don't fragment your span cardinality into an unnavigable mess.
Common rollout mistakes
- Instrumenting everything at once instead of starting with the two or three services in your slowest, most-reported-on request path.
- No context propagation across an async boundary, like a message queue, which silently breaks the trace into two disconnected pieces at exactly the point where debugging usually gets hardest.
- Treating trace data as a debugging tool only, instead of also using span duration data to catch a slow creep in latency before it becomes an incident.
- Skipping sampling strategy entirely until a storage bill or performance overhead forces the conversation, instead of deciding on one before rollout.
A worked example: finding the actual slow hop
Say a checkout request takes three seconds end to end, and without tracing, the working theory rotates between the payment provider, the database, and the recommendations service, none of it based on real evidence. A request-level trace across the checkout path shows the actual breakdown: 200 milliseconds in the payment call, 150 in the database write, and 2.6 seconds in a synchronous call to the recommendations service that the checkout flow never actually needed a response from.
That's the entire value proposition of tracing in one example: turning a guessing exercise across three plausible suspects into a direct look at where the time actually went, without adding logging to any of the three services individually first.
Cost control beyond sampling rate
Sampling rate isn't the only lever on tracing cost. Span attribute cardinality, how many distinct values a given attribute can take, drives storage cost independently of how many traces you keep, since a high-cardinality attribute like a raw request body fragment attached to every span multiplies storage far faster than the trace volume alone would suggest. Review which attributes are actually queried during an investigation versus which were added speculatively, and drop the ones nobody's used in the last few months of incidents.
What Good Looks Like
A good OpenTelemetry rollout instruments service boundaries first, picks a sampling strategy deliberately instead of by accident, and names spans by operation so traces stay searchable as volume grows.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Should we trace every request or sample?
Sample, almost certainly, once you're past a small amount of traffic; tracing every request gets expensive in both storage and per-request overhead. Tail-based sampling, which decides after a request completes based on whether it was slow or errored, is usually worth the added buffering complexity because it keeps the traces you'd actually want to look at.
Why do our traces look fragmented across an async boundary?
The most common cause is missing context propagation across the boundary, commonly a message queue, where the trace context isn't passed along with the message. Without it, OpenTelemetry has no way to connect the span before the queue to the span after it, and you get two disconnected traces instead of one.
How should we name spans so they're actually searchable later?
Name spans after the operation, like a create-order call, not after a dynamic value like a specific order ID baked into the name. Put dynamic values in span attributes instead, where they stay searchable without fragmenting your span cardinality into thousands of effectively unique, unaggregatable span names.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Continuous Device Verification for a Zero-Trust API
How continuous device and identity verification actually works in a zero-trust architecture, and where to draw the line for a small engineering team.
Rolling Out Zero Trust in Production Without a Broad Outage
A checklist for rolling out stricter API authentication and authorization in production, and the pitfalls that turn a rollout into an incident.
Instrumenting Tracing Without Drowning in Spans
Turning on auto-instrumentation everywhere produces a bill bigger than the incident it was meant to debug. A rollout order that avoids that.
Distributed Locking With Redis: Where Redlock Actually Falls Short
A practical guide to distributed locks with Redis, including where the Redlock algorithm's guarantees break down and when to use a database lock instead.
Where Caching Helps a Zero Trust API and Where It Creates Risk
Comparing where caching genuinely speeds up a zero trust API against where it creates a real revocation and permission risk.
How to Audit Whether Your APIs Actually Enforce Zero Trust
A step-by-step method for testing whether your APIs enforce zero trust in practice, not just on paper, and what to do with what you find.