The OpenTelemetry Rollout Order That Keeps the Trace Bill Sane
OpenTelemetry solves a real problem, tracing a single request as it crosses a dozen services so you can find where the latency actually lives, but a naive rollout tends to produce either too little useful data or a trace storage bill nobody approved. The gap between those two outcomes is almost entirely about sampling strategy and instrumentation order.
This is a runbook for adopting OpenTelemetry across a microservices stack without either problem: what to instrument first, how to sample sensibly, and where the cost actually comes from.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Instrument the request path first, not everything at once
The instinct to instrument every service simultaneously produces a flood of trace data before anyone has built the dashboards or alerts to use it, and the rollout stalls under its own volume. Start with the services on your single most important request path, the one most likely to be the target of the next incident review, get end-to-end traces flowing and useful there, then expand service by service. A working trace across three services beats a broken, half-instrumented trace across thirty.
Sampling strategy is the single biggest cost lever
Head-based sampling, deciding whether to keep a trace at the very first service, is simple but throws away traces before you know if they were interesting, meaning a rare, expensive slow request has the same chance of being dropped as a fast, boring one. Tail-based sampling, deciding after the full trace completes, lets you deliberately keep every error trace and every genuinely slow trace while sampling normal traffic down aggressively, which is almost always the right tradeoff, but it requires buffering full traces somewhere before the sampling decision, which has its own memory and infrastructure cost to plan for.
Propagating context correctly across service and language boundaries
A trace breaks, silently becoming several disconnected traces instead of one, whenever context propagation fails: a message queue that doesn't carry trace headers, an async job that starts a new context instead of inheriting the parent's, or a language boundary where two different OpenTelemetry SDKs handle propagation slightly differently. Test context propagation explicitly across every async boundary and message queue in your system rather than assuming it works because it works for synchronous HTTP calls, which is usually the only path anyone actually tests.
What actually drives the trace storage bill
Cardinality is the usual culprit: attaching a high-cardinality attribute, a full user ID or a raw request body, to every span multiplies your storage cost in a way that's hard to predict from a dashboard until the invoice arrives. Set clear guidelines for what belongs as a span attribute versus what belongs in a linked log, and review cardinality on your highest-volume services specifically, since that's where an unbounded attribute does the most damage.
A worked example: the attribute that quietly multiplied the bill
Picture a team that adds the raw request body as a span attribute on their busiest endpoint to help debug a tricky class of bug, ships it, and moves on. Weeks later the trace storage bill has grown well past what traffic growth alone would explain, and nobody connects it to that one debugging attribute because it looked harmless in a local test with a handful of requests. At production volume, a single high-cardinality attribute on a high-traffic span multiplies stored data far faster than intuition suggests, which is exactly why cardinality review belongs in the pull request process, not in a retrospective after the invoice arrives.
Making the data actually useful to on-call engineers
Tracing infrastructure with no dashboards, alerts, or runbook links tied to it becomes a tool three people know how to use and nobody else touches during an incident. Build a small number of high-value views, the request-path latency breakdown for your top three services, an error-rate-by-service view, before expanding to more, and link directly to a trace search filtered to the relevant service from your existing alerting tool so an on-call engineer doesn't need to learn a new interface mid-incident. Whichever observability backend you land on for storing and querying traces, that link-from-alert step is the one teams skip most often.
Roll out tracing in this order:
- Pick the single most important request path and instrument only the services on it, until end-to-end traces work there.
- Build a small set of dashboards, such as a request-path latency breakdown and an error-rate-by-service view, before adding more services.
- Test context propagation across every queue, async job, and language boundary so traces don't split into disconnected fragments.
- Move from head-based to tail-based sampling so every error trace and unusually slow trace is kept while routine traffic is sampled down.
- Set span attribute guidelines that keep high-cardinality values, like full user IDs and raw request bodies, out of spans, and review your busiest services regularly.
What Good Looks Like
A working tracing rollout instruments one critical request path fully before expanding, uses tail-based sampling to keep error and slow traces deliberately, and reviews attribute cardinality on the highest-volume services to keep storage cost predictable.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
Should we start with head-based or tail-based sampling?
Head-based is simpler to stand up and reasonable for an initial rollout on a low-stakes service. Move to tail-based sampling once you're instrumenting anything where catching every error and slow trace matters, since head-based sampling can miss the exact traces you'd most want to investigate.
How many services should we instrument in the first rollout phase?
Start with the services on a single critical request path, often three to six services, rather than the whole system. Get that path genuinely useful with real dashboards before expanding, since a wide but shallow rollout tends to stall without producing anything anyone actually uses.
What's the most common reason a trace shows gaps between services?
Broken context propagation across an async boundary, most often a message queue or background job that doesn't carry trace headers forward. Test this explicitly for every queue and async job in the system rather than assuming it works because synchronous HTTP calls do.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Datadog vs New Relic vs Dynatrace: Cloud Observability Platforms Compared
Compare Datadog, New Relic, and Dynatrace for cloud observability: log ingestion costs, distributed tracing, APM overhead, and MTTR compression.
Instrumenting Tracing Without Drowning in Spans
Turning on auto-instrumentation everywhere produces a bill bigger than the incident it was meant to debug. A rollout order that avoids that.
A Worksheet for Deciding What to Instrument With OpenTelemetry First
A simple worksheet for prioritizing which services get OpenTelemetry instrumentation first, based on incident history and blast radius, not ease of setup.
Setting Up OpenTelemetry So Traces Actually Connect
A worksheet for rolling out OpenTelemetry: context propagation across queues, sampling that keeps errors, and controlling cardinality before costs spike.
Getting Real Value Out of OpenTelemetry Instead of Just Installing It
Why instrumenting every service with OpenTelemetry isn't the same as being able to debug a real production issue, and what to fix first to close that gap.
Setting Up Distributed Tracing Without Drowning in Spans
A practical guide to rolling out OpenTelemetry distributed tracing: what to instrument first, and how to keep trace data useful instead of overwhelming.