Engineering Leadership & Technical HiringPlaybook3 min readUpdated September 2026

Setting Up OpenTelemetry So Traces Actually Connect

The most common OpenTelemetry rollout failure isn't missing instrumentation, it's broken context propagation: a trace that looks complete inside one service and silently starts over the moment a request crosses a message queue or an async job boundary. The result is a set of disconnected trace fragments that don't answer the question tracing was supposed to answer: where did this specific slow request actually spend its time.

Work through this as a worksheet against your own system, one section at a time.

Worksheet Step One: Map Where Context Actually Crosses a Boundary

List every place a request's trace context needs to survive a hop it might not survive automatically: HTTP calls between services usually propagate context fine with standard OpenTelemetry instrumentation, but a message published to a queue, a background job enqueued for later, or a call into a different runtime often drops it unless you explicitly inject the trace ID into the message and extract it on the other side. Go through your architecture diagram and mark every queue, cron job, and cross-runtime call as a place to verify, not assume, propagation works.

Worksheet Step Two: Auto-Instrumentation vs. Manual Spans

Auto-instrumentation for common frameworks and libraries, your web framework, your database driver, your HTTP client, gets you a usable baseline with minimal code changes and should be your starting point everywhere it's available. Manual spans are worth the extra work only around business logic that's actually hard to reason about from the automatic spans alone, a complex calculation, a multi-step workflow inside one function. Instrumenting everything manually from the start is a common overcorrection that produces more maintenance burden than insight.

Worksheet Step Three: Pick a Sampling Strategy That Keeps What Matters

Head-based sampling, deciding whether to keep a trace at the moment it starts, is simple but risks dropping exactly the slow or erroring traces you most want to see, since the decision happens before you know the outcome. Tail-based sampling, deciding after the trace completes, lets you keep every error and every slow trace while sampling down the high-volume, uninteresting, fast successful ones, but requires a collector that can buffer and evaluate complete traces before deciding, which is a heavier infrastructure lift. Most teams past early scale need tail-based sampling to keep tracing both useful and affordable.

Worksheet Step Four: Correlate Traces With Logs and Metrics, Not Just Each Other

A trace ID that also appears in your structured logs for the same request turns "here's a slow span" into "here's the slow span, and here's the exact log lines from that request, including the ones from a service that isn't even instrumented yet." This correlation is usually a small logging configuration change, injecting the active trace ID into your logger's context, but it's the step that makes tracing genuinely useful for debugging instead of just a pretty flame graph nobody references during an actual incident.

Worksheet Step Five: Control Cardinality Before It Controls Your Bill

Attaching high-cardinality values, a full request body, a raw user ID, an unbounded free-text field, as span attributes on every request multiplies your tracing backend's storage and cost far faster than most teams expect, and it's the single most common reason a team turns tracing off a few months after turning it on. Set explicit conventions for what's allowed as a span attribute before instrumentation spreads across your codebase, and review new high-cardinality attributes in code review the same way you'd review a new database index.

Worksheet Step Six: Decide Who Owns the Collector Pipeline

An OpenTelemetry Collector sitting between your services and your tracing backend gives you a single place to apply sampling rules, redact sensitive attributes before they leave your infrastructure, and switch backends later without touching application code. Someone needs to own its configuration and its own uptime, since a collector that falls behind or crashes silently drops trace data with no obvious symptom beyond gaps that are easy to miss until you go looking for a specific trace that isn't there.

What a Working Setup Actually Feels Like Day to Day

Once propagation, sampling, and correlation are all working together, the real test is whether an engineer debugging a slow request reaches for the trace first, instead of grepping through logs across five services by hand. If your team still defaults to log grepping during an incident months after rolling out tracing, that's a signal something in the setup, usually propagation gaps or missing log correlation, is quietly undermining the tool's usefulness even though the dashboards technically show data flowing.

Here is the whole worksheet as a checklist:

  1. Map every boundary where trace context may not survive on its own, such as message queues and background jobs.
  2. Start with auto-instrumentation, then add manual spans where the baseline misses what matters.
  3. Choose a sampling strategy that keeps slow and erroring traces, using tail-based sampling or an error-aware override.
  4. Put the trace ID in your structured logs so a slow span links to that request's log lines.
  5. Keep high-cardinality values, such as raw user IDs or full request bodies, out of span attributes.
  6. Decide who owns the Collector pipeline for sampling, redaction, and backend changes.
Executive Capability Standard

What Good Looks Like

Every request that crosses a service boundary should carry one trace ID from the edge through to the database and back.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Pick one slow endpoint and manually trace its calls through your existing logs before instrumenting anything.
2. Do Manually:Read raw span data for a single request before trusting a flame graph's summarized view of it.
3. Delegate:Give one team ownership of the shared OpenTelemetry SDK configuration so instrumentation doesn't drift service to service.
4. Automate:Sample intelligently: keep every error and slow trace, sample down the rest, so cost doesn't force tracing off later.
5. Buy:Consider a managed tracing backend with tail-based sampling built in once running your own collector infrastructure outweighs the savings.

How to Get Started

Frequently Asked Questions

Do we need a dedicated observability team to run OpenTelemetry well?

Not to start. One engineer owning the shared SDK configuration and instrumentation conventions is enough for most teams initially; a dedicated observability function becomes worth it once you're running tail-based sampling infrastructure and cardinality management across many services.

Why does our trace stop at the message queue instead of continuing into the consumer?

Context isn't being propagated through the message itself. You need to explicitly inject the trace context into the message's headers or metadata when publishing, and extract it when the consumer picks the message up; this doesn't happen automatically the way HTTP propagation usually does.

How do we decide what to keep when tail-based sampling isn't feasible yet?

Start with a fixed percentage sample plus a rule that always keeps traces containing an error status, even with head-based sampling most collectors support an error-aware override. It's a rougher approximation of tail-based sampling's benefits but far better than pure random sampling alone.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides