Model Context Protocol & Agentic ArchitecturePlaybook3 min readUpdated September 2026

Building a Tracing Convention Your Team Will Actually Follow

The first service usually gets solid tracing coverage: someone spends a focused week instrumenting it, naming spans thoughtfully, and the resulting traces are genuinely useful for debugging. Then a second team instruments their service with a different naming scheme, a third skips attributes the first two considered mandatory, and months later the traces exist everywhere but don't connect into a coherent picture of a request across services.

The fix isn't a better tracing tool. It's a small, written set of conventions decided before the second team starts, so every service's spans are searchable and comparable instead of each one reflecting whoever instrumented it first. This works through building that worksheet.

Why Tracing Rollouts Stall After the First Service

Distributed tracing's whole value proposition is following one request across every service it touches. That only works if every service names its spans consistently enough to correlate, tags them with enough shared attributes to filter across services, and samples at a rate that doesn't silently drop the trace right when you need it during an incident.

Without a written convention, each team instruments based on what made sense to them individually, which is reasonable in isolation and fragmenting in aggregate. The second and third services end up with traces that are individually fine but useless for tracing a request that crosses all three, which is usually the exact case you need tracing for most.

Worksheet: Naming Spans So They're Searchable Later

Write down a naming pattern before anyone instruments a second service. A pattern that includes the operation type and the resource it acts on, a consistent verb-plus-resource format applied the same way everywhere, searches far better later than each engineer choosing a name that made sense to them at the time.

Decide, specifically, what goes in the span name versus what goes in an attribute. The span name should stay low-cardinality and consistent across calls, while a specific ID or parameter value belongs in an attribute, not the name itself, or you'll end up with thousands of unique span names that can't be grouped. Write this down somewhere every engineer instrumenting a new service will actually see it, not just in a document nobody opens again.

Worksheet: Deciding Which Attributes Are Mandatory

Pick a small, mandatory set of attributes every span across every service must include, and keep it genuinely small, five or fewer, so it's realistic to enforce. The usual candidates: a request or correlation ID that ties spans together across services, the service name and version, and the environment. Everything beyond that list is optional and service-specific.

The mistake teams make here is either mandating too many attributes, which nobody actually fills in consistently, or mandating none, which means cross-service queries have nothing reliable to filter on. A short mandatory list that's actually enforced beats a long recommended list that's actually ignored.

Setting a Sampling Rate You Can Afford

Sampling every single request generates a volume of trace data that gets expensive fast and often exceeds what your tracing backend can usefully store or query. Sampling too aggressively means the specific slow or failed request you need to debug during an incident was never captured.

A common middle ground: sample a small baseline percentage of all traffic for general visibility, and always capture traces for requests that errored or exceeded a latency threshold, regardless of the baseline sampling decision. This combination keeps volume manageable day to day while making sure the traces you actually need during an incident are the ones least likely to have been sampled away.

Rolling Out to a Second Service Without Redoing the Work

Once the conventions exist, rolling out to each additional service should be mechanical, not a fresh design exercise:

  • Give the new team the written naming and attribute worksheet directly, not a link to the first team's code to reverse-engineer conventions from.
  • Have someone who instrumented the first service review the second team's initial spans before they merge, the same way you'd review any other shared convention.
  • Confirm a request that crosses both services actually correlates into one connected trace, not two disconnected ones, before calling the rollout done.
  • Update the worksheet itself if the second team's real use case surfaces a gap the first service never hit, so the convention improves instead of staying frozen at version one.

A tool comparison like Datadog vs. New Relic vs. Dynatrace matters less than this convention work. Any of those platforms can ingest and display traces well; none of them will fix inconsistent span naming for you.

Executive Capability Standard

What Good Looks Like

Good distributed tracing means a request that crosses multiple services correlates into one connected trace with consistent, searchable span names, not a set of disconnected per-service traces.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Review how the first instrumented service names its spans and tags its attributes, and write that pattern down before a second team starts.
2. Do Manually:Manually check that a request crossing two services actually correlates into one trace before considering either service's instrumentation complete.
3. Delegate:Assign one engineer to own the tracing convention document and review new services' initial instrumentation against it.
4. Automate:Add a lint or CI check that flags span names or missing mandatory attributes that don't match the convention before code merges.
5. Buy:Bring in outside observability expertise to design the initial convention if no one on the team has run a multi-service tracing rollout before, since fixing an inconsistent rollout later costs far more than getting it right the first time.

How to Get Started

Frequently Asked Questions

How many attributes should be mandatory on every span?

Keep it to five or fewer: typically a correlation ID that ties spans together across services, the service name and version, and the environment. A longer mandatory list sounds thorough but rarely gets filled in consistently, which defeats the purpose of making it mandatory at all.

What sampling rate should we start with?

A small baseline percentage of all traffic for general visibility, combined with always capturing traces for anything that errored or exceeded a latency threshold. That combination keeps storage costs manageable while making sure the traces you actually need during an incident weren't sampled away.

Do we need a written convention if we're only instrumenting one service right now?

Write it down anyway, before a second team starts. It's much easier to define naming and attribute conventions once, early, than to retrofit consistency across two or three services that were each instrumented independently by different people with different assumptions.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides