Building a Tracing Convention Your Team Will Actually Follow
The first service usually gets solid tracing coverage: someone spends a focused week instrumenting it, naming spans thoughtfully, and the resulting traces are genuinely useful for debugging. Then a second team instruments their service with a different naming scheme, a third skips attributes the first two considered mandatory, and months later the traces exist everywhere but don't connect into a coherent picture of a request across services.
The fix isn't a better tracing tool. It's a small, written set of conventions decided before the second team starts, so every service's spans are searchable and comparable instead of each one reflecting whoever instrumented it first. This works through building that worksheet.
Why Tracing Rollouts Stall After the First Service
Distributed tracing's whole value proposition is following one request across every service it touches. That only works if every service names its spans consistently enough to correlate, tags them with enough shared attributes to filter across services, and samples at a rate that doesn't silently drop the trace right when you need it during an incident.
Without a written convention, each team instruments based on what made sense to them individually, which is reasonable in isolation and fragmenting in aggregate. The second and third services end up with traces that are individually fine but useless for tracing a request that crosses all three, which is usually the exact case you need tracing for most.
Worksheet: Naming Spans So They're Searchable Later
Write down a naming pattern before anyone instruments a second service. A pattern that includes the operation type and the resource it acts on, a consistent verb-plus-resource format applied the same way everywhere, searches far better later than each engineer choosing a name that made sense to them at the time.
Decide, specifically, what goes in the span name versus what goes in an attribute. The span name should stay low-cardinality and consistent across calls, while a specific ID or parameter value belongs in an attribute, not the name itself, or you'll end up with thousands of unique span names that can't be grouped. Write this down somewhere every engineer instrumenting a new service will actually see it, not just in a document nobody opens again.
Worksheet: Deciding Which Attributes Are Mandatory
Pick a small, mandatory set of attributes every span across every service must include, and keep it genuinely small, five or fewer, so it's realistic to enforce. The usual candidates: a request or correlation ID that ties spans together across services, the service name and version, and the environment. Everything beyond that list is optional and service-specific.
The mistake teams make here is either mandating too many attributes, which nobody actually fills in consistently, or mandating none, which means cross-service queries have nothing reliable to filter on. A short mandatory list that's actually enforced beats a long recommended list that's actually ignored.
Setting a Sampling Rate You Can Afford
Sampling every single request generates a volume of trace data that gets expensive fast and often exceeds what your tracing backend can usefully store or query. Sampling too aggressively means the specific slow or failed request you need to debug during an incident was never captured.
A common middle ground: sample a small baseline percentage of all traffic for general visibility, and always capture traces for requests that errored or exceeded a latency threshold, regardless of the baseline sampling decision. This combination keeps volume manageable day to day while making sure the traces you actually need during an incident are the ones least likely to have been sampled away.
Rolling Out to a Second Service Without Redoing the Work
Once the conventions exist, rolling out to each additional service should be mechanical, not a fresh design exercise:
- Give the new team the written naming and attribute worksheet directly, not a link to the first team's code to reverse-engineer conventions from.
- Have someone who instrumented the first service review the second team's initial spans before they merge, the same way you'd review any other shared convention.
- Confirm a request that crosses both services actually correlates into one connected trace, not two disconnected ones, before calling the rollout done.
- Update the worksheet itself if the second team's real use case surfaces a gap the first service never hit, so the convention improves instead of staying frozen at version one.
A tool comparison like Datadog vs. New Relic vs. Dynatrace matters less than this convention work. Any of those platforms can ingest and display traces well; none of them will fix inconsistent span naming for you.
What Good Looks Like
Good distributed tracing means a request that crosses multiple services correlates into one connected trace with consistent, searchable span names, not a set of disconnected per-service traces.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How many attributes should be mandatory on every span?
Keep it to five or fewer: typically a correlation ID that ties spans together across services, the service name and version, and the environment. A longer mandatory list sounds thorough but rarely gets filled in consistently, which defeats the purpose of making it mandatory at all.
What sampling rate should we start with?
A small baseline percentage of all traffic for general visibility, combined with always capturing traces for anything that errored or exceeded a latency threshold. That combination keeps storage costs manageable while making sure the traces you actually need during an incident weren't sampled away.
Do we need a written convention if we're only instrumenting one service right now?
Write it down anyway, before a second team starts. It's much easier to define naming and attribute conventions once, early, than to retrofit consistency across two or three services that were each instrumented independently by different people with different assumptions.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Datadog vs New Relic vs Dynatrace: Cloud Observability Platforms Compared
Compare Datadog, New Relic, and Dynatrace for cloud observability: log ingestion costs, distributed tracing, APM overhead, and MTTR compression.
Rolling Out Agentic Workflows Without Breaking Production
A practical rollout checklist for shipping an AI agent to production, from a shadow-mode test run through the guardrails that catch it if it misbehaves.
Build vs. Buy for Verifying Every Device That Connects In
What zero-trust device and identity verification actually requires, what a platform gives you over a homegrown check, and how to decide between them.
Choosing a Distributed Lock: Redis, Redlock, Postgres, or etcd
A comparison of single-node Redis locks, Redlock, Postgres advisory locks, and etcd for coordinating work across multiple application instances.
Caching Context So Your Agents Don't Pay for It Twice
Comparing where caching actually helps an agentic system, from prompt caching to tool result caching, and where it introduces stale-data risk instead.
Instrumenting Tracing Without Drowning in Spans
Turning on auto-instrumentation everywhere produces a bill bigger than the incident it was meant to debug. A rollout order that avoids that.