Getting Started With OpenTelemetry on a Small Team
OpenTelemetry is an open standard for collecting traces, metrics and logs from your services and sending them to any compatible backend. A small team should start with traces on one important request path, using auto-instrumentation and a collector, before touching anything else.
This plan gets you to a useful trace in a day and explains the choices that affect cost and lock-in later.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
What are the pieces, in plain terms?
Four ideas cover most of what you need:
- Signals: traces (the path of one request across services), metrics (numbers over time) and logs (event records). Traces are the best starting point because they show where time goes.
- SDK and instrumentation: libraries in your app that create spans. Auto-instrumentation covers common frameworks and clients with little code, and you add manual spans for your own business operations.
- Context propagation: a trace ID passed between services in a header so spans join into one trace. If a service drops the header, the trace breaks.
- Collector: a process that receives telemetry, batches and filters it, and forwards it to a backend. It keeps backend credentials and routing out of your application code.
The standard protocol between them is OTLP, which is why you can switch backends without re-instrumenting.
How to get your first trace in a day
Pick the request path customers care about most, such as sign-in or checkout, and follow these steps:
- Choose one service on that path and enable the language's auto-instrumentation using environment variables where possible.
- Set a clear service name and an environment attribute, since every dashboard and query depends on them.
- Run a collector next to the service, or as a small shared deployment, and point the SDK at it.
- Send data to a backend, even a free tier or a local viewer at first.
- Trigger the request and confirm you see spans for the incoming call, database queries and outgoing HTTP calls.
- Repeat for the next service on the path and verify propagation by checking that both appear in one trace.
Resist instrumenting everything at once. One complete, trustworthy trace beats ten partial ones.
Which backend should you send data to?
Because the format is open, backend choice is reversible. Datadog and New Relic both accept OpenTelemetry data, and you can also use open-source stores you operate yourself. For a small team, the deciding factors are usually who will run it, how you want to search and alert, and what the bill looks like as volume grows. The comparison is in Datadog vs New Relic vs Dynatrace.
Whichever you pick, keep exporter settings in the collector, not scattered through application code. That way, switching or adding a second destination is a config change. Confirm current ingest pricing and limits in a demo, since plans change.
How do you keep the cost under control?
Telemetry volume tends to grow faster than expected. Manage it deliberately:
- Sample traces. Keep all errors and slow requests, and a fraction of normal ones. Tail-based sampling in the collector can make that decision after seeing the whole trace.
- Drop noisy spans such as health checks and static asset requests.
- Watch attribute cardinality. Putting a user ID or full URL with parameters into a metric label can explode series counts. Keep those in traces, not in metric dimensions.
- Set retention to what you use. Few people query traces older than a couple of weeks.
- Set a monthly volume alert with your backend so a bad deploy doesn't produce a surprise bill.
Also keep sensitive data out of telemetry. Don't record request bodies, tokens or personal data in span attributes unless you've reviewed the privacy implications with your counsel.
What should you do after the first trace works?
Turn the data into decisions. Build one dashboard per critical path with request rate, error rate and latency percentiles, and add a deploy marker so changes are visible against the graphs. Add alerts tied to customer symptoms and link them to your paging setup; tie them to service level objectives if you have them. Then extend to metrics and logs, using trace IDs in log lines to jump between them.
Keep an eye on delivery health too. DORA's 2024 report puts change failure rate at 5 percent in its highest-performing cluster and 40 percent in its lowest1, and traces make it faster to see whether a release caused a regression. Related reading includes distributed tracing with OpenTelemetry for DevSecOps and the developer platform version.
What Good Looks Like
Every service on your critical request paths emits connected traces with consistent names, and telemetry cost is monitored and sampled.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
What is OpenTelemetry?
It's an open standard, with SDKs and a collector, for producing and shipping traces, metrics and logs. Because the format is vendor-neutral, you can change the backend without rewriting your instrumentation.
Do I need a collector?
Not strictly, since SDKs can export directly. But a collector lets you batch, filter, sample and route data, and keeps backend credentials out of application code, so most teams add one early.
Where should a small team start?
With traces on one important request path, using auto-instrumentation. Confirm the trace spans every service on that path, then expand to more services, metrics and logs.
How do I avoid vendor lock-in?
Instrument with OpenTelemetry, use OTLP, and keep exporter configuration in a collector. Then switching backends is a configuration change rather than a code rewrite.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Change failure rate by DORA performance cluster. DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.
Related Guides
Datadog vs New Relic vs Dynatrace: Cloud Observability Platforms Compared
Compare Datadog, New Relic, and Dynatrace for cloud observability: log ingestion costs, distributed tracing, APM overhead, and MTTR compression.
The OpenTelemetry Rollout Order That Keeps the Trace Bill Sane
A practical runbook for adopting OpenTelemetry tracing across a microservices stack: what to instrument first, sampling strategy, and cost control.
Instrumenting Tracing Without Drowning in Spans
Turning on auto-instrumentation everywhere produces a bill bigger than the incident it was meant to debug. A rollout order that avoids that.
Getting Real Value Out of OpenTelemetry Instead of Just Installing It
Why instrumenting every service with OpenTelemetry isn't the same as being able to debug a real production issue, and what to fix first to close that gap.
Rolling Out OpenTelemetry Without Drowning Your Team in Spans
A practical rollout sequence for OpenTelemetry distributed tracing across a real time pipeline, including where to instrument first and how to control cost.
Setting Up OpenTelemetry So Traces Actually Connect
A worksheet for rolling out OpenTelemetry: context propagation across queues, sampling that keeps errors, and controlling cardinality before costs spike.