Cloud Observability & APM Platforms3 min readUpdated September 2026

Datadog vs New Relic for AI Agent Pipelines

Inference jobs fail in ways a CPU graph does not explain: a slow model provider, a truncated response, a queue backing up behind a retry loop. Deciding between Datadog and New Relic for an AI or workflow automation agency comes down to custom-metrics cost, how well each traces calls across third-party model APIs, and how each handles bursts of short-lived worker processes.

Token and latency telemetry is custom instrumentation either way. Neither platform tracks it for you out of the box.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Why a Stuck Agent Run Looks Fine on a CPU Graph

A multi-step agent pipeline can sit idle waiting on a slow model response, retry a failed tool call three times, or loop on a malformed output, and none of that shows up as elevated CPU or memory. The signal that actually matters, time spent waiting on each model call and how often a step needs a retry, has to be instrumented explicitly by wrapping every provider call with a span and a small set of custom attributes. Datadog's APM makes it straightforward to attach custom span tags for token counts and provider name; New Relic supports the same pattern through its OpenTelemetry integration, which is a better fit if your agency already standardized on OpenTelemetry to stay portable across clients.

Without that instrumentation, a client asking why their automation is slow today gets an answer built from guesswork instead of a trace that shows exactly which step in the pipeline is waiting on what.

Tracing a Pipeline Across Two or Three Model Providers

An agency running the same workflow against more than one model provider, for cost or reliability reasons, needs a trace that follows a single job across provider boundaries without losing context when a call gets retried on a fallback provider. Datadog's distributed tracing handles that cleanly once the retry logic is instrumented to propagate the original trace ID. New Relic's tracing does the same, but a fallback-provider retry that spins up in a separate worker process more often needs a small amount of manual trace-context passing to avoid showing up as an unrelated, disconnected span.

Instrument an agent pipeline in this order:

  1. Wrap every model provider call in a span so the time spent waiting on each call shows up in your traces.
  2. Attach custom attributes to each span, such as provider name, token counts, and retry count, because neither platform records them automatically.
  3. Propagate the original trace ID when a call retries on a fallback provider, so one job stays a single connected trace.
  4. Track latency and error rate for each provider separately, and run a synthetic check against every provider on a schedule.

Bursty Workers and Why Host-Based Pricing Gets Painful

A workflow automation agency's compute pattern is often bursty by design: a client's scheduled job spins up fifty short-lived workers for ten minutes, then scales back to zero. Datadog's per-host pricing model was built for steadier workloads, and a bursty fleet of ephemeral workers can generate a monitoring bill that swings around unpredictably from one billing period to the next. New Relic's usage-based pricing tracks the actual volume of data you send rather than the number of hosts running at any given moment, which tends to match a bursty compute pattern more closely.

Ask either vendor how their pricing treats a worker that lives for ten minutes versus one that runs all month, since the answer changes which model fits an agency whose infrastructure footprint looks completely different from one hour to the next.

Recovering Fast When a Model Provider Degrades

DORA's research reports a recovery time under an hour for the fastest teams and up to a month for the slowest1, and an agency whose pipelines depend on a third-party model provider needs a similar recovery discipline for a dependency it does not control. A synthetic check that calls each provider on a schedule and alerts on elevated latency, not just a hard failure, catches a degrading provider before every client's pipeline backs up behind it at once.

New Relic's synthetic monitors and Datadog's are functionally similar here; the difference is mostly in how quickly your team can write the check in the query language it already knows, and whether the on-call rotation trusts the alert enough to act on it at two in the morning.

Matching the Platform to a Small, Fast-Moving Team

Teams in DORA's highest-performing cluster keep their change failure rate near 5%, well below the roughly 40% seen in the lowest-performing cluster2, and an agency shipping new agent workflows every week benefits from staying closer to that better end of the range, since a bad prompt change or a broken tool integration is easy to ship quickly and easy to miss without deploy markers tied to specific workflow versions. Whichever platform you pick, tag every deploy with the workflow version it shipped, so a spike in retries or token cost can be traced back to the exact prompt or pipeline change that caused it, ideally within minutes of the change going live rather than after a client reports it.

Executive Capability Standard

What Good Looks Like

An agency that has this under control can trace any agent run across every model provider it calls, tags each deploy with the workflow version that shipped it, and keeps its change failure rate near the better end of the DORA range2 even while shipping new workflows weekly.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Learn which parts of an agent pipeline, model latency, retry counts, tool-call failures, are invisible to standard infrastructure metrics and need custom instrumentation.
2. Do Manually:Manually wrap your highest-traffic workflow's provider calls with spans and cost attributes first, before trying to instrument every pipeline at once.
3. Delegate:Delegate ownership of the instrumentation standard, span names, tag conventions, to one engineer, so every new workflow follows the same pattern.
4. Automate:Automate synthetic checks against every model provider you depend on, and alert on rising latency, not only on a hard failure.
5. Buy:Buy a platform-wide plan once you are running enough concurrent client workflows that manual log review across pipelines is no longer realistic.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Frequently Asked Questions

Does Datadog or New Relic track LLM token usage automatically?

No. Both require you to instrument token counts and cost yourself, usually by wrapping each provider call with a span and attaching token and cost as custom attributes. Some community integrations exist for popular model SDKs, but treat token and cost tracking as custom instrumentation your team owns, not a feature that comes free with either platform.

How do we avoid a huge bill from a bursty worker fleet?

Usage-based pricing generally tracks a bursty compute pattern more predictably than per-host pricing, since it scales with data volume rather than the number of hosts running at any moment. Either way, review your retention settings for short-lived workers specifically, since keeping high-resolution metrics for workers that live ten minutes rarely earns its cost.

What should we monitor if we use more than one model provider?

Track latency and error rate per provider separately, not blended together, and set up a synthetic check that calls each provider on a schedule so you catch a degrading provider before a real job fails. A blended metric across providers hides which one is actually causing trouble.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Failed deployment recovery time by DORA performance cluster (upper bound, days). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.
  2. Change failure rate by DORA performance cluster. DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.

Related Guides