ObservabilityPlaybook3 min readUpdated September 2026

Datadog Bill Too High? Where the Money Goes and How to Cut It

To cut a Datadog bill, open the usage breakdown, find the two or three products that dominate it, and reduce what you send: fewer indexed logs, fewer high-cardinality custom metrics and fewer idle hosts or tests. Do this before you negotiate or switch vendors.

Monitoring platforms generally meter several different things at once, so a bill can grow without any single decision. The steps below are ordered by how often they produce the largest savings for a small engineering team.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Which line items usually drive the bill?

Start with the usage and cost view in your account and sort by spend. Products are metered differently, so read the current metering definitions on your own invoice rather than assuming. Typical culprits, roughly in order of how often they dominate:

  • Log volume: ingestion and, separately, indexing and retention. Debug output and health-check noise are the classic offenders.
  • Custom metrics: billed by unique combinations of metric name and tag values. One tag holding a user ID or request ID can multiply a single metric into thousands.
  • Infrastructure hosts and containers: ephemeral hosts, forgotten staging machines and agents installed on boxes nobody looks at.
  • APM and traces: span volume and retention, often left at defaults.
  • Synthetic tests: frequency multiplied by locations multiplied by test count.

Write down the top three with their monthly amounts. That list is your work plan, and it stops you from optimizing a line that is, for example, 3% of the bill.

How do you cut log costs without going blind?

Separate what you need to search from what you need to keep. The steps:

  1. Find the noisiest services and log statuses. Say one service emits 60% of your volume and most of it is INFO lines from a health check.
  2. Drop known-useless lines at the source or with an exclusion filter before indexing, such as health checks, load balancer probes and verbose library output.
  3. Index errors and warnings at full fidelity, and sample routine INFO lines.
  4. Send everything else to cheap archive storage instead of the searchable index, so you can rehydrate it during an investigation.
  5. Shorten retention for the searchable index to what you actually query, often days rather than weeks.

The logging strategy guide covers what to log in the first place, which is cheaper than filtering it later.

How do you fix custom metric and tag sprawl?

Custom metrics often grow silently. Audit them the same way you'd audit a database index list:

  • List custom metrics by volume and check when each was last used in a dashboard, monitor or query.
  • Remove tags with unbounded values, such as user ID, session ID, request ID, full URL path with IDs or timestamps. Put those in logs or traces, where they belong.
  • Replace many near-identical metrics with one metric and a small, bounded tag.
  • Delete metrics created for a one-off investigation.

Say a latency metric is tagged by endpoint, status code and customer ID. If you have 40 endpoints, 6 status codes and 2,000 customers, that's 480,000 possible series. Dropping the customer tag takes it to 240, and you can still find slow customers through traces.

What about hosts, traces and synthetic tests?

These are smaller wins but quick ones:

  • Hosts: compare the list of hosts reporting to your cloud inventory. Remove agents from decommissioned or unused machines, and check that autoscaling doesn't leave phantom hosts counted long after they're gone.
  • Traces: lower the sampling rate for high-volume, healthy endpoints, and keep every error and slow request.
  • Synthetic tests: does a homepage check really need to run every minute from ten locations? Reduce frequency and locations for low-risk pages. The uptime monitoring checklist explains what deserves the tightest checks.
  • Dashboards and monitors: delete the ones nobody has opened in a quarter. They usually pin metrics that you pay to keep.

Assign a named owner to each service's telemetry budget. Costs drop faster when someone is accountable.

When does it make sense to compare another platform?

Optimize first, then compare. Otherwise you carry the same bad habits to a new vendor and find that the new bill grows the same way. If you still want to look, run the comparison on your real workload: same services, same retention, same alerts, for a representative month. Price models differ by what they meter, so a tool that's cheaper for log-heavy teams may cost more for metric-heavy ones. New Relic is one alternative many teams evaluate, and the platform comparison covers how to line them up.

Factor in the cost of migrating dashboards, alerts and on-call habits. Say a switch saves 10%: that rarely justifies a full re-instrumentation. When you do renew, negotiate on the reduced baseline you now have, not the inflated one, and confirm current pricing terms directly with the vendor.

Executive Capability Standard

What Good Looks Like

Every telemetry stream has an owner and a budget, and the monthly bill is reviewed against what people actually query.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Learn how each product on your invoice is metered: logs, custom metrics, hosts, traces and tests.
2. Do Manually:Rank the top three cost drivers and remove one source of waste in each this month.
3. Delegate:Give each service team an owner and a monthly telemetry budget review.
4. Automate:Add exclusion filters, tag allowlists and usage alerts that flag sudden growth before the invoice arrives.
5. Buy:Compare platforms on a real month of your workload, once you've trimmed what you send.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Datadog

Fits as the platform whose usage breakdown, exclusion filters and metric tools you'll use for these steps.

Visit Datadog→
New Relic

Fits as an alternative to compare on your own workload, after you have trimmed what you send.

Visit New Relic→

Frequently Asked Questions

Why is my Datadog bill so high?

Usually because of log volume, high-cardinality custom metrics, or too many hosts and tests. Open the usage breakdown, sort by cost and look at the top few items. One or two products typically account for most of the invoice.

What is metric cardinality and why does it cost money?

Cardinality is the number of unique combinations of a metric's name and tag values. Each combination is billed as its own series, so a tag with thousands of values, like a customer ID, can multiply cost dramatically.

Can I keep all my logs but pay less?

Often yes. Archive full logs to inexpensive storage and index only what you search regularly, such as warnings and errors. You can rehydrate archived logs when you need to investigate an incident.

Should we switch monitoring vendors to save money?

Not before trimming what you send. A new vendor bills a different set of dimensions, and migration takes real time. Optimize first, then compare vendors on your actual workload if the bill is still a problem.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides