Enterprise DevSecOps & Automated CompliancePlaybook3 min readUpdated September 2026

Where a Log Aggregation Bill Actually Goes, Traced Line by Line

A log aggregation bill that's grown well past what anyone expected is rarely one big problem. It's usually a handful of specific, traceable sources adding up quietly, each individually reasonable at the moment someone added it. Here's how to trace a bill line by line back to its actual sources, and which ones are worth keeping.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Why start with volume by source instead of total cost?

Most log aggregation platforms let you break ingested volume down by service or log stream. Pull that breakdown before looking at the total bill. In a typical teardown, a small number of sources, often two or three, account for most of the volume, and they're rarely the ones anyone would guess without looking.

The usual surprises: a health check endpoint logging every single request at a normal log level instead of being excluded or sampled, and a third-party library's debug logging left on because nobody explicitly turned it off after using it to troubleshoot something months earlier.

Separate debug noise from what you'd actually use during an incident

For each high-volume source, ask a direct question: during the last real incident that touched this service, did anyone actually query these specific logs, or did the investigation use metrics and traces instead. Debug-level logging that's never queried during an incident is pure cost with no offsetting value, and it's usually the largest single category once you actually trace it.

This is different from asking whether the logs could theoretically be useful someday. Almost any log could theoretically help someday. The question that actually reduces cost is whether it has helped recently, for this specific service, at this specific volume.

Ask these questions about each high-volume log source:

  • Did anyone query these logs during the last real incident that touched this service, or did the investigation rely on metrics and traces?
  • Does the source carry security or audit value that justifies longer retention on a search-optimized tier?
  • Would sampling keep enough signal to spot a pattern or a rate change at a fraction of the ingestion cost?
  • What would the service owner actually reach for during an incident, and does the tier match their answer?

Route by value, not by convenience

Not every log needs the same retention or the same storage tier. Security and audit-relevant logs justify longer retention and a search-optimized tier because they're queried unpredictably, often long after the fact. High-volume application debug logs are better suited to a cheap, short-retention tier, since their value drops sharply once the specific incident they might help with is resolved.

Most teams route everything through one pipeline at one retention setting because it was simpler to set up that way initially, not because every log genuinely needs the same treatment. Splitting the routing by value is usually the single biggest lever on a bloated bill.

Should you sample logs instead of dropping them?

For genuinely high-volume, low-value logs, sampling (keeping a percentage of events rather than either all of them or none of them) preserves enough signal to spot a pattern or a rate change without paying full ingestion cost. A health check log sampled down to a small fraction still tells you if something's broken across the board, at a fraction of the ingestion cost of logging every single check.

Be deliberate about what you sample, though. Anything tied to a security or compliance requirement should stay fully logged regardless of volume; sampling is a cost lever for operational debug noise, not a shortcut for audit-relevant events.

Building a lightweight review into the pipeline itself

Once the bill is under control, the same drift that caused it the first time will happen again unless something catches it early. Add a monthly check of volume by source, the same query used for the initial teardown, and treat any source whose volume jumps sharply between months as worth a quick look before it becomes next quarter's surprise line item.

The conversation worth having with each service owner

Cost reduction lands better as a collaborative question than a top-down mandate: ask each service owner directly what they'd actually reach for during an incident on their service, and set that source's tier based on their answer rather than a blanket policy applied without their input. An owner who knows their service's real failure modes will usually give a more accurate answer than a generic retention policy could guess at.

This also builds the habit of thinking about log value at write time, not just during a periodic cleanup, which is where the healthiest version of this practice ends up: new high-volume logging gets a deliberate tier decision from the start, instead of defaulting to the most expensive option and waiting for the next teardown to notice.

Executive Capability Standard

What Good Looks Like

Good log aggregation practice means every high-volume source has a known owner, a value assessment against real incident use, and a retention tier that matches that value rather than a single default setting.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Read your aggregation platform's documentation on volume-by-source reporting and sampling options before assuming a cost problem requires switching vendors.
2. Do Manually:Pull the volume-by-source breakdown and trace the top three sources by hand, checking each against recent incident history.
3. Delegate:Assign each high-volume log source an owner responsible for justifying its retention tier and volume at the next review.
4. Automate:Add a monthly automated report of volume by source with month-over-month change, so drift gets caught before it compounds into a large bill.
5. Buy:Consider a dedicated log management or observability pipeline tool once you're routing enough distinct sources that manual tier assignment stops scaling.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Frequently Asked Questions

How much can a teardown like this typically reduce a log aggregation bill?

The reduction depends on how much debug logging has piled up unchecked, but most teams find a meaningful drop once they trace the bill. Two or three sources usually dominate volume, and they're rarely the ones anyone would guess without looking. Start the trace with volume by source, not the total.

Is it safe to just turn off a high-volume debug log source entirely?

Check with whoever owns that service first, since a log that looks unused from the aggregation platform's query history might still be pulled directly from the source during specific, infrequent troubleshooting. Downgrade to a cheaper tier or sample it before deleting it outright.

Should this review happen before or after choosing a log aggregation vendor?

Before, ideally. Understanding your actual volume by source and value gives you a real number to negotiate pricing against, instead of estimating usage from a rough guess that a vendor's sales process will happily round up.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides