Where Your Log Aggregation Bill Is Actually Going
Log aggregation bills tend to grow quietly until someone finally looks at the invoice, and the reflex response, cut retention across the board, often removes exactly the logs you'd need six weeks from now during an investigation. A more useful approach starts with understanding where the volume actually comes from before deciding what to cut.
Here's a worked walkthrough of that process, using the same steps you can run against your own aggregation dashboard this week.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
How do you break a log aggregation bill down by source?
Most log aggregation costs concentrate in a small number of sources: a chatty debug-level logger left on in production, a health check endpoint logging every single call, or a single high-traffic service generating the bulk of total volume. Pull a breakdown by source before making any retention or sampling decision, because cutting evenly across all sources treats a one-line debug log the same as an audit trail, when they don't remotely carry the same value.
Look for these common sources of outsized log volume:
- A chatty debug-level logger left on in production after a debugging session that was never reverted.
- A health check endpoint that logs a line on every single call.
- A single high-traffic service that generates the bulk of total volume.
- A retention policy applied evenly to every source, so low-value debug lines are kept as long as audit trails.
Which logs need long retention and which don't?
Security and compliance-relevant logs, the kind a platform like Vanta or CrowdStrike expects a feed of, often need retention measured in months or years. Routine application debug logs rarely need more than a few weeks, since their entire value is in helping someone debug something that just happened, not something from six months ago. Applying one retention policy to both categories is the single most common source of overspend, because the expensive-to-keep category is usually the smaller one by volume.
Sample high-volume, low-value log lines instead of dropping them entirely
A health check endpoint that logs a line on every successful call, potentially thousands of times an hour, is rarely useful in that volume; sampling it, logging one in every hundred successful calls while keeping every failure, preserves the ability to spot a trend without paying to store and index near-identical lines. Apply sampling selectively to specifically identified high-volume, low-value sources, not broadly, since sampling something that actually matters trades cost savings for blind spots during an incident.
Fix the loggers generating noise instead of just filtering their output
A verbose logger left at debug level in production is often an oversight from a debugging session that never got reverted, not a deliberate choice. Say a single service is responsible for a large share of your total log volume through an overly chatty debug logger; fixing the log level at the source removes that cost entirely, rather than paying to ingest, store, and then filter out the same noise downstream at the aggregation layer.
Recheck the breakdown after each change
After adjusting retention, sampling, or log levels, pull the source breakdown again and confirm the savings actually materialized and nothing important quietly disappeared along with the noise. A change that looked right on paper can behave differently once it's live, particularly with sampling, where an edge case you didn't anticipate can end up systematically excluded rather than randomly sampled.
When you recheck the breakdown, compare more than the total. Confirm the noisy source you targeted actually shrank, then confirm that the logs you would need in an investigation are still arriving, such as errors and security events tied to a customer report. For example, if you began sampling a health check, look at a recent incident window and make sure failures were all kept and only the routine successes were thinned. A common mistake is celebrating a lower invoice without testing what disappeared. If a source vanished entirely, treat that as a bug in the change, not as savings.
Put someone in charge of the trend, not just the one-time cut
A single cost-cutting pass fixes the bill for one billing cycle; new noisy sources and forgotten debug flags reappear on their own over time as the codebase changes. Give a specific person or team ownership of reviewing the source breakdown on a recurring basis, the same way you would for any other infrastructure cost, so the bill doesn't quietly creep back to where it started six months after the initial cleanup.
Set a budget per service, not just a total
A single company-wide log budget makes it hard to tell which team's change actually caused a jump in the bill, since the total moves for reasons unrelated to any one team's decisions. Give each service or team a rough log volume budget of its own, visible to that team specifically, so the person who added the chatty logger is also the person who sees the number move and can connect the two directly instead of the finding surfacing weeks later in a company-wide review nobody traces back to its source.
What Good Looks Like
Efficient log aggregation means retention and sampling decisions are made per source based on actual value, not applied evenly across everything, with security and compliance-relevant logs kept at full fidelity and noisy, low-value sources fixed at their origin rather than filtered downstream.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Vanta expects a feed of specific compliance-relevant log categories as evidence, which is a reason to keep those categories out of any broad retention cut even while trimming everything else.
CrowdStrike's detection relies on a feed of security-relevant events, so cost-cutting on log volume should specifically exclude whatever categories it depends on rather than applying one retention policy everywhere.
Frequently Asked Questions
Is sampling logs risky during an incident?
It can be, if applied to the wrong sources. Sample high-volume, low-value lines like routine health checks, and keep full fidelity on anything you'd actually need during an investigation: errors, security-relevant events, and anything tied to a specific customer-reported issue.
How do we know which logs are actually security or compliance relevant?
Check against whichever compliance framework or customer contract you're tracking evidence for; those typically name specific categories of events, like access changes or authentication attempts, that need longer retention regardless of their volume.
What's the fastest way to find the biggest cost driver?
Pull a breakdown of log volume by source or service for the last billing period. In most teams, one or two sources account for a disproportionate share of the total, and that's where to focus first rather than spreading effort evenly.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Cutting Your Log Aggregation Bill Without Losing the Logs You Need
How to reduce a runaway log aggregation bill without cutting the specific logs you'd actually need during your next real incident.
Where a Log Aggregation Bill Actually Goes, Traced Line by Line
A cost teardown of a typical log aggregation bill, showing which log volume is worth paying for and which is silently expensive debug noise.
Keeping Inference Log Volume From Outrunning Your Budget
Why inference logging costs grow faster than traffic, and practical ways to sample, structure, and trim logs without losing what you need to debug a failure.
Why Your Log Bill Grows Faster Than Your Traffic
Log volume usually grows faster than the traffic producing it. Where that gap actually comes from, and the retention and sampling changes that close it.
Your Log Bill Is Growing Because Nobody Decided What to Keep
Log volume usually grows because every team logs everything by default. Here are three ways to cut the bill without losing the logs you'll actually need.
Cutting Log Aggregation Costs Without Losing Signal
How to cut log aggregation costs with tiered storage, sampling and retention rules, while keeping the logs you need during an incident.