Model Context Protocol & Agentic ArchitecturePlaybook3 min readUpdated September 2026

Where Your Log Aggregation Bill Is Actually Going

Log aggregation bills tend to grow quietly until someone finally looks at the invoice, and the reflex response, cut retention across the board, often removes exactly the logs you'd need six weeks from now during an investigation. A more useful approach starts with understanding where the volume actually comes from before deciding what to cut.

Here's a worked walkthrough of that process, using the same steps you can run against your own aggregation dashboard this week.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

How do you break a log aggregation bill down by source?

Most log aggregation costs concentrate in a small number of sources: a chatty debug-level logger left on in production, a health check endpoint logging every single call, or a single high-traffic service generating the bulk of total volume. Pull a breakdown by source before making any retention or sampling decision, because cutting evenly across all sources treats a one-line debug log the same as an audit trail, when they don't remotely carry the same value.

Look for these common sources of outsized log volume:

  • A chatty debug-level logger left on in production after a debugging session that was never reverted.
  • A health check endpoint that logs a line on every single call.
  • A single high-traffic service that generates the bulk of total volume.
  • A retention policy applied evenly to every source, so low-value debug lines are kept as long as audit trails.

Which logs need long retention and which don't?

Security and compliance-relevant logs, the kind a platform like Vanta or CrowdStrike expects a feed of, often need retention measured in months or years. Routine application debug logs rarely need more than a few weeks, since their entire value is in helping someone debug something that just happened, not something from six months ago. Applying one retention policy to both categories is the single most common source of overspend, because the expensive-to-keep category is usually the smaller one by volume.

Sample high-volume, low-value log lines instead of dropping them entirely

A health check endpoint that logs a line on every successful call, potentially thousands of times an hour, is rarely useful in that volume; sampling it, logging one in every hundred successful calls while keeping every failure, preserves the ability to spot a trend without paying to store and index near-identical lines. Apply sampling selectively to specifically identified high-volume, low-value sources, not broadly, since sampling something that actually matters trades cost savings for blind spots during an incident.

Fix the loggers generating noise instead of just filtering their output

A verbose logger left at debug level in production is often an oversight from a debugging session that never got reverted, not a deliberate choice. Say a single service is responsible for a large share of your total log volume through an overly chatty debug logger; fixing the log level at the source removes that cost entirely, rather than paying to ingest, store, and then filter out the same noise downstream at the aggregation layer.

Recheck the breakdown after each change

After adjusting retention, sampling, or log levels, pull the source breakdown again and confirm the savings actually materialized and nothing important quietly disappeared along with the noise. A change that looked right on paper can behave differently once it's live, particularly with sampling, where an edge case you didn't anticipate can end up systematically excluded rather than randomly sampled.

When you recheck the breakdown, compare more than the total. Confirm the noisy source you targeted actually shrank, then confirm that the logs you would need in an investigation are still arriving, such as errors and security events tied to a customer report. For example, if you began sampling a health check, look at a recent incident window and make sure failures were all kept and only the routine successes were thinned. A common mistake is celebrating a lower invoice without testing what disappeared. If a source vanished entirely, treat that as a bug in the change, not as savings.

Put someone in charge of the trend, not just the one-time cut

A single cost-cutting pass fixes the bill for one billing cycle; new noisy sources and forgotten debug flags reappear on their own over time as the codebase changes. Give a specific person or team ownership of reviewing the source breakdown on a recurring basis, the same way you would for any other infrastructure cost, so the bill doesn't quietly creep back to where it started six months after the initial cleanup.

Set a budget per service, not just a total

A single company-wide log budget makes it hard to tell which team's change actually caused a jump in the bill, since the total moves for reasons unrelated to any one team's decisions. Give each service or team a rough log volume budget of its own, visible to that team specifically, so the person who added the chatty logger is also the person who sees the number move and can connect the two directly instead of the finding surfacing weeks later in a company-wide review nobody traces back to its source.

Executive Capability Standard

What Good Looks Like

Efficient log aggregation means retention and sampling decisions are made per source based on actual value, not applied evenly across everything, with security and compliance-relevant logs kept at full fidelity and noisy, low-value sources fixed at their origin rather than filtered downstream.

Building The Capability (5-Stage Skill Ladder)

1. Learn:pull a breakdown of your current log volume by source and identify the one or two sources driving most of the cost
2. Do Manually:review retention settings by hand for your top few log sources and adjust the ones mismatched to their actual value
3. Delegate:give a specific owner responsibility for the logging pipeline's cost and retention policy, separate from individual service teams
4. Automate:set per-source retention and sampling rules in your aggregation pipeline so new noisy sources don't silently inflate the bill the same way
5. Buy:compliance platforms like Vanta and endpoint tools like CrowdStrike both expect a feed of specific log categories, which is worth accounting for explicitly when you're deciding what not to cut

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Frequently Asked Questions

Is sampling logs risky during an incident?

It can be, if applied to the wrong sources. Sample high-volume, low-value lines like routine health checks, and keep full fidelity on anything you'd actually need during an investigation: errors, security-relevant events, and anything tied to a specific customer-reported issue.

How do we know which logs are actually security or compliance relevant?

Check against whichever compliance framework or customer contract you're tracking evidence for; those typically name specific categories of events, like access changes or authentication attempts, that need longer retention regardless of their volume.

What's the fastest way to find the biggest cost driver?

Pull a breakdown of log volume by source or service for the last billing period. In most teams, one or two sources account for a disproportionate share of the total, and that's where to focus first rather than spreading effort evenly.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides