Your Log Bill Is Growing Because Nobody Decided What to Keep
Log aggregation bills tend to grow quietly and then show up as a surprise line item, because logging is one of the few systems where every team can add volume without asking anyone. A new debug line here, a verbose library defaulting to info level there, and eighteen months later you're paying to store and index a volume nobody actually queries.
The fix isn't logging less everywhere. It's deciding, deliberately, what actually needs to be searchable in real time, what can sit in cheaper storage for the rare audit or incident, and what shouldn't be logged at all.
Why Log Volume Grows Without Anyone Deciding It Should
Nobody sits down and decides to triple the logging budget. It happens one small addition at a time: a new service ships with its framework's default logging level, a library that logs every request body for debugging gets left on past the debugging session, a team adds detailed tracing during an incident and forgets to remove it afterward.
Each individual addition looks negligible. The aggregate, across dozens of services over a year, is what actually shows up on the invoice, and it's very hard to trace back to any single decision because no single decision caused it.
A useful first step is to stop looking at the total bill and start ranking log sources by volume. For example, sort your services by lines written per day and read the top few. Teams often find that one framework default, one leftover debug session or one endpoint logging a full request body accounts for a large share of the total. Fix those sources first, then write down an owner and a retention rule for each one so the same growth does not quietly return. Ranking by source turns a vague cost complaint into a short, assignable list.
Approach One: Sampling, and Where It Actually Works
Sampling keeps a fraction of log lines instead of all of them, and it works well for high-volume, low-value events: successful health checks, routine request logs on endpoints that rarely error. It works badly for anything you need complete for compliance or precise incident reconstruction, since a sampled log is missing exactly the line you might need most.
Apply sampling selectively, by log source, not as a single global percentage across everything. A payment service and a health check endpoint have very different tolerances for missing a specific line, and treating them identically either under-samples the health checks or over-samples the payment logs.
Approach Two: Tiered Retention Instead of One Retention Policy
Most logging platforms charge far more to keep data in hot, instantly searchable storage than in cold, cheaper archival storage. Keep recent logs, the last week or two, in hot storage where you actually need fast search during an active incident, and move older logs to cold storage where they're still retrievable for an audit or a rare historical investigation but no longer costing full price to sit there unread.
Say your team currently keeps ninety days of everything in hot storage: moving anything past two weeks to cold storage is often the single largest lever available, since the vast majority of log queries happen against the most recent few days, not against data from two months ago.
Approach Three: Deciding What Shouldn't Be Logged at All
Some volume isn't a retention problem, it's a logging problem: a verbose debug line that never gets read, a full request body logged for a high-traffic endpoint where the body adds no diagnostic value over the status code and latency you're already capturing. Cutting that at the source costs nothing to store or index, because it's never generated in the first place.
Review your highest-volume log sources specifically for this, not your logging policy in the abstract. A single high-traffic endpoint logging an unnecessary field can dwarf the volume of every other service combined.
Picking the Right Mix for Your Team
These three approaches aren't mutually exclusive, and most teams end up combining them: sample the noisy, low-value sources, tier retention so only recent data sits in expensive hot storage, and trim the specific high-volume fields that never actually get queried. Start with whichever lever is cheapest to pull for your setup, tiered retention usually requires the least code change, and measure the bill impact before layering on the next one.
Whatever mix you choose, keep an explicit, written retention policy per log source rather than letting each new service default to whatever its framework ships with. That's the actual root cause of the slow creep in the first place.
Match each cost lever to the log source it fits:
- Sample high-volume, low-value events such as successful health checks and routine request logs on endpoints that rarely error.
- Tier retention so only the last week or two sits in hot, searchable storage and older logs move to cold storage.
- Stop generating verbose debug lines and full request bodies that nobody ever queries.
- Keep error logs, traces on critical paths and compliance-relevant logs fully retained, in the cheapest tier that meets the requirement.
What Good Looks Like
Good log management means every log source has a deliberate, written decision behind it: what's sampled, how long it stays in hot storage before moving to cold storage, and what shouldn't be logged at all, rather than every service defaulting to its framework's out of the box settings.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Will sampling logs make it harder to debug a production incident?
It can, if applied to the wrong sources. Reserve full retention for anything you'd need complete during an incident, error logs and traces on your critical paths, and sample only the high-volume, low-value noise like successful health checks. Applied selectively, sampling removes bill, not debugging capability.
How much can tiered retention actually save?
It depends heavily on your current setup, but moving anything past one to two weeks out of hot, instantly searchable storage is often the largest single lever available, since most logging platforms price hot storage well above cold archival storage. Measure your own query patterns before assuming a specific savings number.
Should compliance-relevant logs ever be sampled?
No. Anything you're required to retain completely for an audit or a compliance obligation should stay fully retained, in whatever storage tier is cheapest while still meeting the retention requirement. Sampling is for volume you don't strictly need complete, not for volume you're contractually or legally required to keep.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Cutting Your Log Aggregation Bill Without Losing the Logs You Need
How to reduce a runaway log aggregation bill without cutting the specific logs you'd actually need during your next real incident.
Where a Log Aggregation Bill Actually Goes, Traced Line by Line
A cost teardown of a typical log aggregation bill, showing which log volume is worth paying for and which is silently expensive debug noise.
Cutting Log Volume Without Losing the Logs You Need
A worked walkthrough for reducing log aggregation cost on a real-time pipeline by cutting volume deliberately instead of just raising a retention limit.
Why Your Log Bill Grows Faster Than Your Traffic
Log volume usually grows faster than the traffic producing it. Where that gap actually comes from, and the retention and sampling changes that close it.
Where Your Log Aggregation Bill Is Actually Going
A worked look at where a log aggregation bill actually comes from, and which cuts save real money without losing the logs you'd need during an incident.
Cutting Log Aggregation Costs Without Losing Signal
How to cut log aggregation costs with tiered storage, sampling and retention rules, while keeping the logs you need during an incident.