Cutting Your Log Aggregation Bill Without Losing the Logs You Need
The safest way to cut a log aggregation bill is to find which logs actually earn their storage cost and trim the rest, rather than cutting volume across the board. The bill tends to grow quietly until someone finally looks at the invoice, and much of it is noise nobody has ever queried.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
How do you find out which logs you actually query?
Most logging platforms can show you query patterns over the last few months. Pull that report before you touch retention or sampling settings. Logs that have never been queried, not once, in three months are a very different cost problem than logs your on-call team pulls up during every incident.
This step alone usually finds the biggest, safest cuts: debug-level logging left on in a service nobody's actively working on, or a verbose third-party library logging at a level nobody asked for.
Separate Retention by How the Logs Get Used
Not every log needs the same retention window. Logs used for active debugging need to be fast to search but don't need to live for a year. Logs tied to compliance or security investigation need longer retention but get queried rarely, so they can live in cheaper, slower storage.
A single retention policy applied to everything is usually either too expensive for the debugging logs or too short for the compliance ones. Splitting them by actual use is often a bigger cost lever than any sampling change.
When should you sample logs and when should you keep everything?
A service producing millions of near-identical log lines an hour doesn't need every single one stored to be useful, sampling ten percent still shows you the pattern. A service that logs rarely, an error path that fires once a day, should never be sampled, because you can't afford to lose the one instance that mattered.
Apply sampling per log source based on its actual volume and criticality, not as one global percentage across your whole system.
Don't Cut the Logs You'll Need When Something's Actually Down
Aggressive log cuts feel safe until the exact incident happens where you needed the log line you just deleted. Keep reduced but tightly targeted verbose logging for genuinely high-risk paths, authentication, payments, anything touching money or access control, even if it costs more, because sacrificing your uptime budget during an incident you can't debug is a worse trade than the storage bill1.
The cost of an unresolved incident, extended downtime while you're flying blind, is almost always higher than the storage savings from cutting that one path's logging.
For example, a team samples a chatty service heavily to cut its bill, which is sensible. It also trims verbose logging on its login flow to save a little more. Weeks later, a sign-in incident lasts far longer than it should, because the log lines that would have shown the failing step were sampled away. The extra downtime costs more than the storage saved. Keep sampling on volume-heavy sources where the pattern matters, and protect the handful of paths where a single event decides whether you can diagnose an outage.
Revisit the Cuts Every Time the Architecture Changes
A logging policy tuned for last year's traffic pattern and service mix goes stale as your system changes. A new service that starts at high volume, or an old one that gets deprecated, both shift where your logging budget should actually go. Review the split every couple of quarters rather than setting it once and forgetting it.
Get Engineering Buy-In Before You Cut Anything
A cost cut that lands without warning, where an engineer discovers mid-incident that the log line they needed was sampled away last month, burns trust fast and makes the next cost conversation much harder to have.
Share the proposed cuts with the team before making them, specifically calling out anything touching a path someone actively debugs. A quick review catches a genuinely needed log before it's gone, and it means the team understands why the bill changed instead of being surprised by a gap during the next incident.
This also makes the next round of cuts easier to propose. A team that trusts the process because the last round didn't quietly remove anything they needed is far more receptive the second time around than one that has already been burned once before and now questions every single proposed cut on principle, whether that specific cut is genuinely warranted or not at all, every single time the topic comes up again.
Cut log costs safely in this order:
- Pull the query report for the last few months and list the sources nobody has queried.
- Sample or drop those unqueried sources first, such as leftover debug logging or noisy third-party libraries.
- Split retention by use, keeping debugging logs fast to search and compliance logs in cheaper, slower storage.
- Sample high-volume, near-identical sources one by one, and never sample rare error paths.
- Keep targeted verbose logging on authentication, payments, and access control paths, then share the plan with engineers before cutting.
What Good Looks Like
An efficient logging setup splits retention and sampling by how each log source is actually used, and protects full logging on the paths where losing a single event would hurt during a real incident.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
How do we know if a log source is safe to sample instead of keeping fully?
Ask whether losing ninety percent of the individual events would still let you diagnose a problem in that area. If the value is in the pattern (request volume, error rate) rather than any single event, sampling is usually safe. If each event is individually meaningful, like an auth failure, keep it all.
Should compliance-related logs use the same retention as debugging logs?
No. Compliance and security logs typically need longer retention but get queried far less often, so they belong in cheaper, slower storage with a longer window. Debugging logs need fast search but a much shorter useful life.
What's the fastest safe cut for a team with a runaway logging bill?
Find log sources that have never actually been queried in the last few months and either sample or drop them first. That's almost always a bigger, safer win than tightening retention on logs your team actively uses.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
Your Log Bill Is Growing Because Nobody Decided What to Keep
Log volume usually grows because every team logs everything by default. Here are three ways to cut the bill without losing the logs you'll actually need.
Cutting Log Volume Without Losing the Logs You Need
A worked walkthrough for reducing log aggregation cost on a real-time pipeline by cutting volume deliberately instead of just raising a retention limit.
Where a Log Aggregation Bill Actually Goes, Traced Line by Line
A cost teardown of a typical log aggregation bill, showing which log volume is worth paying for and which is silently expensive debug noise.
Why Your Log Bill Grows Faster Than Your Traffic
Log volume usually grows faster than the traffic producing it. Where that gap actually comes from, and the retention and sampling changes that close it.
Cutting Log Aggregation Costs Without Losing Signal
How to cut log aggregation costs with tiered storage, sampling and retention rules, while keeping the logs you need during an incident.
Where Your Log Aggregation Bill Is Actually Going
A worked look at where a log aggregation bill actually comes from, and which cuts save real money without losing the logs you'd need during an incident.