Developer Productivity & Platform EngineeringPlaybook3 min readUpdated September 2026

Cutting Your Log Aggregation Bill Without Losing the Logs You Need

The safest way to cut a log aggregation bill is to find which logs actually earn their storage cost and trim the rest, rather than cutting volume across the board. The bill tends to grow quietly until someone finally looks at the invoice, and much of it is noise nobody has ever queried.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

How do you find out which logs you actually query?

Most logging platforms can show you query patterns over the last few months. Pull that report before you touch retention or sampling settings. Logs that have never been queried, not once, in three months are a very different cost problem than logs your on-call team pulls up during every incident.

This step alone usually finds the biggest, safest cuts: debug-level logging left on in a service nobody's actively working on, or a verbose third-party library logging at a level nobody asked for.

Separate Retention by How the Logs Get Used

Not every log needs the same retention window. Logs used for active debugging need to be fast to search but don't need to live for a year. Logs tied to compliance or security investigation need longer retention but get queried rarely, so they can live in cheaper, slower storage.

A single retention policy applied to everything is usually either too expensive for the debugging logs or too short for the compliance ones. Splitting them by actual use is often a bigger cost lever than any sampling change.

When should you sample logs and when should you keep everything?

A service producing millions of near-identical log lines an hour doesn't need every single one stored to be useful, sampling ten percent still shows you the pattern. A service that logs rarely, an error path that fires once a day, should never be sampled, because you can't afford to lose the one instance that mattered.

Apply sampling per log source based on its actual volume and criticality, not as one global percentage across your whole system.

Don't Cut the Logs You'll Need When Something's Actually Down

Aggressive log cuts feel safe until the exact incident happens where you needed the log line you just deleted. Keep reduced but tightly targeted verbose logging for genuinely high-risk paths, authentication, payments, anything touching money or access control, even if it costs more, because sacrificing your uptime budget during an incident you can't debug is a worse trade than the storage bill1.

The cost of an unresolved incident, extended downtime while you're flying blind, is almost always higher than the storage savings from cutting that one path's logging.

For example, a team samples a chatty service heavily to cut its bill, which is sensible. It also trims verbose logging on its login flow to save a little more. Weeks later, a sign-in incident lasts far longer than it should, because the log lines that would have shown the failing step were sampled away. The extra downtime costs more than the storage saved. Keep sampling on volume-heavy sources where the pattern matters, and protect the handful of paths where a single event decides whether you can diagnose an outage.

Revisit the Cuts Every Time the Architecture Changes

A logging policy tuned for last year's traffic pattern and service mix goes stale as your system changes. A new service that starts at high volume, or an old one that gets deprecated, both shift where your logging budget should actually go. Review the split every couple of quarters rather than setting it once and forgetting it.

Get Engineering Buy-In Before You Cut Anything

A cost cut that lands without warning, where an engineer discovers mid-incident that the log line they needed was sampled away last month, burns trust fast and makes the next cost conversation much harder to have.

Share the proposed cuts with the team before making them, specifically calling out anything touching a path someone actively debugs. A quick review catches a genuinely needed log before it's gone, and it means the team understands why the bill changed instead of being surprised by a gap during the next incident.

This also makes the next round of cuts easier to propose. A team that trusts the process because the last round didn't quietly remove anything they needed is far more receptive the second time around than one that has already been burned once before and now questions every single proposed cut on principle, whether that specific cut is genuinely warranted or not at all, every single time the topic comes up again.

Cut log costs safely in this order:

  1. Pull the query report for the last few months and list the sources nobody has queried.
  2. Sample or drop those unqueried sources first, such as leftover debug logging or noisy third-party libraries.
  3. Split retention by use, keeping debugging logs fast to search and compliance logs in cheaper, slower storage.
  4. Sample high-volume, near-identical sources one by one, and never sample rare error paths.
  5. Keep targeted verbose logging on authentication, payments, and access control paths, then share the plan with engineers before cutting.
Executive Capability Standard

What Good Looks Like

An efficient logging setup splits retention and sampling by how each log source is actually used, and protects full logging on the paths where losing a single event would hurt during a real incident.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Pull a query report from your logging platform to see which log sources are actually being used before cutting anything.
2. Do Manually:Manually split retention into a short, fast tier for debugging logs and a longer, cheaper tier for compliance logs.
3. Delegate:Assign an engineer to own logging cost review each quarter as the architecture and traffic pattern shifts.
4. Automate:Automate volume-based sampling rules per log source instead of a single global sampling rate.
5. Buy:Bring in a specialist to review your logging architecture if the bill has grown large enough that manual tuning no longer keeps pace.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

ClickUp

Track the quarterly logging cost review as a recurring ClickUp task tied to your architecture changes.

Visit ClickUp→
Trainual

Document which log sources are protected from sampling and why in Trainual so the reasoning survives a team change.

Visit Trainual→

Frequently Asked Questions

How do we know if a log source is safe to sample instead of keeping fully?

Ask whether losing ninety percent of the individual events would still let you diagnose a problem in that area. If the value is in the pattern (request volume, error rate) rather than any single event, sampling is usually safe. If each event is individually meaningful, like an auth failure, keep it all.

Should compliance-related logs use the same retention as debugging logs?

No. Compliance and security logs typically need longer retention but get queried far less often, so they belong in cheaper, slower storage with a longer window. Debugging logs need fast search but a much shorter useful life.

What's the fastest safe cut for a team with a runaway logging bill?

Find log sources that have never actually been queried in the last few months and either sample or drop them first. That's almost always a bigger, safer win than tightening retention on logs your team actively uses.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides