Engineering Leadership & Technical HiringPlaybook3 min readUpdated September 2026

Your Log Bill Is Growing Because Nobody Decided What to Keep

Log aggregation bills tend to grow quietly and then show up as a surprise line item, because logging is one of the few systems where every team can add volume without asking anyone. A new debug line here, a verbose library defaulting to info level there, and eighteen months later you're paying to store and index a volume nobody actually queries.

The fix isn't logging less everywhere. It's deciding, deliberately, what actually needs to be searchable in real time, what can sit in cheaper storage for the rare audit or incident, and what shouldn't be logged at all.

Why Log Volume Grows Without Anyone Deciding It Should

Nobody sits down and decides to triple the logging budget. It happens one small addition at a time: a new service ships with its framework's default logging level, a library that logs every request body for debugging gets left on past the debugging session, a team adds detailed tracing during an incident and forgets to remove it afterward.

Each individual addition looks negligible. The aggregate, across dozens of services over a year, is what actually shows up on the invoice, and it's very hard to trace back to any single decision because no single decision caused it.

A useful first step is to stop looking at the total bill and start ranking log sources by volume. For example, sort your services by lines written per day and read the top few. Teams often find that one framework default, one leftover debug session or one endpoint logging a full request body accounts for a large share of the total. Fix those sources first, then write down an owner and a retention rule for each one so the same growth does not quietly return. Ranking by source turns a vague cost complaint into a short, assignable list.

Approach One: Sampling, and Where It Actually Works

Sampling keeps a fraction of log lines instead of all of them, and it works well for high-volume, low-value events: successful health checks, routine request logs on endpoints that rarely error. It works badly for anything you need complete for compliance or precise incident reconstruction, since a sampled log is missing exactly the line you might need most.

Apply sampling selectively, by log source, not as a single global percentage across everything. A payment service and a health check endpoint have very different tolerances for missing a specific line, and treating them identically either under-samples the health checks or over-samples the payment logs.

Approach Two: Tiered Retention Instead of One Retention Policy

Most logging platforms charge far more to keep data in hot, instantly searchable storage than in cold, cheaper archival storage. Keep recent logs, the last week or two, in hot storage where you actually need fast search during an active incident, and move older logs to cold storage where they're still retrievable for an audit or a rare historical investigation but no longer costing full price to sit there unread.

Say your team currently keeps ninety days of everything in hot storage: moving anything past two weeks to cold storage is often the single largest lever available, since the vast majority of log queries happen against the most recent few days, not against data from two months ago.

Approach Three: Deciding What Shouldn't Be Logged at All

Some volume isn't a retention problem, it's a logging problem: a verbose debug line that never gets read, a full request body logged for a high-traffic endpoint where the body adds no diagnostic value over the status code and latency you're already capturing. Cutting that at the source costs nothing to store or index, because it's never generated in the first place.

Review your highest-volume log sources specifically for this, not your logging policy in the abstract. A single high-traffic endpoint logging an unnecessary field can dwarf the volume of every other service combined.

Picking the Right Mix for Your Team

These three approaches aren't mutually exclusive, and most teams end up combining them: sample the noisy, low-value sources, tier retention so only recent data sits in expensive hot storage, and trim the specific high-volume fields that never actually get queried. Start with whichever lever is cheapest to pull for your setup, tiered retention usually requires the least code change, and measure the bill impact before layering on the next one.

Whatever mix you choose, keep an explicit, written retention policy per log source rather than letting each new service default to whatever its framework ships with. That's the actual root cause of the slow creep in the first place.

Match each cost lever to the log source it fits:

  • Sample high-volume, low-value events such as successful health checks and routine request logs on endpoints that rarely error.
  • Tier retention so only the last week or two sits in hot, searchable storage and older logs move to cold storage.
  • Stop generating verbose debug lines and full request bodies that nobody ever queries.
  • Keep error logs, traces on critical paths and compliance-relevant logs fully retained, in the cheapest tier that meets the requirement.
Executive Capability Standard

What Good Looks Like

Good log management means every log source has a deliberate, written decision behind it: what's sampled, how long it stays in hot storage before moving to cold storage, and what shouldn't be logged at all, rather than every service defaulting to its framework's out of the box settings.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Identify your highest-volume log sources and check whether anyone is actually querying that volume regularly.
2. Do Manually:Manually move one high-volume, low-value source to sampled or shorter hot retention and measure the bill impact.
3. Delegate:Assign an engineer ownership of a written retention policy per log source that new services are expected to follow.
4. Automate:Configure tiered retention rules in your logging platform so aging logs move to cold storage automatically without manual cleanup.
5. Buy:Move to a logging platform with native tiered storage and sampling controls if your current tooling can't express these policies directly.

How to Get Started

Frequently Asked Questions

Will sampling logs make it harder to debug a production incident?

It can, if applied to the wrong sources. Reserve full retention for anything you'd need complete during an incident, error logs and traces on your critical paths, and sample only the high-volume, low-value noise like successful health checks. Applied selectively, sampling removes bill, not debugging capability.

How much can tiered retention actually save?

It depends heavily on your current setup, but moving anything past one to two weeks out of hot, instantly searchable storage is often the largest single lever available, since most logging platforms price hot storage well above cold archival storage. Measure your own query patterns before assuming a specific savings number.

Should compliance-relevant logs ever be sampled?

No. Anything you're required to retain completely for an audit or a compliance obligation should stay fully retained, in whatever storage tier is cheapest while still meeting the retention requirement. Sampling is for volume you don't strictly need complete, not for volume you're contractually or legally required to keep.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides