Why Your Log Bill Grows Faster Than Your Traffic
Traffic doubles and the logging bill triples. That gap is common enough to have a predictable cause: most log volume growth isn't more requests, it's more verbosity per request, debug lines left on in production, retry loops logging every attempt, a new service that logs at a level nobody reviewed before it shipped.
Fixing it isn't about logging less useful information. It's about being deliberate for the first time about what gets kept, at what detail, for how long.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Why does log volume grow faster than traffic?
A service logging at debug level in production, because nobody changed the default after the initial setup, can easily produce ten times the volume of the same service logging only what an operator would actually need during an incident. Audit your highest volume log sources first, not your newest ones: the biggest single fix is usually one or two chatty services logging far more than anyone reads, not a broad problem across the whole fleet.
Does sampling cut log costs more than trimming verbosity?
Reducing log detail per line saves a little. Sampling, keeping every error and a representative slice of successful requests instead of logging every single one, changes the actual cost curve, because the volume driver in most systems is the high frequency, low information successful request, not the rare, information dense error. Set your sample rate so every error is always kept in full and successful requests are sampled at whatever rate your budget allows.
Match retention tiers to how often you actually query old data
Most log queries happen against the last few days; queries against data older than a month are rare and usually driven by a specific investigation, not routine monitoring. A three tier setup, hot storage for the last week where every field is instantly searchable, warm for the last quarter at lower cost and slower query speed, cold archival storage beyond that, matches spend to actual access patterns instead of paying hot storage prices for data nobody's queried in months.
Structured logging pays for itself at query time
Free text log lines are cheap to write and expensive to search: finding every request that hit a specific error requires a fuzzy text search across everything. Structured, field based logging costs a little more discipline upfront but turns that same search into an exact filter, which matters directly during an incident when a 99.95% uptime target leaves roughly 4.38 hours a year of downtime budget to work with across every incident combined1, and every minute spent searching unstructured text is a minute of that budget spent on the search instead of the fix.
The common mistake: keeping everything at full detail 'just in case'
Logging everything indefinitely feels safer than deciding what to drop, but it's a decision by default, not a real strategy, and it's usually the single biggest line item on the bill. The honest question isn't whether a piece of data might someday be useful, almost anything might be, it's how often you've actually gone looking for data that old, and at what level of detail. Most teams have never once queried logs from six months ago at full per-request granularity.
A rough decision rule for what to sample versus keep whole
Keep every error, every security relevant event, and every request touching payment or auth flows in full, regardless of volume. Sample routine successful requests aggressively once volume is high enough that full retention is clearly the biggest cost driver. Review the rule quarterly rather than setting it once, since what counts as routine shifts as the product changes, and a rule set a year ago may be sampling away exactly the traffic that's become worth watching closely now.
Use these rules to decide what to keep whole:
- Keep every error in full regardless of volume, since errors are rare and carry the most information.
- Keep every security relevant event, and every request touching payment or auth flows, in full.
- Sample routine successful requests aggressively once full retention is clearly the biggest cost driver.
- Review the rule quarterly rather than setting it once, since what counts as routine shifts as the product changes.
A worked example: one service, one setting, most of the bill
Say a single payments service is logging at debug level and accounts for sixty percent of total log volume, while every other service combined accounts for the rest. Dropping that one service to info level in production, keeping debug available on demand for local debugging only, can cut total volume by more than half without touching retention or sampling anywhere else. Before building a broader sampling and tiering strategy, check whether your bill actually has a concentration problem like this one: it's a much faster fix than redesigning the whole pipeline.
What Good Looks Like
Good log aggregation means sampling routine traffic while keeping every error in full, retention tiers matched to actual query patterns, and structured fields instead of free text for anything you'd search during an incident.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
tracking the sampling and retention rule review on a recurring schedule in a tool like ClickUp keeps it from drifting back to 'log everything, just in case'
Vanta can hold the evidence of what's retained and for how long when an auditor asks about your logging and monitoring controls
Frequently Asked Questions
Is reducing log verbosity enough to control costs?
It helps, but sampling matters more. High frequency, low information successful requests are usually the real volume driver, not the level of detail in each line. A team that trims verbosity but logs every successful request in full will still see costs grow faster than traffic.
Should errors ever be sampled the same way as successful requests?
No. Errors are rare and information dense, exactly what you want in full when investigating an incident. Sample the high volume, low information traffic instead, successful requests under normal conditions, and keep every error and every security relevant event at full detail regardless of volume.
How do we know if our retention tiers are set correctly?
Check how often queries actually reach into your warm and cold tiers versus how much you're paying to keep data there. If cold storage is rarely queried but represents a large share of spend, your retention window for hot storage is probably too short, or your cold tier is holding data nobody needs at that granularity.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
Cutting Your Log Aggregation Bill Without Losing the Logs You Need
How to reduce a runaway log aggregation bill without cutting the specific logs you'd actually need during your next real incident.
Where a Log Aggregation Bill Actually Goes, Traced Line by Line
A cost teardown of a typical log aggregation bill, showing which log volume is worth paying for and which is silently expensive debug noise.
Your Log Bill Is Growing Because Nobody Decided What to Keep
Log volume usually grows because every team logs everything by default. Here are three ways to cut the bill without losing the logs you'll actually need.
Cutting Log Aggregation Costs Without Losing Signal
How to cut log aggregation costs with tiered storage, sampling and retention rules, while keeping the logs you need during an incident.
Where Your Log Aggregation Bill Is Actually Going
A worked look at where a log aggregation bill actually comes from, and which cuts save real money without losing the logs you'd need during an incident.
Cutting Log Costs Without Losing What Security Needs
How to reduce a runaway log aggregation bill without deleting the specific log data your security and audit needs actually depend on.