Cutting Log Costs Without Losing What Security Needs
Log aggregation bills tend to grow quietly until someone notices they now rival the infrastructure spend they're supposed to help monitor. The instinct is to cut volume across the board, which is exactly how a team accidentally deletes the one log category a security review needed six months later. A better approach treats different log categories differently instead of applying one retention and sampling policy to everything.
Should every log category get the same retention and sampling?
Security relevant events, authentication attempts, permission changes, access to sensitive data, need longer retention and closer to full capture, because you can't predict which one will matter for an investigation months from now. High volume, low signal logs, routine health check pings or verbose debug output from a stable service, are the better candidates for aggressive sampling or a short retention window. Sorting your log sources into these categories before touching any settings prevents an across the board cut from quietly gutting your audit trail. Write the category assignments down somewhere durable, not just in someone's head, so the next person adjusting retention settings doesn't have to rediscover which sources are load-bearing for security.
Sample debug and info level logs, not security or audit events
Sampling, keeping a representative fraction of high volume log lines instead of every single one, works well for routine application chatter where the pattern matters more than any single line. It works badly for security and audit events, where a single dropped entry is exactly the one an investigation needs. Set sampling rates per log level and per source, not as one global percentage, so a debug log flood gets trimmed while an authentication log stays complete.
For example, suppose a stable service writes a debug line for every request, and it makes up most of your log volume. Sampling that source to keep a representative fraction cuts cost with little loss, because the pattern matters more than any single line. Authentication events from the same service should stay fully captured. The common mistake is setting one global sampling percentage, which trims the noisy source but also drops the security events an investigation will need. Set rates per source and per level, and write the reasoning next to the setting so the next person keeps the distinction.
Push cold logs to cheaper storage instead of deleting them outright
A tiered retention model, recent logs in your primary aggregation platform where they're fast to query, older logs moved to cheaper object storage where they're still retrievable but not actively indexed, usually cuts cost more than shortening retention outright. This matters most for audit and security logs, where a compliance requirement or a customer contract might specify keeping data for a year or more even though nobody expects to query most of it after the first few weeks.
Structure logs so filtering doesn't require reading everything first
Unstructured text logs force your aggregation platform to index and store the full line even when you only ever query a few fields out of it. Structured, field based logging lets you drop or sample on specific fields, keep the error field but sample the request body, for instance, which gives you far more precise control over cost than an all or nothing decision per log line. This is also the change that makes actual incident investigation faster, since a structured query beats grepping through raw text either way.
How do you find the log source that is driving your bill?
A surprising share of runaway log volume traces back to a single misconfigured or forgotten source: a debug flag left on in production, a library that logs every request at info level by default, or a duplicate log shipper sending the same events twice after a migration. Reviewing your top volume sources by percentage, not just total bytes, tends to surface these quickly, since the offending source is often a small piece of infrastructure producing a disproportionate share of total volume. Make this review a standing item on your monthly infrastructure check rather than something only prompted by a surprising invoice, since the flag or duplicate shipper is usually cheap to fix once someone actually notices it.
Review the policy against an actual investigation, not just the bill
Once a year, or after any real security incident, walk through what you'd need to investigate a hypothetical breach using only the logs your current retention and sampling policy would have kept. If the answer is that a key piece of evidence would already be gone or sampled out, that's the signal to adjust the policy for that specific log category, even if it costs more, rather than discovering the gap during a real investigation.
Steps for cutting log cost safely:
- Sort log sources into security relevant and high volume, low signal categories, and write the assignments down.
- Sample debug and info logs per level and per source, and never sample authentication or audit events.
- Move older logs to cheaper object storage instead of deleting them, especially audit and security logs.
- Use structured logging so you can drop or sample specific fields rather than whole lines.
- Review your top sources by share of total volume, looking for forgotten debug flags or duplicate shippers.
What Good Looks Like
A good log strategy keeps security and audit events complete while sampling or tiering high volume, low signal logs, based on a written policy per log category rather than one global retention rule.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Is it safe to sample security and audit logs the same way we sample debug logs?
No. Sampling works for high volume, low signal logs where a representative fraction still shows the pattern, but a security or audit trail needs to be complete, since any single dropped event could be the one that matters during an investigation.
What's the fastest way to find where a log aggregation bill is actually going?
Review your top sources by percentage of total volume, not just raw byte count. A single misconfigured or forgotten debug flag is a common cause of runaway volume and usually shows up quickly once you sort sources that way.
Does moving old logs to cheaper storage hurt our ability to investigate an incident?
Not if it's done as a tiering strategy rather than deletion: the logs stay retrievable, just not actively indexed for fast search. It does mean a query against old data takes longer, which is a reasonable tradeoff for logs you rarely need but still have to retain.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Continuous Device Verification for a Zero-Trust API
How continuous device and identity verification actually works in a zero-trust architecture, and where to draw the line for a small engineering team.
Cutting Your Log Aggregation Bill Without Losing the Logs You Need
How to reduce a runaway log aggregation bill without cutting the specific logs you'd actually need during your next real incident.
Where a Log Aggregation Bill Actually Goes, Traced Line by Line
A cost teardown of a typical log aggregation bill, showing which log volume is worth paying for and which is silently expensive debug noise.
Why Your Log Bill Grows Faster Than Your Traffic
Log volume usually grows faster than the traffic producing it. Where that gap actually comes from, and the retention and sampling changes that close it.
Cutting Log Aggregation Costs Without Losing Signal
How to cut log aggregation costs with tiered storage, sampling and retention rules, while keeping the logs you need during an incident.
Where Your Log Aggregation Bill Is Actually Going
A worked look at where a log aggregation bill actually comes from, and which cuts save real money without losing the logs you'd need during an incident.