A FinOps Checklist for Teams Before Their First Big Cloud Bill
Most engineering teams don't have a cost problem so much as a visibility problem: nobody owns cloud spend, so nobody notices when a staging environment scales like production or a data export job runs nightly against a warehouse nobody queries anymore. Cost optimization done well is closer to auditing than engineering, and most of the savings come from the audit, not from rewriting code.
Run this checklist before your cloud bill becomes something a board member asks about, not after. The order matters: teams that jump straight to buying reserved capacity or a FinOps tool before doing the basic cleanup below usually end up locking in the waste instead of removing it.
Tag everything, then find what isn't tagged
You can't optimize spend you can't attribute to a team, service or environment. Before anything else:
- Require a cost-center or team tag on every new resource at creation time, enforced by policy, not by asking nicely
- Pull a report of untagged resources and treat that list as your first cleanup target
- Separate spend by environment (production, staging, dev) so you can see when a non-production environment is costing nearly as much as production
- Tag by service or feature, not just by team, once a team owns more than a handful of resources, so you can see which specific product feature is actually expensive to run
Untagged and overscaled staging environments are one of the most common findings in a first cost audit, because nobody revisits them once they're set up. A load-testing environment spun up for a launch two quarters ago and never torn down is a specific, recurring example: it sits there scaled to handle a traffic spike that happened once, quietly billing every month since.
How do you find cloud resources nobody uses?
Every cloud provider has a version of this list, for instance: unattached storage volumes, idle load balancers, databases with no connections in 30 days, and compute instances running at under 5 percent utilization around the clock. These accumulate because turning something off feels riskier than leaving it running, even when the reason it was created no longer applies.
Set a recurring calendar reminder, monthly is enough, to pull this list and require an owner to justify each item or approve deleting it. Without an owner attached, this list regenerates itself every month.
Before deleting anything flagged as idle, check for a dependency that only runs occasionally, a nightly backup job, a monthly reporting export, a disaster-recovery replica that's supposed to sit unused until it's needed. "Idle for 30 days" and "safe to delete" aren't the same thing, and treating them as the same thing is how a cost cleanup turns into an incident.
Should you right-size before buying reserved capacity?
Reserved instances and committed-use discounts save real money, but only on infrastructure that's actually sized correctly. Committing to a discount on an oversized instance locks in the waste for a year or three instead of removing it. Right-size first (many teams find they can drop one or two instance sizes on database and application tiers without a measurable performance change), then commit.
Say your team is running a $40,000-a-month compute bill on instances sized for a traffic spike that happens twice a year. Right-sizing to typical load and scaling up temporarily for the spike, instead of running spike-sized capacity year-round, is usually a larger saving than any reserved-capacity discount on top of the oversized baseline.
Autoscaling is the other lever here, and it's often misconfigured in the direction of waste: a minimum instance count set high out of caution after a past incident, left unchanged long after the underlying issue was fixed. Revisit autoscaling floors and ceilings whenever you revisit reserved capacity; they tend to drift in the same conservative direction for the same reason.
Five mistakes that undo the savings above
- Optimizing compute while ignoring data transfer and storage costs, which often grow faster than compute as a product matures
- Deleting unused resources without checking whether a batch job or backup depends on them, which turns a cost fix into an incident
- Setting autoscaling limits too conservatively out of caution, which reintroduces the same waste you just removed
- Treating cost optimization as a one-time project instead of a recurring review, so the same waste reaccumulates within two quarters
- Cutting infrastructure spend by reducing observability tooling, which then makes the next incident take longer to diagnose and more expensive to fix
What Good Looks Like
Cost optimization means every resource has a tag, an owner, and a periodic review, so waste gets caught in a monthly pass instead of surfacing as a surprise line item.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How much of a typical cloud bill is usually recoverable on a first pass?
For example, teams running their first real cost audit commonly find 15 to 30 percent of spend attributable to idle, oversized or untagged resources, though this varies widely based on how long the account has existed without review.
Should engineering or finance own cost optimization?
Engineering should own the technical fixes (right-sizing, cleanup, tagging enforcement), while finance or a FinOps function should own the recurring reporting and budget accountability. Splitting ownership either way alone tends to stall.
Is it worth adopting a dedicated FinOps tool for a small team?
For example, below roughly $20,000 to $30,000 a month in cloud spend, your cloud provider's native cost explorer and a monthly tagging review usually cover it. A dedicated tool starts paying for itself once spend and account count grow enough that manual review takes real engineering time.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Three Ways to Cut Cloud Spend, and When Each One Works
Rightsizing, committed-use discounts, and architecture changes all cut cloud spend differently. Here's how to pick the right one for your situation.
A CTO's Framework for Cutting Infrastructure Costs
A decision framework for engineering leaders trying to cut cloud and tooling spend without slowing the team down or cutting into future capacity.
Build vs. Buy for Your Security Tooling Stack
A decision framework for when to build DevSecOps tooling in-house versus buying a platform, based on team size, maintenance burden and audit needs.
Where Real-Time Pipeline Costs Actually Come From
The levers that actually move a streaming pipeline's bill: retention, replication, over-provisioned consumers, and cross-zone network traffic.
The Real Cost Drivers in a RAG Pipeline
Embedding calls, index storage, reranking, and padded context each drive RAG cost differently. Here's where to look first before cutting spend.
Cutting the Cost of Running LLM Agents at Scale
Where agent spend actually goes, and the specific changes, not just a cheaper model, that bring the bill down without cutting quality.