Cloud FinOps & Infrastructure ScalingPlaybook3 min readUpdated September 2026

The IaC Setup That Works Until Someone Changes Something by Hand

Infrastructure as code makes a specific promise: the code is the source of truth, and running it reproduces your infrastructure exactly. That promise holds right up until someone, usually under time pressure during an incident, makes a quick manual change in the cloud console to fix something immediately, and doesn't circle back to update the code. From that moment, the code and reality have quietly diverged.

The discipline that actually matters isn't writing the IaC in the first place. It's catching and closing that gap before it compounds.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Drift starts small and compounds silently

One manual change during an incident is understandable and sometimes necessary. The problem is what happens next: nobody updates the code to match, the next person who runs the IaC either reverts the fix unknowingly or the tool reports a change it doesn't actually understand, and now there's uncertainty about what the code will do the next time it runs. Left unaddressed, drift accumulates until the code is no longer a reliable description of what's actually running.

A simple rule keeps one-time exceptions from becoming ongoing drift: any manual console change made during an incident gets a follow-up ticket before the incident is closed. The ticket names who will update the code, and by when. For example, if an engineer opens a security group to restore access at night, the morning handoff includes a task to encode that rule in the code or remove it. This costs a few minutes per incident, and it means the next full reapply is a routine event rather than a surprise that quietly reverts something still in use.

How do you detect infrastructure drift before it bites?

Run a drift detection check on a recurring schedule, comparing what the code declares against what's actually deployed, rather than discovering drift the hard way when an unrelated change to the same resource produces an unexpected result. Most IaC tools support a plan or diff mode that surfaces this without applying anything. Treat any detected drift as something to investigate and resolve deliberately, either by updating the code to match reality or reverting the manual change, not something to silently accept.

A repeatable routine for closing drift looks like this:

  1. Run a scheduled plan or diff that compares what the code declares against what is actually deployed, without applying anything.
  2. Triage each difference by what it touches, so a production firewall rule gets attention before a development resource tag.
  3. Resolve each item deliberately, either by updating the code to match reality or by reverting the manual change.
  4. Before closing any incident that involved a console change, record it and update the code so the exception does not become permanent.

How should you manage the IaC state file?

The state file tracking what the tool believes is deployed can itself become corrupted, out of sync, or locked by a crashed process mid-run, and a state file problem can be more disruptive than the infrastructure issue that triggered the original manual change. Back up state before risky operations, use remote state with locking to prevent two people running changes simultaneously, and know the recovery procedure for a corrupted or lost state file before you need it during an actual incident.

Module structure determines how much a mistake costs

A monolithic configuration where every resource lives in one giant file means a small, well-intentioned change carries the blast radius of the entire infrastructure. Breaking configuration into smaller, focused modules, networking separate from compute separate from data stores, limits how much any single change can affect, and makes it possible to review and approve changes to one area without touching everything else at once.

Tie your change review process to your actual downtime tolerance

A 99.95% uptime target leaves roughly 4.38 hours a year of downtime budget across every incident combined1, which is a useful anchor for how much review a given infrastructure change deserves: a change to a production database's configuration warrants more scrutiny than a change to a development environment's logging settings, and the review process should reflect that difference explicitly rather than treating every pull request the same.

A worked example: the drift that broke a rollback

Say a security group was manually adjusted during an incident to allow emergency access, and the code was never updated to match. Six months later, an unrelated infrastructure change triggers a full reapply, which silently reverts that manual adjustment because the code never knew about it. If that manual change had actually been load-bearing for something still running, the rollback becomes a new incident, caused entirely by drift nobody tracked down after the original fix.

Common mistake: treating every drift the same way

Not all drift carries the same risk: a manually adjusted tag on a development resource is very different from a manually adjusted firewall rule on a production database, and treating both with the same urgency either wastes time chasing harmless drift or, worse, lets genuinely risky drift sit unresolved because it's buried in a long list of low-stakes items. Triage detected drift by what it actually touches before deciding how fast it needs to be reconciled.

Executive Capability Standard

What Good Looks Like

Good infrastructure as code means scheduled drift detection, a documented state file recovery procedure, modules scoped narrowly enough to limit blast radius, and review scrutiny that scales with what a given change could actually break.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Run a drift detection check today and see how far your current code has already diverged from what's actually deployed.
2. Do Manually:Manually reconcile any drift you find, deciding case by case whether to update the code or revert the manual change.
3. Delegate:Assign an owner for the state file backup and recovery procedure so it's documented before it's needed during an incident.
4. Automate:Automate scheduled drift detection so divergence surfaces on a regular cadence instead of being discovered by accident.
5. Buy:Bring in outside expertise to restructure a monolithic configuration into scoped modules if a single change currently touches far more than it should.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Frequently Asked Questions

Is a manual change during an incident always a mistake?

No, sometimes it's the fastest way to stop an active incident, and that's a reasonable tradeoff in the moment. The mistake is not circling back afterward to update the code to match, which is what turns a one-time exception into ongoing, compounding drift.

How often should we run drift detection?

On a recurring schedule, daily or weekly depending on how frequently your infrastructure changes, rather than only when something unexpected happens. Scheduled detection catches drift while it's still small and easy to reconcile, instead of discovering it months later during an unrelated change.

What's the safest way to structure IaC for a growing team?

Break configuration into smaller, focused modules by function, networking, compute, data stores, rather than one large configuration covering everything. That structure limits the blast radius of any single change and makes it possible to review changes to one area without needing full context on the whole system.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides