API Security, Identity & Zero-TrustPlaybook3 min readUpdated September 2026

Terraform vs Pulumi: A Governance Model That Won't Slow You Down

Most infrastructure-as-code programs don't fail because a team picked the wrong tool. They fail because nobody owns the policy layer: what a `terraform apply` is allowed to touch, who reviews it, and what happens when a module drifts from what's actually running in the account.

Terraform and Pulumi both solve the provisioning problem well. Governance is a separate layer you build on top of either one, and the choice between them changes how that layer gets built, not whether you need it.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Terraform vs Pulumi: what the choice actually changes

Terraform uses its own declarative language, HCL, and a large, mature provider ecosystem; almost anything you'd provision already has a Terraform provider. Pulumi lets you write infrastructure in TypeScript, Python, Go, or another general-purpose language, so you get real loops, conditionals, and unit tests instead of HCL's more limited expressions.

If your platform team already lives in TypeScript, Pulumi cuts the context switch. If your infrastructure is mostly standard cloud resources with little custom logic, HCL's constraints rarely bite, and Terraform's larger community means more prior art for edge cases. Neither tool changes your governance obligations: state locking, plan review, and policy checks are still your job either way.

Where the budget actually leaks: drift, orphans, and duplicate modules

Say your team runs $40,000 a month in cloud spend and nobody's audited what's attached to a Terraform or Pulumi stack in the last two quarters. That's a common setup for waste: a staging cluster a deprecated module never tore down, a load balancer created by hand during an incident and never imported into state, three near-identical VPC modules because nobody wanted to touch the one already in prod.

None of it shows up as a single line item. It surfaces only when someone finally runs a drift-detection pass against the live account and finds a cloud bill much higher than the workload actually needs.

Four steps to add policy-as-code without blocking ships

  • Inventory first. Run `terraform plan` (or `pulumi preview`) against every stack and list anything showing unexpected drift before you write a single policy.
  • Pick a policy engine that matches your tool: Open Policy Agent or Conftest for either, Sentinel if you're on Terraform Cloud, or Pulumi's own CrossGuard for Pulumi stacks.
  • Ship policies in warn-only mode for two to three weeks. Log every violation without blocking a merge, and use that log to find rules too strict for how your team actually works.
  • Flip to blocking, one policy at a time, starting with the highest-risk category: public security groups, unencrypted storage, IAM roles with wildcard permissions.

DORA's deploy-frequency data and why governance can't just mean slower

DORA groups engineering orgs into four performance clusters by deployment frequency: the top cluster ships on demand, while low performers can go up to 180 days between releases1. A governance layer that adds a manual approval step to every infrastructure change is a fast way to drift toward that low end.

The fix isn't skipping review, it's moving the review earlier: catch a bad security group in a policy check that runs in seconds during CI, not in a human approval queue that runs once a week.

Common mistakes that turn governance into a bottleneck

  • Writing one global policy set for every environment, so a rule tuned for production blocks a throwaway sandbox stack.
  • Locking down apply access before anyone has mapped which resources are actually safe to change without review.
  • Skipping an exception path, so an on-call engineer fixing a live incident has no way to bypass a policy check that's correctly blocking a routine change but wrongly blocking an emergency one.
  • Buying a policy engine before writing a single rule by hand, so nobody on the team actually understands what the tool is enforcing when it eventually blocks something during an incident.

What a healthy governance layer looks like a year in

The policies that survive past the first quarter tend to be the small set that map directly to something that already caused an incident: a public bucket, an overly permissive IAM role, a security group open to the world. Teams that start instead with a large, generic ruleset borrowed from a compliance framework tend to either abandon it under noise or leave it in warn-only mode indefinitely because nobody has the time to review the backlog of violations.

A useful signpost that governance is working: the policy engine's blocked-attempt log is short and each entry is something an engineer recognizes as a real near-miss, not a wall of false positives nobody reads. If the log is long and nobody reads it, the rules need tuning before they need enforcing more strictly.

Executive Capability Standard

What Good Looks Like

Good IaC governance means every infrastructure change goes through the same reviewed, versioned path as your application code, with automated policy checks that run before an apply, not a postmortem after one.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Run a drift-detection pass against your live accounts and list every resource that isn't accounted for in Terraform or Pulumi state.
2. Do Manually:Require a second engineer to review the plan output on every apply that touches production, and keep a shared log of what got approved and why.
3. Delegate:Give one platform engineer ownership of the module library and the policy rules, so standards don't drift between teams.
4. Automate:Wire a policy engine like OPA, Conftest, or Sentinel into CI so a noncompliant plan fails the build instead of waiting for a reviewer to catch it.
5. Buy:A fractional platform lead earns their cost once you're standardizing IaC governance across multiple teams without the bandwidth to build the tooling in-house; a compliance automation platform can also generate the audit evidence that your policies actually ran.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Frequently Asked Questions

Do we need to migrate from Terraform to Pulumi to get good governance?

No. Governance is a policy-as-code layer you add on top of either tool with something like OPA, Conftest, or Sentinel. Migrating languages only makes sense if Terraform's HCL is genuinely blocking you, for example if you need complex conditional logic or want infrastructure code under the same test suite as your application.

How do we stop state drift without slowing down releases?

Run drift detection on a schedule, such as nightly, rather than gating every deploy on it. Report drift to the owning team with a window to reconcile it, and reserve blocking checks for the small set of policies where an out-of-band change is genuinely dangerous, like a public storage bucket.

What's the fastest way to add policy checks to an existing setup?

Start with Conftest or OPA against your plan output in CI, running in warn-only mode. Pick two or three high-risk rules, like open security groups or unencrypted volumes, and get those clean before adding more. Flip to blocking only after a couple of weeks with no false positives.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.

Related Guides