Production RAG & Vector Data ArchitecturePlaybook3 min readUpdated September 2026

Why Your RAG Infrastructure Drifted From What Terraform Says It Should Be

Infrastructure as code is supposed to mean your Terraform or Pulumi definitions are the single source of truth for what's actually running. In practice, a RAG pipeline's infrastructure, vector database configuration, network rules, inference container settings, drifts from that definition through manual console changes made during an incident, a one-off fix nobody backported into code, or a resource created directly because writing the IaC felt slower that day.

Governance here isn't about writing more Terraform. It's about catching drift before it's the reason a change behaves unexpectedly.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

How does RAG infrastructure drift from its Terraform definition?

During an incident, an engineer opens the cloud console and manually widens a network rule or bumps a vector database's instance size to resolve the immediate problem, then moves on once things are stable. Nobody backports that change into the Terraform definition, because the incident is over and there's a next thing to work on. Weeks later, someone runs a routine plan and it proposes reverting that manual fix back to the old, broken configuration, because as far as the code is concerned, the fix never happened.

How often should you run infrastructure drift detection?

Running a plan only when you're about to make an intentional change means drift sits undetected for however long it's been since the last deliberate change, which for a stable piece of infrastructure could be months. Run a drift detection job on a recurring schedule, independent of whether anyone's actively changing that infrastructure, so a manual change gets flagged and reconciled deliberately, either backported into code or reverted, instead of discovered by accident.

Make the fast path during an incident still go through code

The core problem isn't that engineers make manual changes during incidents, it's reasonable to fix a production issue as fast as possible. It's that the fix doesn't make its way back into IaC afterward. Build incident response practice to include a follow-up step, opening a pull request to codify whatever manual change was made, as a required part of closing out the incident, not an optional cleanup task that competes with the next fire for attention.

For example, an on-call engineer resizes the vector database instance from the console late at night to stop query timeouts, and the incident ticket closes with a note describing what was done. If closing that ticket also requires a pull request that updates the definition, the next plan matches reality. If it does not, the next apply may shrink the instance back to its old size and bring the original outage back, which then looks like a mysterious regression rather than an unrecorded fix. Making the codifying step a required part of incident closure is what prevents that.

Give RAG-specific resources the same review rigor as everything else

Vector database configuration, embedding inference container settings, and network rules for the RAG pipeline specifically sometimes get treated as less critical than core application infrastructure, especially when they were stood up quickly during an initial build. Apply the same IaC review process, pull request, approval, plan review before apply, to these resources as to anything else in production, since a misconfigured vector database or an overly permissive network rule here is just as real a risk as anywhere else.

Watch for state file drift, not just resource drift

Beyond the infrastructure itself drifting from its definition, the IaC state file can drift from reality too, if a resource is deleted outside of Terraform or Pulumi and the state file isn't updated to match, or if two people apply changes to overlapping infrastructure without coordinating and one apply silently clobbers the other's work. Use remote state with locking to prevent concurrent applies from colliding, and periodically reconcile state against actual cloud resources, not only definitions against actual resources.

An IaC governance checklist for RAG infrastructure

  • Is drift detection run on a recurring schedule, independent of active changes?
  • Does incident response practice include a required follow-up step to codify manual fixes?
  • Do RAG-specific resources, vector database, inference containers, get the same review rigor as core application infrastructure?
  • Is remote state used with locking to prevent concurrent applies from colliding?
  • Has anyone reconciled the state file against actual cloud resources recently, not just definitions against resources?

Make drift visible to the whole team, not just whoever runs the check

A drift detection job that emails its results to one person's inbox tends to produce the same outcome as no drift detection at all once that person is busy or out. Post findings somewhere the whole team sees them, and treat an open drift finding as something with the same visibility and urgency as an open bug, rather than a private notification that quietly ages unread.

Executive Capability Standard

What Good Looks Like

Good IaC governance for a RAG system means drift gets detected on a schedule, incident fixes get codified afterward as a required step, and RAG-specific infrastructure gets the same review rigor as everything else.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Understand the specific way drift accumulates, manual incident fixes that never get backported, so you know what governance practice actually needs to catch.
2. Do Manually:Run a manual drift detection check periodically by hand, comparing actual infrastructure against your IaC definitions, even before automating it.
3. Delegate:Make codifying a manual fix a required, assigned step in incident close-out, with an owner, not an optional follow-up.
4. Automate:Schedule automated drift detection and alert when actual infrastructure diverges from IaC definitions, independent of active change activity.
5. Buy:Use a managed drift detection tool if your IaC platform or provider offers one instead of building your own scheduled comparison job.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Frequently Asked Questions

Why does our Terraform plan keep proposing to revert a fix we made?

Because the fix was made manually, through the cloud console, during an incident, and never backported into the IaC definition. As far as the code knows, the manual change never happened, so a plan proposes reverting infrastructure back to what the code still says it should be.

How often should we check for infrastructure drift?

On a recurring schedule, independent of whether anyone's actively making changes, since drift can sit undetected for months on stable infrastructure otherwise. Waiting to discover it only when you happen to run a plan means it's been silently accumulating the whole time in between.

Should incident fixes always go through a pull request before applying?

During the incident itself, fixing it fast through the console is reasonable. The gap most teams miss is the follow-up: codifying that manual change into IaC as a required part of closing out the incident, not an optional task that competes with the next fire for attention.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides