Why Your RAG Infrastructure Drifted From What Terraform Says It Should Be
Infrastructure as code is supposed to mean your Terraform or Pulumi definitions are the single source of truth for what's actually running. In practice, a RAG pipeline's infrastructure, vector database configuration, network rules, inference container settings, drifts from that definition through manual console changes made during an incident, a one-off fix nobody backported into code, or a resource created directly because writing the IaC felt slower that day.
Governance here isn't about writing more Terraform. It's about catching drift before it's the reason a change behaves unexpectedly.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
How does RAG infrastructure drift from its Terraform definition?
During an incident, an engineer opens the cloud console and manually widens a network rule or bumps a vector database's instance size to resolve the immediate problem, then moves on once things are stable. Nobody backports that change into the Terraform definition, because the incident is over and there's a next thing to work on. Weeks later, someone runs a routine plan and it proposes reverting that manual fix back to the old, broken configuration, because as far as the code is concerned, the fix never happened.
How often should you run infrastructure drift detection?
Running a plan only when you're about to make an intentional change means drift sits undetected for however long it's been since the last deliberate change, which for a stable piece of infrastructure could be months. Run a drift detection job on a recurring schedule, independent of whether anyone's actively changing that infrastructure, so a manual change gets flagged and reconciled deliberately, either backported into code or reverted, instead of discovered by accident.
Make the fast path during an incident still go through code
The core problem isn't that engineers make manual changes during incidents, it's reasonable to fix a production issue as fast as possible. It's that the fix doesn't make its way back into IaC afterward. Build incident response practice to include a follow-up step, opening a pull request to codify whatever manual change was made, as a required part of closing out the incident, not an optional cleanup task that competes with the next fire for attention.
For example, an on-call engineer resizes the vector database instance from the console late at night to stop query timeouts, and the incident ticket closes with a note describing what was done. If closing that ticket also requires a pull request that updates the definition, the next plan matches reality. If it does not, the next apply may shrink the instance back to its old size and bring the original outage back, which then looks like a mysterious regression rather than an unrecorded fix. Making the codifying step a required part of incident closure is what prevents that.
Give RAG-specific resources the same review rigor as everything else
Vector database configuration, embedding inference container settings, and network rules for the RAG pipeline specifically sometimes get treated as less critical than core application infrastructure, especially when they were stood up quickly during an initial build. Apply the same IaC review process, pull request, approval, plan review before apply, to these resources as to anything else in production, since a misconfigured vector database or an overly permissive network rule here is just as real a risk as anywhere else.
Watch for state file drift, not just resource drift
Beyond the infrastructure itself drifting from its definition, the IaC state file can drift from reality too, if a resource is deleted outside of Terraform or Pulumi and the state file isn't updated to match, or if two people apply changes to overlapping infrastructure without coordinating and one apply silently clobbers the other's work. Use remote state with locking to prevent concurrent applies from colliding, and periodically reconcile state against actual cloud resources, not only definitions against actual resources.
An IaC governance checklist for RAG infrastructure
- Is drift detection run on a recurring schedule, independent of active changes?
- Does incident response practice include a required follow-up step to codify manual fixes?
- Do RAG-specific resources, vector database, inference containers, get the same review rigor as core application infrastructure?
- Is remote state used with locking to prevent concurrent applies from colliding?
- Has anyone reconciled the state file against actual cloud resources recently, not just definitions against resources?
Make drift visible to the whole team, not just whoever runs the check
A drift detection job that emails its results to one person's inbox tends to produce the same outcome as no drift detection at all once that person is busy or out. Post findings somewhere the whole team sees them, and treat an open drift finding as something with the same visibility and urgency as an open bug, rather than a private notification that quietly ages unread.
What Good Looks Like
Good IaC governance for a RAG system means drift gets detected on a schedule, incident fixes get codified afterward as a required step, and RAG-specific infrastructure gets the same review rigor as everything else.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Documented, enforced change management for infrastructure, exactly what drift detection and a required incident codification step produce, is core evidence Vanta asks for in a SOC 2 audit.
Drata covers the same change-management control category, and a real drift detection process gives you continuously accurate evidence instead of a point-in-time snapshot you have to refresh manually before each audit.
Frequently Asked Questions
Why does our Terraform plan keep proposing to revert a fix we made?
Because the fix was made manually, through the cloud console, during an incident, and never backported into the IaC definition. As far as the code knows, the manual change never happened, so a plan proposes reverting infrastructure back to what the code still says it should be.
How often should we check for infrastructure drift?
On a recurring schedule, independent of whether anyone's actively making changes, since drift can sit undetected for months on stable infrastructure otherwise. Waiting to discover it only when you happen to run a plan means it's been silently accumulating the whole time in between.
Should incident fixes always go through a pull request before applying?
During the incident itself, fixing it fast through the console is reasonable. The gap most teams miss is the follow-up: codifying that manual change into IaC as a required part of closing out the incident, not an optional task that competes with the next fire for attention.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Mapping SOC 2 Controls to a RAG Pipeline's Real Components
SOC 2 auditors ask about access, change management, and vendors in the abstract. Here's what each control actually maps to in a RAG pipeline.
What to Check First in a RAG Pipeline Security Audit
A practical order of operations for auditing a production RAG pipeline: data exposure, prompt injection, access control, logging, and vendor risk.
Where AI Code Review Catches Real Bugs, and Where It Misses
A clear-eyed look at what automated code review reliably catches in pull requests, where it still misses real defects, and how to route the rest to people.
A Go-Live Checklist for Shipping a RAG Pipeline
The specific steps to check before a RAG pipeline goes live: index warm-up, model version pinning, a canary check, and a real rollback plan.
Terraform or Pulumi: Choosing an Infrastructure-as-Code Tool You Won't Rewrite Later
How Terraform's declarative HCL and Pulumi's general-purpose code differ, where each helps governance, and what switching later costs.
Terraform vs. Pulumi for IaC Governance: What Actually Differs in Practice
A practical comparison of Terraform and Pulumi for infrastructure-as-code governance, focused on policy enforcement, drift detection, and team fit.