Building a CI/CD Pipeline for Agent and Tool Code
Prompts and tool definitions change more often than most application code, and they're often edited by people who aren't running the full test suite in their head the way an engineer would for a code change. A CI/CD pipeline for an agentic system needs to treat prompt and tool changes as first-class deploy artifacts, not as configuration that slips past review.
Version prompts and tool schemas the same way you version code
Store prompts and MCP tool definitions in the same repository as the code that uses them, under normal version control, rather than in a separate content management tool that bypasses code review. This gives you a diff on every change, a blame history when something regresses, and the ability to roll back a bad prompt exactly the way you'd roll back a bad deploy.
A practical detail is to keep the prompt text in its own file, not inline in application code, so a wording change produces a small readable diff, and to require the same reviewer sign-off you would ask for on a code change. Tag each release with the prompt and tool schema versions it shipped with, so a trace from production can be matched to the exact text that produced it. When a regression appears, that pairing turns a vague suspicion about the agent into a specific commit to inspect or revert.
How do you run an evaluation suite as a required CI check?
Before a prompt, tool, or model change merges, run it against your evaluation set automatically and block the merge if scores drop below a set threshold. This is the agent-specific equivalent of a unit test suite, and it catches the class of regression that's otherwise invisible until a customer hits it, an agent that quietly got worse at a task it used to handle well.
How should you deploy prompt and tool changes gradually?
Roll a new prompt or tool version out to a small percentage of traffic first, watch tool error rates and fallback rates, then expand. Deployment frequency is a good proxy for how well this discipline is actually working: teams with strong practices ship on demand in small increments, while the slowest go as long as 180 days between releases1. The gap between those two isn't really about tooling, it's about whether small, frequent, closely watched changes are the default or the exception.
Keep a fast, clean rollback path
Because a bad prompt change degrades quality rather than crashing, it can take longer to notice than a normal outage, which makes a fast rollback path even more important. Keep the previous version of every prompt and tool definition one command away from being restored, and practice the rollback occasionally so it's not the first time anyone's run it during an actual incident.
A rollback drill costs an hour every quarter or so and consistently turns up something worth fixing: a script that assumed a manual step nobody remembered, a dependency that wasn't actually pinned to the old version along with the prompt. Finding that during a drill is a minor annoyance; finding it during a real incident is a much longer night.
A pipeline for prompt and tool changes runs in this sequence:
- Keep prompts and tool schemas in the same repository as the code, under normal version control and code review.
- Run the evaluation suite automatically and block the merge if scores fall below your threshold.
- Roll the change out to a small percentage of traffic and watch tool error rates and fallback rates against the control group.
- Expand in stages only if those signals stay normal, and roll back with a single command if they do not.
- Practice the rollback occasionally so it is not the first time anyone has run it during a real incident.
A worked example: catching a bad prompt change before it spread
Say a new prompt version merges after passing the evaluation suite, and rolls out to five percent of traffic as the pipeline requires. Within the first hour, tool error rates for that slice tick up noticeably compared with the control group still on the old prompt, even though the evaluation suite hadn't caught anything wrong. The gradual rollout, not the evaluation suite, is what caught this one: the new prompt had changed how the agent formatted a date field passed into a downstream tool, a case the evaluation set's examples happened not to cover.
Because only five percent of traffic was affected and the rollback took a single command, the fix cost a few dozen confusing tool calls instead of a day's worth across the whole user base. The date-formatting case was added to the evaluation set immediately afterward, closing the specific gap for next time. None of this required anything exotic, just the discipline of watching the right metrics during a small, deliberate rollout instead of trusting a passing test suite to be the final word. A passing evaluation suite tells you the cases you thought to test still work; a gradual rollout with real monitoring tells you about the cases you didn't think to test, which is a different and equally necessary kind of coverage.
What Good Looks Like
A solid CI/CD setup for agentic systems versions prompts and tool schemas as code, blocks merges that regress the evaluation suite, and rolls changes out gradually with a fast, practiced rollback path.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Should prompts live in code or in a separate content tool?
In code, under the same version control and review process as everything else that ships. A separate content tool that bypasses code review is convenient at first but removes the diff history and rollback safety you'd expect for any other production change.
What should block a prompt or tool change from merging?
A drop in your evaluation suite score below a set threshold, the same way a failing test blocks a normal code merge. Without this check, a prompt change that quietly makes the agent worse at a task can ship and go unnoticed for weeks.
How gradually should a new prompt version roll out?
Start with a small percentage of traffic, watch error and fallback rates for a defined window, then expand in stages rather than all at once. Treat it with the same caution as any change that's hard to notice going wrong through normal monitoring alone.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.
Related Guides
Rolling Out Agentic Workflows Without Breaking Production
A practical rollout checklist for shipping an AI agent to production, from a shadow-mode test run through the guardrails that catch it if it misbehaves.
How to Know If Your Agent Is Actually Working
Building an evaluation framework for an AI agent, from the first small test set through catching quality regressions before customers do.
Build vs. Buy for Verifying Every Device That Connects In
What zero-trust device and identity verification actually requires, what a platform gives you over a homegrown check, and how to decide between them.
Building a CI/CD Pipeline That Actually Catches Bugs
How to build a pipeline that blocks real regressions instead of just style errors, from test selection to what actually belongs as a merge gate.
Building a CI/CD Pipeline That Tests Models, Not Just Code
How to extend CI/CD for AI model serving so a prompt or model change is evaluated automatically, not just checked for syntax.
A Worksheet for Sizing Your CI Pipeline's Real Cost
A step-by-step worksheet for pricing out what your automated test pipeline actually costs in compute and engineering wait time, and where to trim it.