Production RAG & Vector Data ArchitecturePlaybook3 min readUpdated September 2026

The Feature Flags Nobody Remembers Turning On

A RAG pipeline tends to accumulate feature flags faster than a typical application: one for a new embedding model rollout, one for a reranker experiment, one for a prompt variant, one for a retrieval parameter test, each added for a specific, time-bound reason and then left in place after the decision was made.

Flag hygiene isn't about having fewer flags. It's about knowing, for every flag that exists, why it's still there and what happens if you remove it.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

How should you tag each flag with a removal condition?

A flag created to test a new reranker should be tagged not just with what it does, but with what has to be true for it to be safe to remove: the experiment concluded, a winner chosen, the losing path deleted. Without that condition written down at creation time, a flag's purpose fades from memory long before the flag itself gets removed, and eventually nobody feels confident deleting something they don't fully understand anymore.

Which feature flags have quietly become permanent configuration?

A flag meant to gate a temporary experiment sometimes becomes the de facto way a permanent decision gets made, an embedding model choice or a retrieval parameter that's actually fixed in practice but still lives behind a flag because removing the flag means touching code nobody wants to touch. If a flag's value hasn't changed in months and isn't expected to, it's not a flag anymore, it's configuration wearing a flag's clothing, and it should be simplified into an actual config value or removed.

Check what happens when two RAG-related flags interact

A flag controlling which embedding model is active and a flag controlling which reranker is active can combine into states nobody tested: the new embedding model with the old reranker, or the reverse, if both flags can be toggled independently. Map out which combinations of your active flags are actually valid, and either prevent invalid combinations in code or explicitly test them, rather than assuming flags that were each tested individually are safe together.

For example, suppose one flag switches between an old and a new embedding model, and another switches between an old and a new reranker. Each was tested on its own, but the new embedding model paired with the old reranker was never run together, and that pairing appears whenever the two flags are rolled out at different speeds. A simple fix is to write down the valid combinations, reject the rest in code, and add a test for each allowed pair. That takes a small amount of effort and removes a whole class of surprises that no single flag review would ever catch.

Clean up losing variants completely, not just the flag

When an experiment concludes and a variant loses, removing the flag without removing the losing code path leaves dead logic that still has to be read, understood, and accounted for by anyone touching that part of the pipeline later. Treat flag removal as removing the flag and the losing branch together, in the same change, so the codebase actually gets simpler after an experiment ends instead of just gaining one more piece of history nobody prunes.

Review the full flag list on a schedule, not just when someone notices

Set a recurring review, monthly or quarterly depending on how fast your team adds flags, where someone goes through the full active list and asks, for each one, whether its removal condition has been met. Flags that survive several reviews without their condition being met are worth a direct question to whoever owns them: is this actually still an open experiment, or has it just been forgotten.

A feature flag hygiene checklist for RAG pipelines

  • Does every flag have a written removal condition from the day it's created, not just a description of what it does?
  • Are any flags actually functioning as permanent configuration rather than a temporary experiment?
  • Have the valid and invalid combinations of interacting flags been mapped and tested?
  • Does removing a flag also remove the losing code path, in the same change?
  • Is there a recurring, scheduled review of the full active flag list?

Name an owner for every flag, not just the review process

A flag list reviewed by a team collectively tends to produce the same outcome every time: everyone assumes someone else knows why a given flag is still there, and nobody removes it. Assign each flag an individual owner at creation time, the person who can actually answer whether its removal condition has been met, so the recurring review has someone specific to ask instead of a group shrug. When that person leaves the team or changes roles, reassign their flags explicitly rather than letting ownership quietly lapse along with them.

Executive Capability Standard

What Good Looks Like

Good flag hygiene for a RAG pipeline means every active flag has a written removal condition, interacting flags have been tested together, and removed flags take their losing code path with them.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Understand the specific ways RAG flags tend to go stale, becoming permanent configuration or combining in untested ways, before setting a review process.
2. Do Manually:Go through your full active flag list by hand at least once, checking each one against a written or reconstructed removal condition.
3. Delegate:Assign one person ownership of the flag review schedule and the question of whether stale flags' owners still consider them active experiments.
4. Automate:Add tooling that flags configuration values that haven't changed in a set period as candidates for simplification out of the flagging system.
5. Buy:Use your feature flag platform's built-in stale-flag detection and usage reporting instead of tracking flag age and combinations manually.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Frequently Asked Questions

Why do RAG pipelines accumulate more feature flags than typical applications?

Because embedding models, rerankers, prompts, and retrieval parameters each get their own experiments, and each experiment tends to start with its own flag. Without a written removal condition at creation time, these flags outlive the experiments they were built for far more easily than in a typical application.

How do you know if a feature flag should actually be removed?

Check whether its written removal condition has been met, the experiment concluded and a winner chosen. If a flag's value hasn't changed in months and isn't expected to, it's functioning as permanent configuration and should be simplified into an actual config value or removed entirely.

What's the risk of removing a flag without removing the losing code path?

The codebase doesn't actually get simpler, it just gains dead logic that everyone touching that area later still has to read and reason about. Treat flag removal and removing the losing branch as one change, not two separate cleanup tasks where the second one never gets around to happening.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides