A Checklist for Cleaning Up Feature Flags Before They Become Their Own Codebase
Clean up a stale feature flag by confirming it has been stable, finding every reference, ruling out kill switches, then removing the flag and its dead branch together. A flag that outlives its purpose is still evaluated on every request, can still be flipped by accident, and adds a path every future engineer has to reason about.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Find flags that have been fully on or fully off for a set period
A flag that's been fully rolled out, or fully rolled back, for a defined stretch of time, say a full quarter, with no changes, is a strong candidate for removal: the decision it was meant to test has effectively already been made. Pull this list directly from your flag management tool's history rather than relying on memory of which flags are "probably done."
The pitfall here is treating a long stretch of stability as proof the flag is safe to remove immediately. It's a strong signal, not a guarantee; check the next few items before deleting anything.
How do you confirm a flag isn't still referenced elsewhere?
Search the codebase for the flag's key, not just in the feature it was originally built for. Flags have a way of getting reused or checked in unrelated code paths added later by someone who didn't know the flag's original intent, and removing it without finding every reference can silently change behavior somewhere nobody's watching.
This is also where old, unrelated experiments sometimes surface: a flag added for a specific A/B test that got wired into an analytics event or a logging condition long after the original test ended, in a way that isn't obvious from the flag's name alone.
Remove the flag and its dead branch in one focused change
Once you're confident a flag is safe to remove, delete both the flag definition and the code path it disabled, not just one or the other. Removing the flag but leaving dead code behind, or leaving the flag defined but unused, both create a smaller version of the same confusion the cleanup was meant to fix.
Keep this as its own pull request, separate from any feature work happening at the same time. A flag removal mixed into an unrelated change is much harder to review carefully, and careful review is exactly what this step needs.
The safe removal order, start to finish:
- Confirm the flag has been fully on or fully off, with no changes, for a defined stretch such as a quarter.
- Search the codebase for the flag key everywhere, not only in the feature it was built for.
- Separate genuine leftovers from deliberate kill switches, and tag the kill switches so they are not removed by mistake.
- Remove the flag definition and its dead code path together in one focused pull request, kept separate from feature work.
Is the flag still serving as a kill switch?
Some flags that look fully and permanently rolled out are deliberately kept as an operational kill switch, a way to quickly disable a feature if something goes wrong, even though the feature itself is fully launched. Removing one of these isn't a cleanup, it's removing a safety mechanism, often without anyone realizing until the next time it's needed and isn't there.
Tag these explicitly in your flag management tool so they're excluded from routine cleanup sweeps, and review them on a separate, longer cadence instead of deleting them alongside genuinely stale flags.
Build the review into a recurring cadence, not a one-time purge
A single big cleanup effort feels satisfying but the flag count creeps back up the same way it did the first time unless the review becomes routine. Add a short flag review to a regular cadence, monthly or quarterly depending on how quickly your team ships flagged changes, and treat a growing count of stale flags as a signal the review cadence itself needs to tighten, not just a backlog to eventually get to.
Naming and defaults that make the next cleanup easier
A flag named after the specific decision it supports, with an expected removal date set at creation time, is far easier to evaluate six months later than a flag named after a generic feature with no context attached. Make an expiry date a required field when a flag is created, even if it's just a rough estimate, since a flag with no expectation of when it should be revisited tends to be the one still sitting there years later.
This small habit at creation time does more to keep the flag count manageable long-term than any single cleanup sweep, because it turns "forgot this existed" into "this is overdue," which is a much easier thing to act on.
What Good Looks Like
Good feature flag hygiene means every flag has a known owner and purpose, stable flags get reviewed for removal on a schedule, and kill switches are explicitly distinguished from flags that are simply done.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Industry-leading platform for Enterprise DevSecOps: Feature Flag Cleanup Protocols.
A flag left flippable in production longer than it should be is also a lever an attacker with any foothold could pull; CrowdStrike's runtime visibility is worth having on unusual flag-driven behavior changes, not just on the cleanup process itself.
Frequently Asked Questions
How many feature flags is too many for a small team?
There's no fixed number that applies to every team, but if nobody can quickly explain what most of the active flags do, that's a clearer signal than any count. A smaller number of well-understood flags is healthier than a larger number nobody fully tracks.
Should flag cleanup be assigned to whoever created the flag?
Not necessarily; that engineer may have moved to a different project or left the team. Assign cleanup ownership to whoever currently owns the feature area the flag touches, since they're best positioned to confirm it's genuinely safe to remove.
What's the safest order to remove a flag in?
Confirm the rollout has been stable, search for all references, distinguish kill switches from genuinely done flags, then remove flag and dead code together in a dedicated, reviewed change. Skipping the reference search is the step most likely to cause an unexpected regression.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
The Feature Flag Cleanup Habit Most Teams Never Build
Why feature flags pile up unused for years, and a simple habit that keeps your flag count from becoming its own source of bugs.
The Feature Flag Graveyard Nobody's Cleaning Up
Feature flags accumulate faster than anyone notices, and the old ones left behind carry a real cost. A checklist for finding and safely deleting them.
A Checklist for Cleaning Up Feature Flags Before They Rot
The common ways feature flags turn into permanent technical debt, and a checklist for cleaning them up before they become a security risk.
Cleaning Up Feature Flags Before They Clean Up You
A checklist for keeping feature flags from piling up into technical debt, including who should own cleanup and what to check before deleting an old flag.
Cleaning Up Feature Flags Before They Become the Bug
A checklist for keeping feature flags around a real-time pipeline from accumulating into their own source of bugs and slow, risky deploys.
The Feature Flags Nobody Remembers Turning On
A checklist for keeping feature flags clean in a RAG pipeline, where flags controlling embedding models, rerankers, and prompts multiply fast.