Cleaning Up Feature Flags Before They Become the Bug
Keep feature flags from becoming bugs by giving each one an expiration date or owner at creation, separating temporary release flags from permanent configuration flags, and deleting the flag code once the outcome is final. Left unmanaged, stale flags turn every deploy into a test of combinations nobody fully understands.
This is a checklist for the specific ways flag hygiene breaks down on a pipeline team, and how to stop it.
Why should every flag get an expiration date at creation?
A flag created without an expected removal date almost never gets removed, because removing a flag competes for time against new feature work and rarely wins on its own. Require an expiration date, or at minimum an owner and a review date, at the moment a flag is created, not as a cleanup step attempted later. This single habit prevents most of the long-term accumulation problem before it starts.
Distinguish Release Flags From Permanent Configuration Flags
Not every flag is meant to be temporary. A release flag, gating a specific feature rollout, should be removed once the feature is fully shipped or fully rolled back. A permanent configuration flag, like a regional feature toggle or a customer-tier gate, is meant to stay. Treating both categories the same way, either never cleaning up or trying to expire everything, causes real problems in both directions. Tag flags by category explicitly so cleanup tooling knows which ones are actually stale.
For example, a team has a release flag for a new deduplication step and a permanent configuration flag that enables a regional data path. After the deduplication rollout finishes, nobody removes the release flag, and months later a change to the regional flag interacts with the dead branch in a way no test covers. The decision rule is to ask, for each flag, whether its outcome could ever change back. If the answer is no, it is a finished release flag and gets deleted. If the answer is yes, it is configuration that stays, with clear documentation of what it controls.
How do you audit feature flag combinations in production?
A pipeline with even a modest number of active flags has an exponential number of possible on/off combinations, and most teams only ever test a handful of them in practice. The dangerous state isn't any single flag, it's two flags that were never tested together in production because the second one shipped after the first one had already been forgotten about. Periodically audit which combinations are actually live in production, not just which flags exist.
Tie Cleanup to Deploy Frequency, Not a Separate Initiative
Teams that deploy often naturally accumulate flags faster, since flags are how frequent, small, safe deploys become possible in the first place. DORA's research groups engineering teams into deployment frequency clusters, from teams shipping multiple times a day at the fastest end to teams releasing as infrequently as every 180 days at the slowest1. A team near the faster end of that range needs flag cleanup built into its regular cadence, as a normal part of shipping, rather than a quarterly cleanup sprint that always loses to feature work.
Make Stale Flags Visible, Not Just Discoverable
A flag dashboard that requires someone to go looking rarely gets checked. Surface stale flags, ones past their expiration date or untouched for a long stretch, somewhere the team already looks, such as a recurring note in a standing engineering meeting or a comment automatically added to relevant pull requests. Visibility without effort is what actually gets cleanup done; a report nobody opens accomplishes nothing regardless of how accurate it is.
Delete the Flag Code, Not Just Turn the Flag Off
Setting a flag permanently to one value and leaving the conditional branching in place isn't cleanup, it's a slightly quieter version of the same problem: the dead branch still has to be read, reasoned about, and accidentally preserved by anyone touching that file later. Once a flag's outcome is final, remove the conditional entirely and delete the code path that's no longer reachable, not just the flag definition. A codebase full of permanently-true or permanently-false conditionals is nearly as confusing to a new reader as one full of active flags.
Assign the deletion as its own small, trackable task at the moment a flag's outcome becomes final, rather than hoping it happens as a byproduct of someone eventually noticing. A flag ticket that closes with 'shipped' and reopens with 'cleanup' as a separate, explicit step is far more likely to actually get the dead code removed than one that treats cleanup as implied.
A flag lifecycle checklist:
- Record an expiration date, or at minimum an owner and a review date, at the moment the flag is created.
- Tag the flag as a release flag or a permanent configuration flag so cleanup tooling knows which ones are stale.
- Surface flags past their expiration date somewhere the team already looks, such as a standing meeting or pull request comments.
- Audit which flag combinations are actually live in production, not only which flags exist.
- When the outcome is final, delete the conditional and the unreachable code path as its own tracked task.
What Good Looks Like
Good flag hygiene means every flag has an owner, a category, and an expiration date or review date set at creation, stale flags are surfaced automatically where the team already looks, and combinations of live flags get audited, not just individual flags in isolation.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Should every feature flag have an expiration date?
Release flags gating a specific rollout should. Permanent configuration flags, like regional toggles or tier gates, shouldn't be forced onto the same schedule. The key is tagging flags by category at creation so your cleanup process treats each type correctly instead of applying one rule to both.
How often should we audit which flag combinations are actually live in production?
Monthly is a reasonable baseline for a pipeline with an active set of flags, more often if you're shipping and flagging new features frequently. The goal is catching combinations that were never intentionally tested together, not just tracking individual flags in isolation.
What's the biggest reason flag cleanup gets skipped?
It competes directly against feature work for engineering time and usually loses, especially when nobody owns it explicitly. Assigning ownership and making stale flags visible in a place the team already looks, rather than a dashboard nobody opens, is what actually gets cleanup prioritized.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.
Related Guides
The Feature Flag Cleanup Habit Most Teams Never Build
Why feature flags pile up unused for years, and a simple habit that keeps your flag count from becoming its own source of bugs.
Blue-Green, Canary, or Rolling: Deploying Stream Processors
A decision guide to rolling, blue-green, and canary deploys for stateful stream processors, plus the rollback plan most teams never actually test.
A Checklist for Cleaning Up Feature Flags Before They Become Their Own Codebase
A checklist for finding and safely removing stale feature flags, and the pitfalls that turn a routine cleanup into a production incident.
Stale Feature Flags Are Technical Debt With a Kill Switch
A feature flag left in code after launch is a branch nobody tests and a rollback path nobody trusts. Here is a checklist for keeping flags from piling up.
Verifying Every Service That Talks to Your Pipeline
Which parts of zero-trust verification to build and which to buy, so every producer and consumer on a streaming pipeline proves its identity.
The Feature Flag Graveyard Nobody's Cleaning Up
Feature flags accumulate faster than anyone notices, and the old ones left behind carry a real cost. A checklist for finding and safely deleting them.