A Checklist for Cleaning Up Feature Flags Before They Rot
Feature flags rot when they outlive their purpose, leaving untested code paths and forgotten access rules in your codebase. A flag is meant to be temporary, so cleanup means giving each one an expiry date, deleting dead branches, and sweeping on a regular schedule.
Give every flag an expiry date when it's created, not after
A flag created without an expected removal date almost never gets one added later, because by the time someone notices it's stale, nobody remembers why it exists or whether it's safe to remove. Require an expected resolution date as part of creating any flag, a rollout flag that becomes permanent code within a month, an experiment flag that resolves in a quarter, and track flags past their date the same way you'd track an overdue task, not as background noise. Make the date a required field in whatever tool creates the flag, so skipping it simply isn't an option under deadline pressure.
Why are access gating flags riskier than other flags?
Some feature flags aren't really feature flags, they're informal access controls: an internal only flag, a beta customer allowlist, an emergency kill switch for a risky feature. These deserve more scrutiny than a simple UI toggle, because a stale access gating flag is effectively an undocumented permission system that a security review won't find unless someone specifically looks for it. Tag these differently in your flag system and review them on their own schedule, separate from routine feature flags.
Remove the dead code path, not just the flag check
Deleting a flag definition while leaving both branches of the old if statement in the codebase doesn't actually clean anything up, it just makes the dead branch harder to find later. Removing a resolved flag means picking a winner, deleting the losing code path entirely, and confirming the remaining tests still cover what's left. This is more work than deleting a flag record, which is exactly why it's the step that gets skipped under time pressure and the one worth explicitly checking for in a review.
For example, suppose a checkout flag has been on for everyone for months. Removing only the flag record leaves an old branch that no test exercises and no reader can tell is dead. Removing the flag properly means keeping the winning branch, deleting the losing one, and running the tests that cover what remains. A common mistake is reviewing only the flag definition in the pull request. Ask reviewers to check that the losing branch is gone too, since that is the step that gets skipped when a deadline is close.
Watch for flags that combine into untested states
A codebase with a dozen active flags doesn't have a dozen possible states, it has up to two to the twelfth, and almost none of those combinations have ever actually been tested or even considered. This is where flags quietly become a reliability and security risk: an old flag combined with a new one can produce a code path nobody designed for, including one that accidentally bypasses an access check that assumed the old flag would never coexist with the new one. Keeping the total count of live flags low is a more effective mitigation than trying to test every combination.
How often should you sweep for stale flags?
On a set cadence, pull every flag past its expected resolution date and force a decision: ship it permanently and remove the flag, kill the feature and remove the flag, or extend the date with an explicit written reason. A flag with an expired date and no owner willing to make a call is itself a finding, since it usually means the feature's outcome was never actually decided, just left running indefinitely by default.
Put a hard cap on how many flags can be active at once
A quarterly sweep catches existing rot, but a cap on total active flags changes behavior going forward: once the team is near the limit, adding a new flag means someone has to actually resolve an old one first, rather than the count creeping upward indefinitely. Pick a number that fits your team's size and stick to it, treating a flag creation request that would exceed the cap as a prompt to look at what's already stale, not as friction to route around. This one change tends to do more to prevent future rot than any amount of after the fact cleanup discipline.
A flag cleanup checklist:
- Require an expected resolution date when a flag is created, as a required field in your flag tool.
- Tag access gating flags separately and review them on their own schedule.
- When a flag is resolved, pick a winner and delete the losing code path entirely.
- Keep the number of live flags low, since combinations of flags create untested states.
- Sweep flags past their date each quarter and decide: ship, kill, or extend with a written reason.
- Cap the number of active flags so adding a new one means resolving an old one.
What Good Looks Like
Good flag hygiene means every flag has an expiry date set at creation, access gating flags are tracked separately from feature flags, and resolved flags get their dead code path removed, not just the flag record.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How long should a typical feature flag stay active before it's considered stale?
It depends on the flag's purpose: a gradual rollout flag should usually resolve within a month, while a longer running experiment might reasonably run a quarter. The key is setting that expectation when the flag is created, not deciding after the fact that it's been too long.
Why are access gating flags more risky than regular feature flags?
Because they function as an informal permission system that a normal security review won't catch unless someone specifically audits flags for this purpose. A stale flag that was meant to gate a beta feature to a handful of internal users can quietly become a forgotten access rule nobody remembers to revoke.
Is it enough to just delete old flag definitions during cleanup?
No. Deleting the flag record without removing the dead code branch it used to gate leaves that code path in the codebase, untested and undocumented. A real cleanup removes the losing branch entirely, not just the flag that used to point to it.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Continuous Device Verification for a Zero-Trust API
How continuous device and identity verification actually works in a zero-trust architecture, and where to draw the line for a small engineering team.
The Feature Flag Cleanup Habit Most Teams Never Build
Why feature flags pile up unused for years, and a simple habit that keeps your flag count from becoming its own source of bugs.
A Checklist for Cleaning Up Feature Flags Before They Become Their Own Codebase
A checklist for finding and safely removing stale feature flags, and the pitfalls that turn a routine cleanup into a production incident.
Rolling Out Zero Trust in Production Without a Broad Outage
A checklist for rolling out stricter API authentication and authorization in production, and the pitfalls that turn a rollout into an incident.
The Feature Flag Graveyard Nobody's Cleaning Up
Feature flags accumulate faster than anyone notices, and the old ones left behind carry a real cost. A checklist for finding and safely deleting them.
How to Audit Whether Your APIs Actually Enforce Zero Trust
A step-by-step method for testing whether your APIs enforce zero trust in practice, not just on paper, and what to do with what you find.