Data Engineering & Real-Time Event StreamsPlaybook3 min readUpdated September 2026

Cleaning Up Feature Flags Before They Become the Bug

Keep feature flags from becoming bugs by giving each one an expiration date or owner at creation, separating temporary release flags from permanent configuration flags, and deleting the flag code once the outcome is final. Left unmanaged, stale flags turn every deploy into a test of combinations nobody fully understands.

This is a checklist for the specific ways flag hygiene breaks down on a pipeline team, and how to stop it.

Why should every flag get an expiration date at creation?

A flag created without an expected removal date almost never gets removed, because removing a flag competes for time against new feature work and rarely wins on its own. Require an expiration date, or at minimum an owner and a review date, at the moment a flag is created, not as a cleanup step attempted later. This single habit prevents most of the long-term accumulation problem before it starts.

Distinguish Release Flags From Permanent Configuration Flags

Not every flag is meant to be temporary. A release flag, gating a specific feature rollout, should be removed once the feature is fully shipped or fully rolled back. A permanent configuration flag, like a regional feature toggle or a customer-tier gate, is meant to stay. Treating both categories the same way, either never cleaning up or trying to expire everything, causes real problems in both directions. Tag flags by category explicitly so cleanup tooling knows which ones are actually stale.

For example, a team has a release flag for a new deduplication step and a permanent configuration flag that enables a regional data path. After the deduplication rollout finishes, nobody removes the release flag, and months later a change to the regional flag interacts with the dead branch in a way no test covers. The decision rule is to ask, for each flag, whether its outcome could ever change back. If the answer is no, it is a finished release flag and gets deleted. If the answer is yes, it is configuration that stays, with clear documentation of what it controls.

How do you audit feature flag combinations in production?

A pipeline with even a modest number of active flags has an exponential number of possible on/off combinations, and most teams only ever test a handful of them in practice. The dangerous state isn't any single flag, it's two flags that were never tested together in production because the second one shipped after the first one had already been forgotten about. Periodically audit which combinations are actually live in production, not just which flags exist.

Tie Cleanup to Deploy Frequency, Not a Separate Initiative

Teams that deploy often naturally accumulate flags faster, since flags are how frequent, small, safe deploys become possible in the first place. DORA's research groups engineering teams into deployment frequency clusters, from teams shipping multiple times a day at the fastest end to teams releasing as infrequently as every 180 days at the slowest1. A team near the faster end of that range needs flag cleanup built into its regular cadence, as a normal part of shipping, rather than a quarterly cleanup sprint that always loses to feature work.

Make Stale Flags Visible, Not Just Discoverable

A flag dashboard that requires someone to go looking rarely gets checked. Surface stale flags, ones past their expiration date or untouched for a long stretch, somewhere the team already looks, such as a recurring note in a standing engineering meeting or a comment automatically added to relevant pull requests. Visibility without effort is what actually gets cleanup done; a report nobody opens accomplishes nothing regardless of how accurate it is.

Delete the Flag Code, Not Just Turn the Flag Off

Setting a flag permanently to one value and leaving the conditional branching in place isn't cleanup, it's a slightly quieter version of the same problem: the dead branch still has to be read, reasoned about, and accidentally preserved by anyone touching that file later. Once a flag's outcome is final, remove the conditional entirely and delete the code path that's no longer reachable, not just the flag definition. A codebase full of permanently-true or permanently-false conditionals is nearly as confusing to a new reader as one full of active flags.

Assign the deletion as its own small, trackable task at the moment a flag's outcome becomes final, rather than hoping it happens as a byproduct of someone eventually noticing. A flag ticket that closes with 'shipped' and reopens with 'cleanup' as a separate, explicit step is far more likely to actually get the dead code removed than one that treats cleanup as implied.

A flag lifecycle checklist:

  1. Record an expiration date, or at minimum an owner and a review date, at the moment the flag is created.
  2. Tag the flag as a release flag or a permanent configuration flag so cleanup tooling knows which ones are stale.
  3. Surface flags past their expiration date somewhere the team already looks, such as a standing meeting or pull request comments.
  4. Audit which flag combinations are actually live in production, not only which flags exist.
  5. When the outcome is final, delete the conditional and the unreachable code path as its own tracked task.
Executive Capability Standard

What Good Looks Like

Good flag hygiene means every flag has an owner, a category, and an expiration date or review date set at creation, stale flags are surfaced automatically where the team already looks, and combinations of live flags get audited, not just individual flags in isolation.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Audit your current flag inventory and count how many have no owner, no category, and no expiration or review date attached.
2. Do Manually:Add an owner, category, and expiration or review date to every actively used flag, starting with the oldest ones.
3. Delegate:Assign an engineer to own flag hygiene as a recurring responsibility, reviewing stale flags on a set cadence.
4. Automate:Build tooling that flags stale entries automatically in a place the team already looks, such as a standing meeting note or pull request comment.
5. Buy:Adopt a dedicated feature-flag platform with built-in stale-flag detection if your current flag count has outgrown what a manual process can track.

How to Get Started

Frequently Asked Questions

Should every feature flag have an expiration date?

Release flags gating a specific rollout should. Permanent configuration flags, like regional toggles or tier gates, shouldn't be forced onto the same schedule. The key is tagging flags by category at creation so your cleanup process treats each type correctly instead of applying one rule to both.

How often should we audit which flag combinations are actually live in production?

Monthly is a reasonable baseline for a pipeline with an active set of flags, more often if you're shipping and flagging new features frequently. The goal is catching combinations that were never intentionally tested together, not just tracking individual flags in isolation.

What's the biggest reason flag cleanup gets skipped?

It competes directly against feature work for engineering time and usually loses, especially when nobody owns it explicitly. Assigning ownership and making stale flags visible in a place the team already looks, rather than a dashboard nobody opens, is what actually gets cleanup prioritized.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.

Related Guides