Data Engineering & Real-Time Event StreamsPlaybook3 min readUpdated September 2026

Rotating Credentials on a Live Pipeline Without an Outage

To rotate credentials on a live pipeline without an outage, overlap the old and new credential, confirm every client has picked up the new one, and only then revoke the old one. This matters more than on most services because broker certificates, connector keys, and registry tokens sit on long-lived connections.

Here's a way to do it without an outage, for the three credential types that actually show up in a streaming stack, plus the difference between a routine rotation and one forced by a suspected leak.

How do you rotate broker certificates without an outage?

Brokers and clients that authenticate with mutual TLS hold that connection open for a long time, so swapping a certificate abruptly, revoking the old one before every client has picked up the new one, disconnects everything still using it. Issue the new certificate, deploy it to clients, and confirm every client has actually reconnected with it before revoking the old one.

A client inventory matters here: you need a way to confirm every producer and consumer has picked up the new certificate, not just that you deployed it, since a client that's crashed or stuck on an old connection won't rotate on its own.

Connector API keys: rotate through a secrets manager, not a config file

A connector's third-party API key hardcoded into a config file or environment variable makes rotation a deploy event, which is slower than it needs to be and easy to forget on a regular cadence. Route these through a secrets manager that supports rotation without a full redeploy, so the connector picks up the new key on its next scheduled refresh instead of waiting for someone to notice the old key is stale.

Set a rotation cadence and actually follow it, since an API key that's never rotated is functionally a permanent credential, and a permanent credential that leaks has no expiration working in your favor.

For example, a connector pulls enrichment data from a third party using an API key stored in a config file. Nobody rotates it because rotation means a deploy, so the key quietly becomes a permanent credential. Moving the key into a secrets manager with live rotation turns the change into a scheduled, low-drama task: the connector refreshes the key on its next cycle, and you check its logs for a successful call with the new key before disabling the old one. If that check fails, the old key is still there to fall back on, which is the whole point of overlapping.

Schema registry credentials: check every service that writes schemas, not just producers

It's easy to remember the producers that read from the schema registry regularly and forget the one-off migration script or admin tool that only writes a schema change once a quarter. Before rotating a schema registry credential, inventory every service and script that touches it, not just the ones in daily use, since a forgotten one will fail silently the next time someone runs it, often months later when nobody remembers the rotation happened at all.

Practice the rotation before you need to do it under pressure

The first time you rotate a given credential type shouldn't be during an actual suspected compromise. Run a practice rotation on a non-production environment for each credential type above, on a schedule, so the team has already worked out the client inventory and overlap steps before a real incident forces the pace.

Time the practice run and write down what actually happened, including anything that didn't go the way the plan assumed. That record is what turns a rotation from something one engineer remembers how to do into something the whole team can execute correctly, regardless of who's on call that day.

How is an emergency rotation different from a scheduled one?

A scheduled rotation, planned days in advance with plenty of time to confirm every client has picked up the new credential, is a different problem from an emergency rotation forced by a suspected leak, where you want the old credential dead as fast as possible even at the cost of a brief disruption. Decide in advance which of your credential types can tolerate a hard, immediate cutover in an emergency and which genuinely need the overlap approach even under pressure, so that decision isn't being made for the first time during an actual incident.

For anything that can't tolerate a hard cutover safely, know in advance roughly how long a full, careful rotation actually takes, so an incident response conversation about acceptable exposure time is grounded in a real number instead of a guess made under pressure.

Use this sequence for a scheduled rotation, and adapt it for an emergency:

  1. Inventory every producer, consumer, connector, script, and admin tool that uses the credential, including ones that only run occasionally.
  2. Issue the new credential and deploy it while the old one still works.
  3. Confirm every client has actually picked up the new credential, not just that it was deployed.
  4. Revoke the old credential only after that confirmation.
  5. Time a practice run in a non-production environment and write down anything that didn't go as planned.
Executive Capability Standard

What Good Looks Like

Rotation is production-ready when every credential type has a documented client inventory, an overlap step that avoids a hard cutover, and a rotation that's been practiced outside of an actual incident.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Inventory every credential type your pipeline uses and every service or script that holds each one.
2. Do Manually:Run one practice rotation for your highest-risk credential type on a non-production environment.
3. Delegate:Assign a platform engineer to own a rotation calendar and the client inventory for each credential type.
4. Automate:Move connector API keys into a secrets manager that supports rotation without a full redeploy.
5. Buy:Bring in a security specialist to design the rotation process if a past incident already showed gaps in it.

How to Get Started

Frequently Asked Questions

How often should broker certificates actually be rotated?

Follow your organization's standard certificate lifetime policy, but the more important habit is having a tested, repeatable rotation process, not a specific interval. A short certificate lifetime with an untested rotation process is worse than a longer one your team has actually practiced rotating without an outage.

What's the biggest risk when rotating a broker certificate?

Revoking the old certificate before every client has actually reconnected with the new one. This is why a client inventory matters: without one, you're guessing whether rotation is complete instead of confirming it, and a client stuck on the old certificate will simply disconnect the moment you revoke it.

Should connector API keys go through the same process as broker certificates?

The principle is the same (avoid a hard cutover that breaks something still using the old credential) but the mechanism differs. A secrets manager that supports live rotation is usually a better fit for API keys than the certificate overlap process above, since most connectors can pick up a new key on their next scheduled refresh.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides