Rotating API Keys With Zero Downtime: A Step-by-Step Playbook
To rotate an API key without downtime, make the new key valid before you retire the old one, move every consumer to the new key, confirm nothing still uses the old one, and only then revoke it. The overlap window is what prevents an outage.
The steps differ depending on whether you consume a key from another provider or issue keys to your own customers. Both are covered below, with the mistakes that cause most rotation incidents.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Why do rotations cause outages?
Rotations fail for predictable reasons: the old key is revoked before all consumers have the new one, a service caches the secret at startup and never reloads it, or a forgotten script uses the key from a laptop or a cron job. Each of these is a coordination problem, not a cryptography problem.
Outages also cost real budget. A 99.9 percent availability target allows about 8.76 hours of downtime a year1, so a rotation that breaks production for even a quarter of an hour uses a visible slice of it. The fix is a design that can't fail closed: keep two valid keys during the transition and prove the old one is idle before you kill it.
How do you rotate a key you use from a third-party provider?
Follow this sequence:
- Inventory consumers. List every service, job, pipeline and person that uses the key. Search the repositories and your secrets store.
- Create the new key at the provider without disabling the old one. Confirm the provider allows two active keys at once. If it doesn't, plan a brief switch window and pick a quiet time.
- Store the new key in your secrets manager as a new version, not by overwriting the old one in place.
- Roll it out in stages. Update a low-risk consumer first, then verify requests succeed with the new key. Continue service by service.
- Watch for the old key. Check provider logs for last-used times on the old key.
- Revoke the old key only after it has shown no traffic for a period you set, such as a full business cycle.
- Record the date, the owner and the next rotation due.
Secrets tools such as Doppler, HashiCorp Vault or AWS Secrets Manager are options for storing versioned secrets and getting a new value to your services, which can make step four safer. Confirm in a demo how each fits your stack.
How do you let your own customers rotate keys safely?
If you issue API keys, build rotation into the product so customers never face a cliff:
- Allow multiple active keys per account. Customers create key two, deploy it, then delete key one.
- Give each key an identifier and a label, and show last-used time, so customers can see stragglers.
- Support expiry with a grace period, so an expiring key still works briefly and returns a clear warning header or message.
- Notify before expiry by email and in the dashboard.
- Store only hashes of keys on your side and show the full value once at creation.
- Offer an emergency revoke that takes effect immediately for a leaked key.
The common mistake is a single key per account, where rotation means a hard cutover and downtime for every integration.
Which application patterns make rotation painless?
Design services so a new secret doesn't require a redeploy:
- Read secrets at runtime, from the secrets manager or a mounted file, instead of baking them into images or build-time variables.
- Reload on change or on failure. If a request fails authentication, fetch the latest secret once and retry before returning an error.
- Accept two values on the receiving side for shared secrets, such as webhook signing keys, by verifying against both the current and the previous key during a transition.
- Use short-lived credentials where the platform allows it, so rotation happens automatically and nobody handles a long-lived key.
- Keep secrets out of logs and error messages, since a rotated key that leaked into logs stays exposed.
For picking a store, see the comparison of HashiCorp Vault, AWS Secrets Manager and Doppler.
What if a key leaks and you must rotate now?
An emergency rotation compresses the same steps, so rehearse it before you need it:
- Revoke first if the exposure is severe, accepting a short outage over continued misuse. If the key can move money or expose customer data, don't wait for a clean overlap.
- Rotate related secrets. A leaked deployment token may expose others in the same environment.
- Check the provider's logs for activity from unfamiliar addresses during the exposure window.
- Remove the secret from history where it appeared, but assume it's compromised anyway, since deleting a commit doesn't undo who saw it.
- Write a short incident note covering how it leaked and what prevents a repeat.
Set a routine schedule too. Many teams rotate on a calendar and after every offboarding of someone with access. For a more careful approach to environment files, see the Doppler vs AWS Secrets Manager comparison for B2B SaaS.
What Good Looks Like
Every secret has an owner, a rotation schedule and a tested overlap procedure, so rotating it never interrupts a request.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Fits when you want one place to manage secret versions across environments; confirm in a demo how it delivers them to your services.
Fits when you want rotation to be as automatic as possible and can run and operate the service yourself.
Fits when your workloads run on AWS and you want secrets managed with your existing IAM permissions; confirm what rotation it supports for your key types.
Frequently Asked Questions
How often should API keys be rotated?
Set a schedule that fits the key's risk, such as every few months for sensitive keys, and rotate immediately after a suspected leak or when someone with access leaves. Automating rotation makes frequent schedules practical and reduces the chance of an outage.
What is the safest way to rotate an API key?
Create the new key while the old one is still valid, move all consumers over in stages, confirm the old key has no traffic, then revoke it. This overlap prevents outages. Store each new value as a new version in a secrets manager.
What if the provider only allows one active key?
Plan a short switch window during low traffic. Prepare the new value in your secrets store, create the key, update consumers immediately and verify. Alternatively, ask whether the provider supports a second credential type or per-service keys that avoid the constraint.
How do I find where an API key is used?
Search repositories, CI variables, secrets stores, infrastructure code and shared documents, then check the provider's usage logs for the key's last-used time and source addresses. Expect surprises, so keep the old key active until logs confirm nothing still calls it.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
HashiCorp Vault vs AWS Secrets Manager vs Doppler: Secrets Platforms
Compare Vault, AWS Secrets Manager, and Doppler for secret sprawl prevention, dynamic credential rotation, Kubernetes injection, and SOC 2 audits.
Doppler or AWS Secrets Manager for a Multi-Cloud SaaS Stack
A criteria-based way for B2B SaaS teams to pick between Doppler and AWS Secrets Manager, based on where your deploys actually run and who audits you.
Shipping API Version Migrations Without a Maintenance Window
A step-by-step approach to migrating API versions and running database or schema changes without a maintenance window or breaking existing clients.
A Runbook for Shipping Breaking API Changes Without Downtime
A step-by-step approach to shipping a breaking API or schema change without a maintenance window, built around parallel versions.
A Runbook for Zero-Downtime Schema Migrations on a Live Database
A step-by-step runbook for running schema migrations against a production database without an outage window, including the rollback checkpoints.
A Runbook for Shipping a New Model Version Without Downtime
A step-by-step way to roll a new model version into production: shadow traffic first, a small canary, clear rollback triggers, and a real cutover.