The First Hour After a Suspected API Key Compromise
The difference between a contained incident and a serious breach usually comes down to how fast the first hour goes, and that speed comes from having a written runbook, not from improvising under pressure. A team writing its incident process for the first time during an actual incident makes worse decisions than one following a plan written calmly in advance.
Here's the sequence for the specific case of a suspected compromised API key or credential.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
What should you do in minute zero of a suspected key compromise?
The moment a credential is suspected of being compromised, revoke it immediately, before you've fully confirmed the suspicion. The cost of a false alarm, a legitimate integration briefly disrupted while you reissue a credential, is far smaller than the cost of leaving a genuinely compromised credential active while you investigate.
Have the revocation mechanism ready before you need it: know exactly which button to click or which command to run, tested in advance, not something to figure out for the first time during the incident itself. Assign this step to a specific on-call role, not "whoever notices first," since an incident discovered at an odd hour needs a clear owner who can act immediately rather than a diffusion of responsibility that delays the one step that matters most in the first few minutes of the whole response.
Contain the blast radius before you start root-causing
Once the immediate credential is revoked, check what else that identity had access to and whether any of it shows signs of misuse: unusual data access patterns, actions outside the caller's normal behavior, requests from unfamiliar locations. Contain anything showing active misuse, additional key revocations, session invalidation, before moving to the slower work of figuring out how the credential was exposed in the first place.
Resist the urge to fully root-cause before containing. It's tempting to understand the whole story first, but containment is time-sensitive in a way root-causing generally isn't, and a partial containment done quickly beats a complete, well-understood response that arrives after more damage has already happened.
For example, suppose a contractor's API key shows up in a public code repository at an odd hour. The on-call owner revokes it within minutes, before anyone has confirmed whether it was used. A teammate then reviews that key's recent activity while a third person drafts the initial customer notice from the template. Nothing here requires knowing the root cause yet. Each person has one job, and the order matters more than the detail. The common mistake is letting the discussion turn into a debate about how the key leaked, which delays the one action that limits damage. If you catch yourself asking why before the key is revoked, stop and revoke.
Who should you notify after a suspected compromise, and how fast?
Decide in advance who needs to know and how fast: engineering leadership immediately, affected customers or partners within a defined window if their data or access was involved, and legal counsel early if there's any chance of a regulatory notification obligation. Waiting to notify until you have a complete picture usually means notifying too late.
A short, honest initial notice, here's what we know, here's what we're doing, more details to follow, is better than delaying notification until every detail is confirmed. Keep a template for this initial notice drafted in advance, since writing careful, accurate communication under time pressure during a real incident is exactly when mistakes in tone or unintentionally overstated claims creep in, and having a starting point to edit is faster than composing something from scratch while the rest of the response is still underway.
Run the postmortem on the process, not just the technical cause
A postmortem that only identifies the technical root cause, a key committed to a public repo, a credential left in an old config file, misses the second question: why did detection take as long as it did, and did the runbook actually work as written. Both questions matter equally for reducing the next incident's impact.
Update the runbook itself based on what the postmortem finds. A runbook that never changes after an incident either means it was already perfect, which is unlikely, or that the postmortem's findings aren't feeding back into the process that's supposed to use them. Schedule the runbook update as a specific task with an owner and a deadline coming out of the postmortem meeting, rather than a general action item that tends to quietly drop off everyone's list once the immediate pressure of the incident has passed.
The first-hour sequence, in order:
- Revoke the suspect credential immediately, before the suspicion is confirmed, using a revocation step that a named on-call owner has already tested.
- Check what else that identity could reach, and contain any sign of misuse with further key revocations or session invalidation before you root-cause.
- Tell engineering leadership right away, send affected customers a short honest initial notice within your defined window, and involve legal counsel early if a regulatory duty is possible.
- Hold a postmortem that asks both how the credential was exposed and why detection took as long as it did.
- Turn the findings into a runbook update with a named owner and a deadline.
What Good Looks Like
Good practice means a tested, written runbook exists for credential compromise specifically, revocation happens before full investigation rather than after, and every incident's postmortem examines the response process itself, not just the technical root cause.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
Should we revoke a credential before we're certain it was actually compromised?
Yes, if there's genuine suspicion. Revoking and reissuing a credential that turns out to have been fine costs far less than leaving a truly compromised one active while you take time to confirm, and reissuing is usually a quick operation if you've built your rotation process well.
How quickly should affected customers be notified after a suspected compromise?
Typically within 24 to 72 hours, once you can say something true and useful, though contracts or applicable law may set a different deadline depending on what data was involved. A short initial notice that says what you know and what you are doing beats waiting for every detail. Check with your attorney on notification obligations specific to your jurisdiction and industry.
What's the most common runbook gap teams discover during their first real incident?
Revocation taking longer than expected because the actual mechanism, who has access to the admin panel, what the command is, was never tested in advance. Running a tabletop exercise that walks through the revocation steps before an incident happens catches this gap for free.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Continuous Device Verification for a Zero-Trust API
How continuous device and identity verification actually works in a zero-trust architecture, and where to draw the line for a small engineering team.
Rolling Out Zero Trust in Production Without a Broad Outage
A checklist for rolling out stricter API authentication and authorization in production, and the pitfalls that turn a rollout into an incident.
How to Audit Whether Your APIs Actually Enforce Zero Trust
A step-by-step method for testing whether your APIs enforce zero trust in practice, not just on paper, and what to do with what you find.
The Real Latency Cost of Zero Trust, and How to Measure It
How to find out how much latency your zero trust controls actually add, which checks are worth the cost, and which ones you can move off the hot path.
Keeping Auth Checks Fast as Your API Traffic Grows
A worked example for keeping zero trust authorization checks fast as request volume grows, and where teams usually add latency without noticing.
How to Actually Compare API Gateway Latency Claims
A method for benchmarking API gateway latency yourself, since vendor numbers rarely reflect what your own policies will cost you in practice.