What to Actually Alert On in a Zero Trust API Setup
A zero trust setup generates a lot of security-relevant events: token rejections, scope mismatches, certificate expirations, policy denials. Most of that is normal background noise, and alerting on all of it trains your team to ignore the channel entirely, which is worse than having no alerts at all.
Build this as a worksheet: one row per signal, with a decision about whether it pages someone, shows up on a dashboard, or just gets logged.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
How do you tell expected denials from anomalous ones?
A single token rejection because a session expired is expected behavior, not an incident. A sudden spike in rejections from one client, or rejections from a client that has never had one before, is a different signal entirely. Build your alerting around the second category (rate of change, first occurrence, unusual concentration) rather than raw counts of denials, which will always be high in a healthy system doing its job.
Write down, per signal, what "normal" looks like on a typical day, so an anomaly detector or a human reviewing a dashboard has something to compare against.
Imagine a client that normally sees a handful of expired-token rejections a day because its refresh logic occasionally lags: that's baseline noise. The same client suddenly generating rejections on every single call is the anomaly, and the difference between the two is the rate of change against its own history, not an absolute number that applies to every client the same way.
Give certificate and credential expiry its own lane
Expiring mTLS certificates and rotating API keys cause a specific, predictable failure mode: everything works fine until a specific date, then a service starts failing every request. This deserves advance warning, days ahead, not a same-day alert when it's already too late to fix calmly. Put expiry warnings on their own schedule-based check rather than mixing them into your general anomaly alerting, since they need a different kind of response (renew before the deadline) than a security anomaly does (investigate now).
A common mistake is treating certificate expiry as a one-time setup task instead of a recurring operational one: the team that provisioned the cert two years ago has often moved on to other work by the time it expires, and the alert fires against an on-call rotation that has no context for what the certificate is even for. Document what each certificate protects, not just when it expires, so whoever gets the alert can act on it without archaeology.
Does every alert have a documented first response?
An alert with no defined first action just interrupts someone at 2 a.m. so they can look at it and go back to sleep. For every alert in your matrix, write one sentence describing what the on-call engineer actually does when it fires: check this dashboard, run this query, page this other team. If you can't write that sentence, the alert isn't ready to page anyone; downgrade it to a dashboard metric until it is.
Do this review quarterly. Alerts that made sense at your old scale often stop making sense as traffic patterns and team structure change.
Build the dashboard around your uptime commitment, not raw counts
Your error budget, and how it maps to your availability target, is the right lens for a security-observability dashboard: it converts a stream of individual denial events into a single number your team can track against a real commitment. The gap between a 99.9% and a 99.99% target is the difference between roughly nine hours and roughly fifty minutes of allowed downtime a year1, and that budget should be visible on the same dashboard as your security denial rate, since a misconfigured policy that starts rejecting valid traffic burns the same budget as an infrastructure outage.
Review the whole matrix after every real incident
After any real security incident, however small, go back to your alerting matrix and ask whether the signal that would have caught it earlier existed, and if it did, why it didn't page anyone. This is a better source of new alerting rules than brainstorming in the abstract, because it's grounded in something that actually happened rather than something that theoretically could.
Keep the review short and specific: one incident, one question (would an alert have caught this sooner, and why didn't the existing ones), one concrete change to the matrix. A long, general postmortem meeting tends to produce vague action items that never turn into an actual new alert; a narrow one produces a rule you can point to next quarter.
Each row in your alerting matrix should record:
- The signal itself and what normal looks like on a typical day, so anomalies have a baseline to compare against.
- Whether the signal pages someone, appears on a dashboard, or is only logged for later review.
- One sentence describing the first action the on-call engineer takes when it fires, or the alert gets downgraded.
- For certificate and credential expiry, a schedule-based warning days ahead instead of a same-day alert.
What Good Looks Like
Good practice means every alert in your zero trust monitoring has a documented first response and a distinction between expected, routine denials and genuinely anomalous ones, so the channel stays worth watching.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
How do we avoid alert fatigue with all these zero trust signals?
Alert on anomalies and rate of change rather than raw event counts, give every alert a documented first response, and delete or downgrade any alert that hasn't led to real action in the last quarter. Fatigue comes from volume without action, not from having security monitoring at all.
Should certificate expiry alerts page someone or just show on a dashboard?
A dashboard with an escalating schedule works better than a single page: a quiet reminder weeks out, a louder one days out, and only an actual page if it's within 24 hours and still unrenewed. Paging weeks in advance for something with a known, distant deadline just trains people to snooze it.
What's a reasonable first metric to put on a zero trust dashboard?
Start with the rate of authorization denials per service over time, split by whether the denial came from a known caller or an unrecognized one. That single view surfaces both misconfigurations (a known caller suddenly failing) and potential probing (an unrecognized caller trying repeatedly) without needing a dozen separate charts.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
Continuous Device Verification for a Zero-Trust API
How continuous device and identity verification actually works in a zero-trust architecture, and where to draw the line for a small engineering team.
Rolling Out Zero Trust in Production Without a Broad Outage
A checklist for rolling out stricter API authentication and authorization in production, and the pitfalls that turn a rollout into an incident.
How to Audit Whether Your APIs Actually Enforce Zero Trust
A step-by-step method for testing whether your APIs enforce zero trust in practice, not just on paper, and what to do with what you find.
The Real Latency Cost of Zero Trust, and How to Measure It
How to find out how much latency your zero trust controls actually add, which checks are worth the cost, and which ones you can move off the hot path.
Keeping Auth Checks Fast as Your API Traffic Grows
A worked example for keeping zero trust authorization checks fast as request volume grows, and where teams usually add latency without noticing.
How to Actually Compare API Gateway Latency Claims
A method for benchmarking API gateway latency yourself, since vendor numbers rarely reflect what your own policies will cost you in practice.