API Security, Identity & Zero-TrustPlaybook3 min readUpdated September 2026

What to Actually Alert On in a Zero Trust API Setup

A zero trust setup generates a lot of security-relevant events: token rejections, scope mismatches, certificate expirations, policy denials. Most of that is normal background noise, and alerting on all of it trains your team to ignore the channel entirely, which is worse than having no alerts at all.

Build this as a worksheet: one row per signal, with a decision about whether it pages someone, shows up on a dashboard, or just gets logged.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

How do you tell expected denials from anomalous ones?

A single token rejection because a session expired is expected behavior, not an incident. A sudden spike in rejections from one client, or rejections from a client that has never had one before, is a different signal entirely. Build your alerting around the second category (rate of change, first occurrence, unusual concentration) rather than raw counts of denials, which will always be high in a healthy system doing its job.

Write down, per signal, what "normal" looks like on a typical day, so an anomaly detector or a human reviewing a dashboard has something to compare against.

Imagine a client that normally sees a handful of expired-token rejections a day because its refresh logic occasionally lags: that's baseline noise. The same client suddenly generating rejections on every single call is the anomaly, and the difference between the two is the rate of change against its own history, not an absolute number that applies to every client the same way.

Give certificate and credential expiry its own lane

Expiring mTLS certificates and rotating API keys cause a specific, predictable failure mode: everything works fine until a specific date, then a service starts failing every request. This deserves advance warning, days ahead, not a same-day alert when it's already too late to fix calmly. Put expiry warnings on their own schedule-based check rather than mixing them into your general anomaly alerting, since they need a different kind of response (renew before the deadline) than a security anomaly does (investigate now).

A common mistake is treating certificate expiry as a one-time setup task instead of a recurring operational one: the team that provisioned the cert two years ago has often moved on to other work by the time it expires, and the alert fires against an on-call rotation that has no context for what the certificate is even for. Document what each certificate protects, not just when it expires, so whoever gets the alert can act on it without archaeology.

Does every alert have a documented first response?

An alert with no defined first action just interrupts someone at 2 a.m. so they can look at it and go back to sleep. For every alert in your matrix, write one sentence describing what the on-call engineer actually does when it fires: check this dashboard, run this query, page this other team. If you can't write that sentence, the alert isn't ready to page anyone; downgrade it to a dashboard metric until it is.

Do this review quarterly. Alerts that made sense at your old scale often stop making sense as traffic patterns and team structure change.

Build the dashboard around your uptime commitment, not raw counts

Your error budget, and how it maps to your availability target, is the right lens for a security-observability dashboard: it converts a stream of individual denial events into a single number your team can track against a real commitment. The gap between a 99.9% and a 99.99% target is the difference between roughly nine hours and roughly fifty minutes of allowed downtime a year1, and that budget should be visible on the same dashboard as your security denial rate, since a misconfigured policy that starts rejecting valid traffic burns the same budget as an infrastructure outage.

Review the whole matrix after every real incident

After any real security incident, however small, go back to your alerting matrix and ask whether the signal that would have caught it earlier existed, and if it did, why it didn't page anyone. This is a better source of new alerting rules than brainstorming in the abstract, because it's grounded in something that actually happened rather than something that theoretically could.

Keep the review short and specific: one incident, one question (would an alert have caught this sooner, and why didn't the existing ones), one concrete change to the matrix. A long, general postmortem meeting tends to produce vague action items that never turn into an actual new alert; a narrow one produces a rule you can point to next quarter.

Each row in your alerting matrix should record:

  • The signal itself and what normal looks like on a typical day, so anomalies have a baseline to compare against.
  • Whether the signal pages someone, appears on a dashboard, or is only logged for later review.
  • One sentence describing the first action the on-call engineer takes when it fires, or the alert gets downgraded.
  • For certificate and credential expiry, a schedule-based warning days ahead instead of a same-day alert.
Executive Capability Standard

What Good Looks Like

Good practice means every alert in your zero trust monitoring has a documented first response and a distinction between expected, routine denials and genuinely anomalous ones, so the channel stays worth watching.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Audit your current alert list and count how many have no documented first response, then read up on error budgets as a framing for security dashboards.
2. Do Manually:Build the alerting worksheet by hand for your five most important services, deciding page, dashboard, or log for each existing signal.
3. Delegate:Give an on-call rotation lead ownership of a quarterly alert review, with authority to delete alerts that haven't led to action.
4. Automate:Wire your alerting rules to rate-of-change and first-occurrence detection instead of static thresholds, so new anomalies surface without manual tuning of every threshold.
5. Buy:Use a vulnerability management platform's own dashboards to cover the parts of this picture that come from known, disclosed weaknesses rather than building that view yourself.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Tenable

Tenable is worth layering in for the slice of this dashboard that tracks known vulnerabilities and exposed attack surface, which is a different signal than live traffic anomalies and shouldn't be built from scratch.

Visit Tenable→

Frequently Asked Questions

How do we avoid alert fatigue with all these zero trust signals?

Alert on anomalies and rate of change rather than raw event counts, give every alert a documented first response, and delete or downgrade any alert that hasn't led to real action in the last quarter. Fatigue comes from volume without action, not from having security monitoring at all.

Should certificate expiry alerts page someone or just show on a dashboard?

A dashboard with an escalating schedule works better than a single page: a quiet reminder weeks out, a louder one days out, and only an actual page if it's within 24 hours and still unrenewed. Paging weeks in advance for something with a known, distant deadline just trains people to snooze it.

What's a reasonable first metric to put on a zero trust dashboard?

Start with the rate of authorization denials per service over time, split by whether the denial came from a known caller or an unrecognized one. That single view surfaces both misconfigurations (a known caller suddenly failing) and potential probing (an unrecognized caller trying repeatedly) without needing a dozen separate charts.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides