API Security, Identity & Zero-TrustPlaybook3 min readUpdated September 2026

How to Audit Whether Your APIs Actually Enforce Zero Trust

A zero trust diagram is easy to draw and hard to prove. The real test isn't whether you wrote a policy, it's whether a request carrying a stolen token, an expired certificate, or the wrong scope actually gets rejected, every time, at every service.

This audit method is built for a team that has some auth in place already and wants to know where it actually breaks, not a from-scratch design exercise.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

How do you map every path into your APIs before testing?

Start with a plain list, not a diagram: every public endpoint, every internal service-to-service call, every admin route, every webhook a partner can hit, and every path CI/CD uses to deploy or roll back. Teams that skip this step end up auditing the five endpoints they remember and missing the internal billing service that still takes an unauthenticated call from "trusted" infrastructure.

Pull this list from your API gateway config, your service mesh (if you run one), and your load balancer rules, not from memory or an old architecture doc. Cross-check it against actual traffic logs for a week; anything receiving requests that isn't on your list is itself a finding.

Two categories get missed almost every time: health-check and metrics endpoints that were never meant to be reachable from outside the cluster but ended up exposed by a misconfigured load balancer rule, and old API versions that engineering assumed were retired but a handful of clients are still quietly calling. Add both to your list explicitly rather than assuming they're covered by "everything else."

How do you test whether zero trust is actually enforced?

Pick ten endpoints across different services and actually attempt the bypass: call one with an expired token, one with the right token but the wrong scope, one over plain HTTP where mTLS should be required, one from an IP outside your expected ranges. Write down what happened, not what should have happened.

This is where most audits find the gap between the architecture doc and production. A service that "requires" a valid JWT will often still process the request and only fail on a later step, quietly leaking data before the rejection.

Find where trust still gets assumed by network location

Zero trust means no service gets a pass because it's "inside the VPC." Go through your service-to-service calls and flag any that skip authentication because the caller is on the same subnet, in the same Kubernetes namespace, or behind the same load balancer. This pattern survives longest in older internal tools: an admin dashboard, a reporting job, a batch script someone wrote three years ago and never touched again.

For each one you find, decide whether it gets a real identity and authentication now or goes on a dated remediation list. Don't leave it undecided.

Score identity separately from the perimeter

A firewall rule is not an identity. Check whether each service authenticates as itself (a workload certificate, a signed service token, a SPIFFE identity) or whether it's sharing a static API key with five other services. Shared keys are the most common finding in this kind of audit: they're easy to rotate on paper and almost never actually rotated, because nobody can list everything that would break.

Separately, check human access: are engineers using personal accounts with MFA to reach production APIs, or is there a shared "ops" credential floating around in a password manager? Those need different fixes.

A quick way to find shared keys without reading every config file is to look at your traffic logs for a single credential making calls from more than one service's network location; that pattern almost always means the key is shared rather than scoped to one workload.

Give every finding an owner and a date, not a slide

An audit that ends in a report nobody reopens didn't happen. Every finding needs a named owner, a target date, and a severity that maps to how exposed it is: an unauthenticated internal-only endpoint is not the same severity as a public API accepting expired tokens. Federal patch guidance for internet-facing, critical vulnerabilities sets a genuinely tight remediation window, and it's a reasonable anchor for how fast your own public-facing findings should move1.

Put the list somewhere your team already looks, a sprint board, not a shared drive folder, and revisit it on a fixed schedule rather than waiting for the next audit to notice it never got fixed.

Run the audit in this order:

  1. List every public endpoint, internal call, admin route, webhook and deployment path, using gateway config and traffic logs rather than memory.
  2. Attempt real bypasses on a sample of endpoints, using expired tokens, wrong scopes, plain HTTP and unexpected IP ranges, and record what happened.
  3. Flag every service call that skips authentication because the caller sits on the same subnet, namespace or load balancer.
  4. Check whether each service authenticates as itself or shares a static API key with other services.
  5. Assign every finding a named owner, a target date and a severity based on how exposed the affected endpoint is.
Executive Capability Standard

What Good Looks Like

Good enforcement means every request, human or machine, internal or external, is authenticated and authorized at the service that handles it, and you have current, testable evidence of that rather than a policy document.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Read through your own gateway, mesh, and IAM configuration until you can list, from memory, which services trust which callers and why.
2. Do Manually:Pick ten endpoints and manually attempt the bypasses described above, recording exactly what happened.
3. Delegate:Assign one senior engineer to own the findings list end to end, including chasing down owners for fixes outside their own team.
4. Automate:Move from a point-in-time spreadsheet to continuous evidence collection so the audit trail updates itself as services change, rather than going stale the day the report ships.
5. Buy:Bring in an outside penetration tester or a fractional security lead for your first fully independent audit, since an internal team auditing its own work misses its own blind spots.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Frequently Asked Questions

How long should a first API security audit take?

For a team with a handful of services, plan on two to three weeks to inventory every path in, test enforcement on a representative sample, and write up findings with owners. A larger, more fragmented service estate takes longer mostly because the inventory step takes longer, not because the testing does.

Is this the same thing as a penetration test?

No. A penetration test tries to find any way in, often through a narrow slice of your system, using techniques an attacker would use. This audit is broader and more mundane: it checks whether the authentication and authorization you believe exists is actually enforced everywhere. Many teams do both, in that order.

Do we need a specialized tool to run this audit?

Not for the first pass. A spreadsheet, your gateway and mesh configs, and a way to send test requests with bad credentials will get you through it. Continuous compliance platforms help later, once you want the evidence trail to stay current between audits instead of going stale the day after you finish.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Security patch remediation SLAs (CISA federal mandates, used as industry norm). CISA Binding Operational Directives 19-02 and 22-01 (CISA briefing hosted at NIST CSRC), 2022.

Related Guides