How to Actually Compare API Gateway Latency Claims
Every API gateway vendor publishes a latency number, and almost none of them reflect what you'll actually see once you've added authentication checks, rate limiting, and request transformation on top of a bare proxy. The only benchmark worth trusting is one you run yourself, against your own policy stack and your own traffic shape. Here's how to set that test up so the result actually means something.
Why do vendor gateway benchmarks tell you so little?
Vendor marketing benchmarks are usually measured with a minimal or empty policy chain, no authentication, no rate limiting, no request or response transformation, because that configuration makes every gateway look fast. Your production traffic will run through none of those bare setups. Before trusting any published number, check what policies were active during the test, and if it isn't disclosed, assume it was closer to a bare proxy than to what you'll actually run.
Build a policy stack that matches what you'll actually deploy
Configure the gateway you're testing with the real set of checks your zero-trust setup requires: token validation, per-client rate limiting, and any request transformation your API depends on. Latency scales with how much work the gateway does per request, so a benchmark run with two of your eventual five policies enabled will understate your real cost meaningfully. If you're comparing gateways, hold the policy stack identical across every candidate so the comparison isolates the gateway itself, not a difference in what each one was configured to check.
What traffic should you use to benchmark an API gateway?
A steady stream of identical requests measures something different from your real traffic, which likely has a mix of payload sizes, occasional bursts, and a spread of endpoints with different policy requirements. Replay a sample of real traffic if you have it logged, or construct a synthetic mix that reflects your actual endpoint distribution, and watch tail latency, the 95th and 99th percentile, not just the average, since a gateway that looks fine on average can still produce a bad tail under the specific load pattern your policies create.
Measure where the time actually goes, not just the total
Total request latency through a gateway is the sum of several distinct costs: the network hop itself, token validation against an identity provider, policy evaluation, and any transformation work. Breaking the total down by stage tells you which part is actually expensive and worth optimizing, whereas a single aggregate number just tells you something is slower than you'd like without pointing at what to fix. A gateway with a genuinely slow token validation call, for instance, might be solved by caching validated tokens briefly rather than switching gateways entirely.
For example, suppose a benchmark shows a gateway adding noticeable latency once token validation is enabled. The stage breakdown reveals that nearly all of it comes from a call to the identity provider on every request. Caching validated tokens briefly may remove most of the cost without changing gateways at all. The common mistake is comparing single totals, which would have sent the team shopping for a new product to fix a configuration problem. Keep the breakdown with the result, so anyone rerunning the test can see which stage moved.
Run the same test again after any policy or scale change
A benchmark result is only valid for the configuration and load it was run under. Adding a new policy, increasing your rate limiting granularity, or roughly doubling request volume are all reasons to rerun the test rather than trust a number from six months earlier. Treat gateway benchmarking as a recurring check tied to real changes, not a one time evaluation you did once during the original vendor selection and never revisited.
Include failure and rejection paths in the test, not only successful requests
A benchmark that only measures the happy path misses a real cost: what a request that fails authentication or hits a rate limit actually costs the gateway to process and reject. If a meaningful share of your production traffic is malformed, unauthenticated, or over quota, and for a public API it usually is, exclude that cost from your benchmark and you'll be surprised by real world latency under load that a clean test never predicted. Add a deliberate mix of rejected requests to your test traffic so the number you get back reflects what the gateway actually has to do all day, not just its best case. A gateway that rejects cheaply under load is worth more in practice than one that only looks fast when every request happens to be valid.
A fair benchmark, step by step:
- Check which policies were active in any vendor number, and assume a bare proxy if that is not disclosed.
- Configure every candidate with the same real policy stack: token validation, per-client rate limiting and any request transformation.
- Replay real or realistic traffic, and record 95th and 99th percentile latency, not just the average.
- Break latency down by stage: network hop, token validation, policy evaluation and transformation.
- Include rejected and unauthenticated requests, then rerun the test after any policy or volume change.
What Good Looks Like
A good gateway benchmark runs your actual policy stack against your real traffic shape, measures tail latency broken down by stage, and gets rerun after any meaningful configuration or scale change.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Why do vendor published latency numbers rarely match what we see in production?
Because they're usually measured with a minimal policy chain, no authentication, rate limiting, or transformation active, which makes every gateway look faster than it will once your real zero-trust policy stack is running on top of it.
Should we compare average latency or tail latency when benchmarking a gateway?
Tail latency, the 95th and 99th percentile, matters more for most production decisions, since it reflects the worst experience a meaningful share of your requests will actually have. A gateway with a good average but a bad tail can still cause real customer facing problems.
How often should we rerun a gateway latency benchmark?
Any time you add a policy, change rate limiting configuration, or meaningfully shift traffic volume. A benchmark result only describes the exact configuration and load it was measured under, so a stale result can be actively misleading.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
The Real Latency Cost of Zero Trust, and How to Measure It
How to find out how much latency your zero trust controls actually add, which checks are worth the cost, and which ones you can move off the hot path.
Continuous Device Verification for a Zero-Trust API
How continuous device and identity verification actually works in a zero-trust architecture, and where to draw the line for a small engineering team.
Rolling Out Zero Trust in Production Without a Broad Outage
A checklist for rolling out stricter API authentication and authorization in production, and the pitfalls that turn a rollout into an incident.
Managing Upstream API Rate Limits Before They Break Production
A practical approach to upstream API quota management: how to track headroom, queue gracefully, and avoid a vendor's rate limit taking down your app.
Keeping Auth Checks Fast as Your API Traffic Grows
A worked example for keeping zero trust authorization checks fast as request volume grows, and where teams usually add latency without noticing.
How to Audit Whether Your APIs Actually Enforce Zero Trust
A step-by-step method for testing whether your APIs enforce zero trust in practice, not just on paper, and what to do with what you find.