How to Actually Benchmark Your API Gateway's Latency
Benchmark an API gateway honestly by testing with your real payload sizes, a backend with realistic latency, and tail percentiles rather than averages, ideally including a run from inside your own network. Vendor-published numbers usually come from synthetic tests that look nothing like production, so only your own traffic pattern shows what users experience.
This is a methodology for doing that honestly, not a comparison of specific products.
Which payload sizes should you use to benchmark an API gateway?
A gateway benchmark using a tiny JSON payload measures gateway overhead in the best possible case, which is rarely the case your production traffic hits. Pull real payload size distributions from your own traffic and test across that range, including your p95 and p99 sizes, not just the median. Gateways that look nearly identical at a small payload size can diverge meaningfully once payloads get larger, especially anything doing request or response transformation.
Include Your Real Backend Latency, Not a Zero-Latency Mock
Benchmarking a gateway against a backend that responds instantly measures the gateway's pure overhead, which is useful information but not the number that matters for your actual latency budget. Test against a backend with realistic response time distributions, including its own tail latency, so the benchmark reflects how the gateway behaves under conditions that resemble production rather than an idealized best case.
Why measure tail latency instead of the average?
An average latency number hides exactly the behavior that causes real problems: the p99 or p999 requests that time out or feel slow to a user even while the average looks fine. Report percentiles, not a single mean number, and pay particular attention to how tail latency changes as you increase concurrent load, since that's usually where gateways start to diverge from each other and from their own low-load numbers.
Treat Uptime During the Load Test as Part of the Result
A gateway that returns fast responses right up until it falls over under sustained load hasn't actually passed the benchmark, it's failed it in a way average latency numbers won't show. Run the load test long enough and hard enough to find that breaking point, and treat the availability during sustained peak load as a first-class result alongside latency. If your ingress point is a single gateway, its downtime during a real traffic spike is effectively your API's downtime, and at a 99.9% availability target you're working with roughly 8.76 hours of allowed downtime a year total, which a gateway failure under load eats into just as much as any other outage1.
Note how the gateway fails, too, not just when. A gateway that degrades gracefully under overload, shedding load or returning fast errors, is a very different operational situation than one that hangs and slowly takes down every client waiting on a response.
Rerun the Benchmark After Any Config or Traffic Shape Change
A benchmark result from six months ago doesn't reflect a gateway with new plugins, new routing rules, or a traffic pattern that's shifted since then. Treat the benchmark as a recurring check tied to meaningful changes, not a one-time comparison you ran once during vendor selection and never revisited. A gateway that was fast at launch can degrade quietly as configuration complexity accumulates, and the only way to catch that is to keep measuring.
Isolate Whether the Gateway or the Network Is the Bottleneck
A latency number measured from a test client sitting outside your infrastructure conflates gateway processing time with plain network round trip, which makes it hard to tell what you'd actually improve by changing gateways. Run at least one version of the benchmark from inside the same network as production traffic, so you can separate gateway overhead from network distance, and report both numbers rather than a single figure that hides which one is actually the larger contributor.
Do the same decomposition for any middleware the gateway runs, such as authentication checks or request transformation. A gateway that looks slow overall might actually have fast core routing and a slow custom plugin, and that distinction changes what you'd actually fix rather than which vendor you'd consider replacing it with.
A benchmark run in outline:
- Pull real payload size distributions from your own traffic and test across the whole range, including the large tail sizes, not just the median.
- Test against a backend with realistic response time distributions, including its own tail latency, rather than an instant mock.
- Report latency as percentiles and watch how the tail changes as concurrent load increases.
- Run the load long enough to find the breaking point, and record availability under sustained peak load as a result.
- Repeat one run from inside the production network to separate gateway overhead from network distance.
What Good Looks Like
A meaningful gateway benchmark uses real payload sizes and realistic backend latency, reports tail percentiles rather than an average, treats availability under sustained load as part of the result, and gets rerun after meaningful configuration or traffic changes rather than trusted as a one-time number.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Why do vendor-published gateway benchmarks often not match what we see in our own environment?
They're usually run with minimal payloads, a zero-latency mock backend, and low concurrent load, none of which resembles real production traffic. Your own benchmark, run against your actual payload sizes, backend latency, and load pattern, is the only number that reliably predicts what your users will experience.
Is average latency a good enough metric to compare gateways?
No. Average latency hides tail behavior, the p99 or p999 requests that are actually slow or time out, which is usually what causes visible problems for users. Report percentiles and pay particular attention to how tail latency changes under sustained load.
How often should we rerun our gateway latency benchmark?
After any meaningful configuration change, new plugin, or noticeable shift in traffic shape, not just once during initial selection. A gateway's real-world latency can drift as configuration complexity accumulates, and only a recurring benchmark catches that drift.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
Where Latency Actually Hides in a Growing Data Pipeline
A walkthrough of where latency hides as a real-time pipeline grows, from producer batching to consumer lag, so you can find your own bottleneck fast.
Blue-Green, Canary, or Rolling: Deploying Stream Processors
A decision guide to rolling, blue-green, and canary deploys for stateful stream processors, plus the rollback plan most teams never actually test.
Verifying Every Service That Talks to Your Pipeline
Which parts of zero-trust verification to build and which to buy, so every producer and consumer on a streaming pipeline proves its identity.
Webhooks, Polling, or a Real Event Stream: Choosing an Integration
A comparison of webhooks, polling, and true event streaming for connecting systems, with the tradeoffs that actually decide which one fits your case.
Stopping a Rate Limited Upstream API From Taking Down Your Pipeline
How to design an ingestion pipeline so a rate limited third party API degrades gracefully instead of cascading into a full outage.
Finding Your Pipeline's Actual Throughput Ceiling
A worked example of finding a real-time pipeline's actual throughput ceiling, and why partition count usually matters more than raw consumer horsepower.