Data Engineering & Real-Time Event StreamsPlaybook3 min readUpdated September 2026

How to Actually Benchmark Your API Gateway's Latency

Benchmark an API gateway honestly by testing with your real payload sizes, a backend with realistic latency, and tail percentiles rather than averages, ideally including a run from inside your own network. Vendor-published numbers usually come from synthetic tests that look nothing like production, so only your own traffic pattern shows what users experience.

This is a methodology for doing that honestly, not a comparison of specific products.

Which payload sizes should you use to benchmark an API gateway?

A gateway benchmark using a tiny JSON payload measures gateway overhead in the best possible case, which is rarely the case your production traffic hits. Pull real payload size distributions from your own traffic and test across that range, including your p95 and p99 sizes, not just the median. Gateways that look nearly identical at a small payload size can diverge meaningfully once payloads get larger, especially anything doing request or response transformation.

Include Your Real Backend Latency, Not a Zero-Latency Mock

Benchmarking a gateway against a backend that responds instantly measures the gateway's pure overhead, which is useful information but not the number that matters for your actual latency budget. Test against a backend with realistic response time distributions, including its own tail latency, so the benchmark reflects how the gateway behaves under conditions that resemble production rather than an idealized best case.

Why measure tail latency instead of the average?

An average latency number hides exactly the behavior that causes real problems: the p99 or p999 requests that time out or feel slow to a user even while the average looks fine. Report percentiles, not a single mean number, and pay particular attention to how tail latency changes as you increase concurrent load, since that's usually where gateways start to diverge from each other and from their own low-load numbers.

Treat Uptime During the Load Test as Part of the Result

A gateway that returns fast responses right up until it falls over under sustained load hasn't actually passed the benchmark, it's failed it in a way average latency numbers won't show. Run the load test long enough and hard enough to find that breaking point, and treat the availability during sustained peak load as a first-class result alongside latency. If your ingress point is a single gateway, its downtime during a real traffic spike is effectively your API's downtime, and at a 99.9% availability target you're working with roughly 8.76 hours of allowed downtime a year total, which a gateway failure under load eats into just as much as any other outage1.

Note how the gateway fails, too, not just when. A gateway that degrades gracefully under overload, shedding load or returning fast errors, is a very different operational situation than one that hangs and slowly takes down every client waiting on a response.

Rerun the Benchmark After Any Config or Traffic Shape Change

A benchmark result from six months ago doesn't reflect a gateway with new plugins, new routing rules, or a traffic pattern that's shifted since then. Treat the benchmark as a recurring check tied to meaningful changes, not a one-time comparison you ran once during vendor selection and never revisited. A gateway that was fast at launch can degrade quietly as configuration complexity accumulates, and the only way to catch that is to keep measuring.

Isolate Whether the Gateway or the Network Is the Bottleneck

A latency number measured from a test client sitting outside your infrastructure conflates gateway processing time with plain network round trip, which makes it hard to tell what you'd actually improve by changing gateways. Run at least one version of the benchmark from inside the same network as production traffic, so you can separate gateway overhead from network distance, and report both numbers rather than a single figure that hides which one is actually the larger contributor.

Do the same decomposition for any middleware the gateway runs, such as authentication checks or request transformation. A gateway that looks slow overall might actually have fast core routing and a slow custom plugin, and that distinction changes what you'd actually fix rather than which vendor you'd consider replacing it with.

A benchmark run in outline:

  1. Pull real payload size distributions from your own traffic and test across the whole range, including the large tail sizes, not just the median.
  2. Test against a backend with realistic response time distributions, including its own tail latency, rather than an instant mock.
  3. Report latency as percentiles and watch how the tail changes as concurrent load increases.
  4. Run the load long enough to find the breaking point, and record availability under sustained peak load as a result.
  5. Repeat one run from inside the production network to separate gateway overhead from network distance.
Executive Capability Standard

What Good Looks Like

A meaningful gateway benchmark uses real payload sizes and realistic backend latency, reports tail percentiles rather than an average, treats availability under sustained load as part of the result, and gets rerun after meaningful configuration or traffic changes rather than trusted as a one-time number.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Pull your real production payload size and backend latency distributions so your next benchmark reflects actual traffic, not a synthetic best case.
2. Do Manually:Run a load test against your gateway using those real distributions and record p50, p95, and p99 latency along with availability under sustained load.
3. Delegate:Assign an engineer to own gateway benchmarking as a recurring practice tied to configuration and traffic changes.
4. Automate:Build the benchmark into your CI or a scheduled job so it reruns automatically rather than depending on someone remembering to do it.
5. Buy:Bring in outside performance engineering help if a benchmark reveals tail latency or stability problems your team doesn't have the bandwidth to diagnose.

How to Get Started

Frequently Asked Questions

Why do vendor-published gateway benchmarks often not match what we see in our own environment?

They're usually run with minimal payloads, a zero-latency mock backend, and low concurrent load, none of which resembles real production traffic. Your own benchmark, run against your actual payload sizes, backend latency, and load pattern, is the only number that reliably predicts what your users will experience.

Is average latency a good enough metric to compare gateways?

No. Average latency hides tail behavior, the p99 or p999 requests that are actually slow or time out, which is usually what causes visible problems for users. Report percentiles and pay particular attention to how tail latency changes under sustained load.

How often should we rerun our gateway latency benchmark?

After any meaningful configuration change, new plugin, or noticeable shift in traffic shape, not just once during initial selection. A gateway's real-world latency can drift as configuration complexity accumulates, and only a recurring benchmark catches that drift.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides