Engineering Leadership & Technical HiringPlaybook3 min readUpdated September 2026

Why Your API Gateway Load Test Doesn't Match Production

Vendor benchmarks for API gateways are almost always measuring the best possible case: a minimal route, no authentication, no rate limiting, run against hardware tuned specifically for the test. That number tells you almost nothing about how the gateway will perform once it's sitting in front of your actual traffic, doing the actual work you need it to do.

A benchmark worth trusting has to run your workload, with your middleware turned on, against your traffic shape, not a stripped down synthetic route designed to look impressive in a chart.

Why the Vendor Number on the Landing Page Doesn't Apply to You

A gateway advertising a huge requests per second figure is almost always measuring a bare route with no authentication check, no rate limiting, no request transformation, and no logging, because every one of those adds latency and every one of those is something you'll actually have turned on in production.

Treat the marketing number as an upper bound under ideal conditions, not a prediction of your own performance. The gap between that number and what you'll actually see is usually driven entirely by which middleware you enable, which is exactly the part a synthetic benchmark leaves out.

Building a Benchmark That Reflects Your Actual Traffic

Start from your real traffic shape: the mix of route types, payload sizes, and middleware your production system actually runs, not a single, simplest possible endpoint. If eighty percent of your traffic hits routes with authentication and rate limiting enabled, benchmark that configuration, not a bare passthrough route that represents almost none of your real requests.

Replay a sample of real production traffic against the gateway if you can, rather than synthetic load, since synthetic traffic generators tend to produce a far more uniform request pattern than real users ever do, and that uniformity can hide latency spikes that only show up under a realistic, uneven load.

What to Measure Beyond Raw Throughput

Throughput alone hides the number that actually matters to your users: tail latency, the ninety-ninth percentile response time, not the average. A gateway with excellent average latency and a bad tail can still produce a steady trickle of slow requests that show up as real user complaints, even while every dashboard built around averages looks fine.

Measure latency under sustained load over a period long enough to reveal degradation, not a short burst test. Some gateways perform well for the first minute and then degrade as connection pools or internal caches fill up, a pattern a thirty second test will never catch.

Comparing Gateways: What Actually Differs Between Them

The meaningful differences between gateways usually aren't in raw throughput, most modern gateways can handle far more traffic than a small or mid sized company will ever send through one. The meaningful differences are in operational fit: how easy the configuration is to reason about and version control, how the gateway behaves under partial failure of an upstream service, and how observable its own internal behavior is when something goes wrong.

Run your benchmark with an upstream service deliberately failing or slow, not just healthy, since gateway behavior under a struggling backend, does it retry sensibly, does it fail fast, does it isolate the failure, matters more day to day than its peak throughput on a good day.

For example, suppose two gateways both clear your projected peak traffic on a healthy backend. Run the same replay again with one upstream service slowed down and watch what each gateway does: does it retry sensibly, fail fast, or let queued requests pile up and drag unrelated routes down with it? The gateway that isolates the struggling backend is usually the better daily choice, even if its peak throughput is a little lower. Also note how easy each configuration was to read and version control during the test, since that is the part your team will live with long after the benchmark chart is forgotten.

Turning Benchmark Results Into a Decision, Not Just a Chart

A benchmark is only useful if it changes a decision. Before running one, write down the specific question it needs to answer: can this gateway handle our projected peak traffic with our full middleware stack enabled and stay under our latency target, not a general "which gateway is fastest."

Re-run the benchmark against any gateway you're seriously considering with the same test harness and the same traffic replay, so the comparison is actually apples to apples. A benchmark run once against one candidate, compared against a vendor's marketing number for another, isn't a comparison at all.

Check a benchmark against this list before you trust its result:

  • It runs your real middleware stack, including authentication, rate limiting, request transformation, and logging, instead of a bare passthrough route.
  • It matches your production mix of route types and payload sizes, ideally by replaying a sample of real traffic.
  • It reports tail latency at the ninety-ninth percentile, not just averages or raw throughput.
  • It holds sustained load long enough to expose connection pool or cache degradation, and repeats the run with a slow or failing upstream.
  • It uses the same harness and traffic replay for every candidate, so the comparison is apples to apples.
Executive Capability Standard

What Good Looks Like

Good gateway benchmarking means testing your actual traffic shape and full middleware stack under sustained load, measuring tail latency and behavior under upstream failure, not just a vendor's best-case throughput number.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Document your actual production traffic shape: route mix, payload sizes, and which middleware runs on each route.
2. Do Manually:Manually run a load test against one gateway candidate with your real middleware enabled and record tail latency, not just average.
3. Delegate:Assign an engineer ownership of building a repeatable benchmark harness that replays realistic traffic against any candidate gateway.
4. Automate:Run the benchmark harness automatically against new gateway versions so a performance regression is caught before it ships to production.
5. Buy:Bring in a performance engineer to design the benchmark if your team has never built a load test that goes beyond a vendor's own numbers.

How to Get Started

Frequently Asked Questions

Should we trust a vendor's published benchmark numbers at all?

Treat them as a rough ceiling under ideal conditions, not a prediction. They're useful for ruling out a gateway that can't even hit your requirements in the best case, but not for choosing between two gateways that both clear that bar, since the real differences show up under your own middleware and traffic shape.

How long should a load test run to be meaningful?

Long enough to catch degradation that only appears after sustained load, connection pool exhaustion or cache pressure, which a short burst test won't reveal. Several minutes of sustained traffic at your target load is a more reliable signal than a thirty second spike test.

Does gateway choice matter much for a small company's traffic volume?

Raw throughput rarely matters at a small company's scale, since most gateways can handle far more than you'll send them. What matters more at that scale is operational fit: how easy it is to configure, debug, and reason about failure, which is worth weighting more heavily than a throughput number you'll never come close to needing.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides