Model Context Protocol & Agentic ArchitecturePlaybook3 min readUpdated September 2026

How to Actually Compare API Gateways on Latency

Published API gateway benchmarks tend to compare best-case numbers under synthetic, uniform load, which is rarely how your own real traffic actually looks day to day. A gateway that wins a published benchmark can still be the wrong choice for a workload with a very different shape, so the more useful comparison is one you run yourself, against your own traffic pattern.

The goal isn't to distrust every published number on principle. It's to know which parts of a benchmark transfer to your situation and which parts were measured under conditions your production traffic will never actually resemble.

What are you trading against gateway latency?

A gateway's added latency at the median is only part of the picture; the tail, the 95th or 99th percentile, matters more for user experience, and it often comes from a different set of features than the median does: request transformation, authentication checks, or rate limiting logic. Compare gateways on the specific features your traffic will actually exercise, not just a bare pass-through benchmark that neither option will resemble once it's actually handling your production requests.

Why test with your own traffic shape instead of synthetic load?

A gateway benchmarked against a steady stream of identical, small requests behaves differently under your real mix of request sizes, payload types, and burstiness. Capture a representative sample of your actual traffic, or as close an approximation as you can build, and replay it against each candidate gateway rather than trusting a vendor's published number, which was almost certainly measured under conditions chosen to look favorable.

Account for what happens under real failure conditions

A gateway's behavior when a backend is slow or unavailable, whether it fails fast, queues requests, or degrades gracefully, matters as much as its best-case latency number, because backends do fail, and the gateway's behavior during that moment is often what a user actually experiences as an outage. Include a deliberate backend-failure scenario in your comparison, not just a happy-path latency test, since this is often where the real difference between two otherwise similar candidates actually shows up.

Include these scenarios in every gateway comparison:

  • A replay of a representative sample of your real traffic, including its mix of request sizes, payload types and bursts.
  • The specific features your traffic exercises, such as request transformation, authentication checks and rate limiting.
  • Tail latency at the 95th and 99th percentile, not just the median.
  • A deliberate backend failure or slow backend, to see whether the gateway fails fast, queues requests or degrades gracefully.

Weigh operational cost alongside raw performance

The gateway with the best latency number isn't automatically the right choice if it requires meaningfully more operational effort to run, upgrade, and debug than an alternative with slightly worse but still acceptable numbers. Factor in your team's existing familiarity with a given gateway's ecosystem and its debugging tools; a small latency advantage rarely outweighs a large gap in how quickly your team can diagnose a problem with it at 2 a.m., when familiarity matters far more than any single number on a benchmark chart.

To weigh operations against speed, score each candidate on a few practical questions your team can answer from experience. How quickly could someone diagnose a failure with it late at night? Does your team already know its configuration and tooling? How painful are upgrades? For example, if one option is slightly faster but nobody on the team has run it, ask whether the difference will ever matter to users more than a slower diagnosis will. A common mistake is treating the fastest benchmark as the answer. Record the tradeoff explicitly so the choice can be defended later and revisited when the traffic mix changes.

Re-benchmark after any meaningful traffic shift

A comparison done when your traffic looked one way can go stale once your traffic mix changes meaningfully, a new client type, a new region, a different payload size distribution. Treat the benchmark as something to redo when the traffic shape changes significantly, not as a one-time decision that holds indefinitely regardless of how the underlying workload evolves.

Write the methodology down so the next comparison is faster

The first time you build a real traffic-shape benchmark takes real effort: capturing a representative sample, building the replay harness, defining the failure scenario. Document exactly how you did it, so the next evaluation, whether it's a routine re-check or a genuinely new candidate gateway, starts from a working methodology instead of being rebuilt from memory by whoever happens to own the decision next time.

Decide upfront what result would actually change your mind

Before running the comparison, write down what result would justify switching gateways versus what result would confirm your current choice is fine. Without that upfront line, it's easy to run a benchmark, get an ambiguous result, and default to whichever option required less migration effort regardless of what the numbers actually showed, which defeats the purpose of running a rigorous comparison in the first place and wastes the real engineering time the benchmark itself cost to build.

Executive Capability Standard

What Good Looks Like

A fair API gateway comparison tests candidates against your own captured traffic shape including peak load, evaluates behavior under backend failure and not just the happy path, and weighs operational cost and team familiarity alongside raw latency numbers rather than treating latency as the only variable that matters.

Building The Capability (5-Stage Skill Ladder)

1. Learn:read the latency breakdown (median versus tail) your current gateway reports and understand which features contribute most to it
2. Do Manually:capture a sample of real traffic and manually replay it against your current gateway to establish a baseline before comparing alternatives
3. Delegate:assign an engineer to own the benchmark methodology so comparisons stay consistent across evaluation rounds
4. Automate:build a repeatable load-testing harness using your captured traffic sample so re-benchmarking after a traffic shift takes hours, not a full new project
5. Buy:most gateway vendors will run a proof of concept against your own traffic if asked directly; use that instead of relying solely on their published benchmark numbers

How to Get Started

Frequently Asked Questions

Are published API gateway benchmarks worth reading at all?

They're a reasonable starting point for narrowing candidates, but treat the specific numbers as directional rather than predictive for your own workload, since they're typically measured under synthetic, favorable conditions that rarely match real production traffic.

How much traffic do we need to capture for a fair comparison?

Enough to represent your real mix of request types, sizes, and burstiness, including your peak traffic periods, not just an average day. A sample that only reflects quiet-hour traffic will understate how each gateway performs when it actually matters most.

Does gateway choice matter more than application-level optimization?

For most teams, no. The gateway is one layer among several a request passes through, and application or backend latency is frequently the larger contributor. Benchmark the gateway specifically, but don't assume it's the primary lever before checking where your actual latency budget is going.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides