Benchmarking API Gateway Latency the Right Way
Most API gateway latency numbers you'll find are measured against a single, simple route with minimal middleware, under ideal network conditions. That's a fine way to compare raw baseline overhead. It's a poor way to predict how a gateway will behave with your actual routing rules, your actual authentication middleware, and your actual traffic pattern under load.
This is a methodology for benchmarking a gateway against conditions that resemble what you'll actually run, so the number you get out means something for your specific decision.
Benchmark Your Actual Configuration, Not a Clean Default
Test with the middleware chain you'll actually run: authentication, rate limiting, request transformation, logging, all enabled together, not a stripped-down default configuration. Each middleware layer adds latency, and a gateway that looks fast with everything disabled can look very different once your real configuration is running.
Use your actual route complexity too. A gateway routing to one backend behaves differently under load than one making routing decisions across dozens of services with path-based rules, and a benchmark on the simple case won't predict the complex one.
Test Under Realistic Concurrency, Not a Single Request
A single request's response time tells you almost nothing about how a gateway behaves under real traffic. Run a sustained load test at the concurrency level your actual traffic reaches, including your peak, not just your average, since latency under load often degrades in a way that a light test never reveals.
Measure the full latency distribution, not just the average. The p99, the slowest one percent of requests, is usually what determines whether real users notice a problem, while an average can look perfectly fine even while a meaningful share of requests are struggling.
Run the same load test against more than one candidate gateway using identical traffic, identical middleware where possible, and identical hardware or instance sizing. A comparison where the conditions differ between candidates isn't really a comparison, even if each individual number looks precise.
A benchmark that predicts real behavior does the following:
- Runs with the full middleware chain you will actually use, including authentication, rate limiting, request transformation and logging, all enabled together.
- Uses your real route complexity and a sustained load at your peak concurrency, not a single request or an average day.
- Reports the full latency distribution, especially the slowest one percent of requests, instead of only the average.
- Gives every candidate gateway identical traffic, middleware and instance sizing, so the comparison is fair.
- Includes a realistic mix of cold and warm requests, and reuses connections the way real clients do.
Account for Cold Starts and Connection Overhead
If your gateway or its backends run on infrastructure with cold starts, serverless functions being the most common case, a benchmark that only measures warm, already-running requests misses a real and sometimes significant chunk of your actual user experience. Include a mix of cold and warm requests if that reflects your real traffic pattern.
Connection reuse matters too. A benchmark that opens a fresh connection for every request measures something different than one that reuses persistent connections the way most real clients and load balancers actually do.
What Latency Numbers Don't Tell You
A gateway can have excellent raw latency and still be the wrong choice if its failure behavior under overload is bad, dropping requests ungracefully instead of shedding load predictably, or if its operational tooling makes debugging a production issue slower than the milliseconds it saved you.
A gateway sitting in the request path for every service is also a single point that affects your overall availability, so weigh its failure behavior, not just its speed, against how much downtime budget you're willing to risk on it.
For example, suppose two gateways land within noise of each other on latency under your real middleware chain. Latency stops being the deciding factor, and overload behavior takes over. Push traffic past your expected peak on purpose and watch what each one does. One may shed load predictably and return clear errors, while another drops requests ungracefully or stalls. Then compare how quickly an engineer can find the cause of a slow request in each product's tooling. The gateway that fails gracefully and is easy to debug is usually the safer choice, even if it is a few milliseconds slower.
Reading Vendor-Published Benchmarks Skeptically
A vendor's own published benchmark is measuring the configuration that makes them look best, which is rarely the configuration you'll actually run. Treat any published number as a starting hypothesis to verify against your own realistic test, not as a number you can plug directly into a capacity or vendor decision.
What to confirm yourself in any trial or demo: latency under your real middleware chain and your real concurrency, not the vendor's clean-room number.
Ask the vendor directly what configuration their published number reflects, and whether they'll support a trial where you can run your own realistic benchmark rather than taking their number as given. A vendor confident in their real-world performance usually has no problem with that request.
What Good Looks Like
Good gateway benchmarking means testing your actual middleware chain and route complexity under realistic concurrency, measuring the full latency distribution rather than just the average, and weighing failure behavior alongside raw speed.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Why do our real gateway latency numbers not match the vendor's published benchmark?
Because vendor benchmarks are usually measured against a minimal configuration, often a single simple route with little or no middleware, which is rarely how you'll actually run it. Authentication, rate limiting, routing complexity, and real concurrency all add latency that a clean-room benchmark doesn't capture.
Should we measure average latency or something else when benchmarking a gateway?
Look at the full distribution, especially the p99, the slowest one percent of requests, rather than relying on the average alone. An average can look healthy even while a meaningful share of real requests are noticeably slow, and that slower tail is usually what users actually notice.
Does gateway benchmarking need to account for cold starts?
Yes, if your gateway or its backends run on infrastructure that has them, serverless functions being the most common case. A benchmark that only measures already-warm requests misses a real part of the user experience, so include a realistic mix of cold and warm requests if that matches your actual traffic.
Is raw latency the only thing that matters when choosing an API gateway?
No. A gateway with excellent raw latency can still be a poor choice if it handles overload badly, dropping requests instead of shedding load predictably, or if its tooling makes debugging a real incident slower. It's also a single point in your request path, so its failure behavior matters as much as its speed.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Finding the Real Source of Latency in a Distributed System
A decision guide for narrowing down whether a slow request is a network problem, a database problem, a queue problem, or your own code.
Keeping API Contracts From Breaking Between Services
A practical standard for versioning, owning, and validating API contracts so one team's change doesn't quietly break three other services.
A Production Deployment Checklist That Actually Catches Problems
A stage-by-stage deployment checklist for distributed systems, covering rollback readiness, dependency ordering, and the checks teams skip under pressure.
A Runbook for When an Upstream API Starts Throttling You
A step-by-step runbook for handling upstream API throttling: detecting it fast, absorbing it without cascading failures, and fixing the root cause.
Verifying Devices Before They Touch Production, Not After
How to build device verification into a zero-trust rollout, what actually counts as a trust signal, and where teams stop checking too early.
Load Testing Numbers That Don't Match What Users Actually Feel
Why a clean throughput benchmark often fails to predict real-world scaling behavior, and how to build one around your real traffic mix and first bottleneck.