Benchmark Your Own Gateway Before You Trust Anyone Else's Numbers
Every API gateway vendor publishes latency numbers, and every one of those numbers was measured under conditions chosen to look good: a synthetic payload, a clean network path, a request shape nothing like what your actual traffic will send through it. That doesn't make the numbers dishonest. It makes them irrelevant to the decision you're actually trying to make.
The only latency number worth trusting is one you measured yourself, against your own traffic.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Vendor numbers are measured on their best day, not yours
A published p50 latency figure usually comes from a minimal request, a small payload, no authentication overhead, run against infrastructure tuned specifically for that benchmark. Your production traffic carries auth checks, payload transformation, rate limiting logic, and a request size distribution the vendor's number never accounted for. Treat any vendor benchmark as a ceiling on what's possible under ideal conditions, not a prediction of what you'll actually see.
Measure the percentile that matches what customers feel
Average latency hides the requests that were genuinely slow behind the much larger number that were fine. p50 tells you about the typical request; p99 tells you about the request your slowest, most frustrated customers are actually experiencing. If you only track one number, track p99, since it's the one most likely to correlate with support tickets and churn, not the one that makes a dashboard look the best.
For example, when comparing two gateways, put the p50 and p99 for each side by side against the same replayed traffic, and note which requests sit in the slow tail. If the slowest requests are the large uploads, the tail belongs to payload handling, not routing, and that calls for a different question to the vendor than a general one about speed. Ask each vendor to explain the tail on your traffic. A vendor that can only point back to its published benchmark hasn't answered the question you are actually deciding on.
Build a benchmark that matches your real traffic shape
Replay a sample of actual production request logs, not a synthetic payload generator, against any gateway you're evaluating, including your current one as the baseline. Match the real distribution of payload sizes, authentication patterns, and endpoint mix, since a gateway that looks fast against small, uniform synthetic requests can behave very differently once it's handling your actual mix of large uploads, small polling requests, and everything in between.
Build a fair gateway benchmark in this order:
- Sample real production request logs to replay, rather than using a synthetic payload generator.
- Match the real mix of payload sizes, authentication patterns and endpoints in the replayed traffic.
- Run the same replay against your current gateway as the baseline for comparison.
- Turn on every feature you'll use in production together, including auth, rate limiting and transformation, instead of testing each alone.
- Compare p99 as well as p50, since the slowest requests drive support tickets and churn.
- Rerun the whole test quarterly at minimum, and after any meaningful change.
Where gateway overhead compounds in ways you don't expect
Each individual feature, auth token validation, rate limiting, request transformation, response caching, adds a small amount of latency on its own, but they compound, and the combined overhead is rarely the simple sum of each feature's advertised cost in isolation. Benchmark with every feature you'll actually run in production turned on together, not one at a time, since a gateway that's fast with authentication alone and fast with rate limiting alone can still be surprisingly slow with both running at once.
Don't let benchmarking distract from patching the gateway itself
A gateway sitting at the edge of your infrastructure, handling every request before it reaches anything else, is also one of your highest value patching targets: a known exploited vulnerability in gateway software carries a fourteen day remediation clock under federal guidance for CVEs assigned since 20211. Benchmarking performance and staying current on security patches are separate workstreams, and it's worth explicitly confirming neither one is quietly starving the other of attention.
A worked example: what a synthetic benchmark missed
Say a synthetic benchmark shows a candidate gateway beating your current one by a wide margin on small, uniform requests. Once you replay real production traffic, including the ten percent of requests carrying large file uploads, the gap narrows sharply or reverses, because the candidate's transformation layer handles large payloads less efficiently than your current setup. That gap only shows up once you test with your actual traffic shape, which is exactly why the synthetic number alone would have led to the wrong decision.
Rerun the benchmark after any meaningful change, not just once
A benchmark run once during initial evaluation goes stale the moment traffic patterns shift, a new endpoint gets added, or the gateway's configuration changes to support a new feature. Treat the replay benchmark as a recurring check, quarterly at minimum, rather than a one-time decision made during procurement. The gateway that won the original comparison can quietly become the slower option eighteen months later if nobody's checked since, and the only way to know is to actually run the test again, not assume the original result still holds.
What Good Looks Like
Good gateway benchmarking means testing against replayed real traffic with every planned feature enabled together, tracking p99 rather than average latency, and keeping the gateway's own patch cadence current alongside the performance work.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
Are vendor-published latency benchmarks worth looking at all?
They're a reasonable starting filter, a gateway with a terrible synthetic number is unlikely to be fast for you either, but treat them as a ceiling under ideal conditions, not a prediction. The only number that actually informs a decision is one measured against your own traffic shape.
Should we benchmark every gateway feature separately or all at once?
Benchmark with every feature you'll actually run in production enabled together. Individual feature overhead compounds in ways that don't show up when each one is tested in isolation, so a combined test is the only one that reflects what production will actually look like.
How much of our real traffic do we need to replay for a meaningful benchmark?
Enough to capture your actual distribution of payload sizes, endpoint mix, and authentication patterns, not a fixed percentage. A smaller sample that faithfully represents the shape of your traffic is more useful than a larger one that's still mostly uniform, synthetic-style requests.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Security patch remediation SLAs (CISA federal mandates, used as industry norm). CISA Binding Operational Directives 19-02 and 22-01 (CISA briefing hosted at NIST CSRC), 2022.
Related Guides
Diagnosing Slow Requests Before You Blame the Database
A step-by-step way to find out whether a slowdown is the network, the app, or the database, before you add caching or upgrade infrastructure to fix it.
Build or Buy for Verifying Every Device That Connects?
How to split device identity from device posture checking, what building either one in house actually costs, and where a platform earns its keep instead.
How to Ship a Risky Change Without a 2am Rollback
A concrete walkthrough of how to plan a risky production deployment: how to split it, what to watch, and when to decide the rollback trigger.
Deciding How to Handle Upstream API Rate Limits Before They Hit You
Choose between a higher API quota, caching and batching, or a queue when a third-party rate limit becomes a real constraint on your product.
Building a Throughput Benchmark You Can Actually Trust
A worksheet approach to benchmarking throughput: what load pattern to test, what to record, and how synthetic benchmarks lie about real capacity.
The API Standards Worth Enforcing, and the Ones That Aren't
Which API integration standards actually prevent problems, which ones are busywork, and how to tell the difference before you write a style guide.