Developer Productivity & Platform EngineeringPlaybook3 min readUpdated September 2026

How to Benchmark an API Gateway Without Fooling Yourself

Gateway benchmarks are easy to run and easy to get wrong. A number from a load test that doesn't resemble your real traffic pattern is worse than no benchmark at all, because it gives you false confidence in a decision that hasn't actually been tested.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

How do you benchmark with your actual request shape?

A generic benchmark hammers a single simple endpoint with identical requests. Real traffic is a mix: some requests hit cache, some don't, payload sizes vary, and some routes do far more work than others. Build your test traffic from a sample of real request logs, not a synthetic script that only exercises the easiest path.

This single change, testing with realistic request shape instead of a generic pattern, is usually the difference between a benchmark that predicts real behavior and one that just produces an impressive-looking number.

Why measure tail latency instead of the average?

An average latency number can look fine while a meaningful share of requests are genuinely slow. Look at your higher percentiles specifically, since those are the requests that actually generate complaints and support tickets, even when the average across everything looks healthy.

A gateway that looks great on average latency but has a long tail of slow outliers is a worse choice in practice than one with a slightly higher average and a much tighter, more predictable tail.

Test Under Realistic Concurrency, Not Just Realistic Volume

Total requests per second matters less than how many of those requests are happening at the same moment your system has to handle other things too, deploys, background jobs, database maintenance. Run at least one benchmark pass concurrent with a simulated deploy or a heavy background job, since that's closer to a real bad day than a clean, isolated test window.

How often you actually face this concurrency depends partly on your deployment frequency, since a higher deployment frequency means more windows where gateway load and deploy load genuinely overlap in production1.

Compare Self-Hosted and Managed Options on Operational Cost, Not Just Speed

A self-hosted gateway can often win a raw latency benchmark while losing badly on the operational cost of running it yourself, patching it, scaling it, and being the one paged when it has a bad night. A managed gateway service can lose the latency benchmark by a small margin while saving significant engineering time better spent elsewhere.

Decide up front which tradeoff matters more for your team's current size and priorities, and weigh the benchmark result against that decision rather than treating raw latency as the only number that counts.

Rerun the Benchmark When Traffic Actually Changes

A benchmark run once at launch goes stale as your traffic pattern shifts, new endpoints get added, and request volume grows. Rerun it whenever traffic composition changes meaningfully, not on a fixed calendar schedule disconnected from what's actually happening to your system.

Write Down What the Benchmark Was Actually Testing

Six months later, nobody remembers exactly what traffic shape or concurrency scenario a benchmark used, which makes the old numbers useless for deciding whether a regression is real or just a different test setup. Record the methodology alongside the results: what sample of requests, what concurrency, what background load was running at the same time.

That record is what turns a one-off benchmark into something you can actually compare against next time, instead of starting the whole design process over from scratch every time a gateway decision comes up again.

A benchmark you can trust follows these steps:

  1. Build test traffic from a sample of real request logs so it reflects cache hits, varied payloads, and uneven route costs.
  2. Report higher percentiles alongside the average, since slow outliers generate the complaints and support tickets.
  3. Run at least one pass concurrently with a simulated deploy or heavy background job.
  4. Weigh self-hosted against managed options on operational cost as well as raw latency.
  5. Record the methodology with the results and rerun the benchmark when traffic composition changes.

Don't Let a Vendor's Own Benchmark Replace Yours

A vendor's published benchmark numbers are usually run against traffic and hardware chosen to make that vendor look good, and there's nothing wrong with that, it's marketing, not deception. The mistake is treating those numbers as a substitute for testing against your own request shape and your own concurrency pattern.

Use a vendor's own numbers as a rough first filter to narrow your options, then always run your own benchmark before committing to a migration, since the gap between a marketing number and your real workload's number can be substantial.

For example, a vendor publishes a strong throughput number measured on a single simple route with identical requests. Your traffic mixes cached reads, large uploads, and a few expensive routes. Use the published figure only to shortlist candidates, then replay a sample of your own request logs against each one. You may find that the vendor with the lower headline number has tighter tail latency on your mix, which is the number your users feel. Keep both results in the write-up so the next comparison starts from a known baseline.

Executive Capability Standard

What Good Looks Like

A meaningful gateway benchmark uses realistic request shape and concurrency, measures tail latency rather than just the average, and gets rerun as traffic composition actually changes.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Pull a sample of real request logs to understand your actual traffic shape before designing a benchmark.
2. Do Manually:Manually run a benchmark using that real request sample and record tail latency, not just the average.
3. Delegate:Assign an engineer ownership of the benchmark methodology so it's reused consistently instead of reinvented each time.
4. Automate:Automate a recurring benchmark run tied to meaningful traffic changes rather than a fixed calendar date.
5. Buy:Bring in outside performance engineering expertise before a gateway migration if your team hasn't run a load test at this scale before.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Frequently Asked Questions

How long should a gateway benchmark run to be meaningful?

Long enough to capture your real traffic variation, typically at least the length of a full business day cycle if your traffic has daily patterns, rather than a short burst test that only captures one moment in time.

Is a higher requests-per-second number always the better gateway choice?

No. A gateway with a lower raw throughput number but tighter, more predictable tail latency under your actual traffic shape is often the better real-world choice than one that wins on a synthetic peak-throughput number alone.

Should we benchmark with production-level traffic volume before switching gateways?

Yes, as close as you safely can. A benchmark at a fraction of real volume can hide problems that only appear under genuine load, like connection pool exhaustion or garbage collection pauses that don't show up in a smaller test.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.

Related Guides