API Gateways, Management & Edge Security3 min readUpdated September 2026

Kong vs Apigee for Agencies Proxying AI Model Endpoints

Proxying model endpoints breaks assumptions ordinary gateways make. Responses stream for minutes at a time, retries cost real money instead of just latency, and every client needs a separate key with its own spend ceiling.

Those requirements reshape the choice between Kong and Google Cloud Apigee for AI and workflow automation agencies. Kong's plugin model lets you write token accounting and streaming behavior yourself in Lua or Go, while Apigee sells partner governance built for an ecosystem most agencies have not built yet.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Why a model proxying gateway breaks ordinary assumptions

A typical API gateway assumes a request finishes in milliseconds and costs the same regardless of what is in the response body. Neither holds for a proxied model call. A streamed response can stay open for minutes, and the cost of a single request varies with how many tokens the model generates, not just whether the call succeeded.

That means timeout defaults, retry logic, and cost accounting all need to be rebuilt around token counts and stream duration rather than request count. A gateway configured for a normal REST API will either time out long generations or fail to catch a client burning through their allotted spend.

A worked example: metering a client's monthly token spend

Say an agency runs automation workflows for a dozen clients, each with a monthly budget for model calls. Every request needs to be tagged with the client's key, the tokens consumed logged against that key, and a hard stop triggered once the client crosses their ceiling for the month.

On Kong, that means a custom plugin reading the token count out of the model's response body, writing it to a counter keyed by client, and checking that counter against a configured limit before allowing the next request through. It is a real build, but a contained one once the pattern is set.

Building streaming and retry handling in Kong

Kong's plugin model gives you control over the entire request and response lifecycle, which matters for streaming because you need to keep the connection open without triggering the gateway's own idle timeout, and you need retry logic that knows a partially streamed response should not simply be retried from scratch.

This is where Kong's flexibility pays off directly: you write the plugin once, in Lua or Go, and it applies to every client and every model endpoint you proxy going forward.

What Apigee's partner governance actually buys you

Apigee was built for API programs with many external partners consuming a stable set of endpoints under negotiated terms, which describes a mature platform business more than most AI automation agencies at launch. Its monetization and developer portal tooling assumes an ecosystem of partners who are not yet a typical automation agency's client base.

The governance features are real, but they solve a problem that shows up later, once an agency has enough clients and enough standardized offerings that self service onboarding starts to matter more than custom build work.

Picking a default for a multi client agency

Most agencies proxying model endpoints for clients are better served starting with Kong, because the actual hard problems, streaming, retries, and per client token accounting, need custom logic regardless of which gateway sits underneath them. Kong lets you write that logic once and keep it in version control alongside the rest of your infrastructure.

Revisit Apigee once you are selling a standardized product to enough clients that a self service portal and formal partner agreements start to matter more than bespoke integration work.

A common mistake: reusing REST timeout defaults for streamed calls

Teams new to proxying model endpoints often leave a gateway's default request timeout in place, tuned for a REST API that responds in under a second. A long running generation then gets killed mid stream, the client sees a truncated response, and the agency pays for tokens that were generated but never delivered.

The fix is straightforward once you know to look for it: set a separate, much longer timeout policy specifically for streaming routes, and make sure your retry logic checks whether a response was cut off by the gateway before assuming the model itself failed. Catching this during testing with a deliberately long prompt is far cheaper than catching it from a client's support ticket.

Settings to review before proxying a streamed model endpoint:

  • Raise the gateway's request timeout well above the REST default, so a long generation is not killed midstream and left truncated.
  • Keep the connection open while a response streams without tripping the gateway's own idle timeout.
  • Make retry logic aware that a partially streamed response should not simply be retried from scratch, since retries cost real money.
  • Tag each request with the client's key, log the tokens consumed against it, and trigger a hard stop once the client crosses their monthly ceiling.
  • Scope keys strictly per client when several clients share one gateway instance, and isolate a client once their traffic justifies it.

Deciding whether a client needs its own gateway instance at all

Not every client relationship needs a dedicated gateway. A small agency running a handful of low volume automations for a client can sometimes route everything through a single shared instance with strict per client key scoping, deferring the cost of a fully isolated deployment.

The threshold to isolate a client onto their own instance is usually reached once their traffic volume is large enough to affect other clients' latency, once their contract requires isolation explicitly, or once their token spend is large enough that a shared instance's blast radius becomes a real risk if that client's automation misbehaves.

Executive Capability Standard

What Good Looks Like

A well managed AI proxying layer attributes every dollar of model spend to the client and workflow that generated it, and stops a runaway workflow before it exceeds that client's budget rather than after.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Audit which of your current automations call a model endpoint directly instead of through a shared gateway layer.
2. Do Manually:Track each client's monthly token spend in a shared sheet pulled from provider dashboards until the volume makes that unworkable.
3. Delegate:Assign one engineer to own the gateway plugin that handles token accounting and per client limits.
4. Automate:Move spend tracking into a Kong plugin or Apigee policy that enforces the ceiling in real time instead of after the fact.
5. Buy:Adopt a managed AI gateway product once your client count outgrows what a custom plugin can comfortably maintain.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Frequently Asked Questions

Can a gateway meter tokens directly, or does that always require custom code?

Neither Kong nor Apigee parses token counts out of a model provider's response body natively, since token accounting formats differ by provider. Both let you write custom logic, a Kong plugin or an Apigee policy, that extracts the count and applies it against a limit. Expect to build this regardless of which gateway you choose.

How should retries work for a streaming model response that fails partway through?

Retrying from the beginning usually means paying for tokens twice and returning a duplicated response to the client. A safer pattern checkpoints how much of the response the client already received and either resumes from there if the model API supports it, or fails cleanly and lets the calling workflow decide whether to retry the whole request.

Do clients need their own API keys, or can one shared key work across an agency's automations?

Separate keys per client are worth the setup cost almost immediately. A shared key makes it impossible to attribute spend, enforce a per client ceiling, or revoke one client's access without affecting everyone else. Issue keys per client from day one, even if the volume seems small enough to share at first.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides