AI Model Serving & Inference OptimizationPlaybook3 min readUpdated September 2026

Where to Terminate TLS in Your Model Serving Path

Terminate TLS at the gateway if your internal network is genuinely isolated, and carry mutual TLS to the model server when it handles sensitive input or serves multiple tenants. Every hop toward the model server is a place traffic could be intercepted if it is unencrypted, so the decision turns on where your trusted network boundary sits.

The decision is really about where you draw the boundary of your trusted network, and how confident you actually are in that boundary holding over time, not just on the day it was first configured.

Should You Terminate TLS at the Gateway?

Ending encryption at your API gateway and running plain, unencrypted traffic between the gateway and internal model servers is the simplest setup, and it is a reasonable choice if that internal network segment is genuinely isolated and access to it is tightly controlled. The risk is that this simplicity depends entirely on that internal isolation holding, and a misconfigured security group or an unexpected peering connection can quietly turn what you assumed was a private segment into something reachable from further away than intended.

Carrying mTLS All the Way to the Model Server

Mutual TLS all the way to the model server means every hop authenticates both directions, so even a service that finds itself on the same network segment cannot simply connect without a valid certificate. This closes the gap that a gateway-only termination leaves open if internal isolation ever fails, at the cost of managing and rotating certificates for every service in the path. For a model server handling sensitive input or serving multiple tenants, that cost is usually worth it.

A Middle Path: Encrypt Everywhere, Authenticate Where It Matters Most

A reasonable middle ground for many teams is encrypting every hop but reserving full mutual authentication for the boundary between your network and the model server itself, where the consequences of a bypass are highest. This avoids the full certificate management burden of mTLS everywhere while still closing the specific gap, an unauthenticated internal caller reaching the model server, that matters most. It also gives you a natural place to start if you are retrofitting encryption onto an existing setup: get the highest-risk boundary right first, then decide deliberately whether the rest of the path needs the same treatment.

Why Does Certificate Rotation Usually Break?

The design decision above matters less than whether certificate rotation actually happens reliably. An expired certificate silently failing over to plain connections, or causing an outage because nothing rotated it in time, is a more common real-world failure than the encryption boundary being wrong in the first place. Automate rotation and alert on an approaching expiration well before it happens, rather than relying on someone remembering a manual renewal date. Test the alert itself periodically too, by checking that it would actually fire with enough lead time to act, not only that it exists somewhere in a configuration file.

For example, list every certificate in the request path with its expiry date and the mechanism that renews it. Any certificate with no automated renewal, or whose expiry alert has never been tested, is your first fix. When a certificate is replaced, confirm the new one is actually presented and that the destination still rejects callers without a valid one, since a rotation can succeed while enforcement quietly changes. Keep the list with your runbook so a different on-call engineer can follow it without reconstructing the setup.

Checking What You Actually Have Today

Trace a real request through your system and confirm, at each hop, whether it is encrypted, and whether the destination is authenticating the caller or just accepting whatever reaches it. Many teams discover during this exercise that an assumed mTLS boundary quietly degraded to plain TLS at some point during a refactor, without anyone deciding that on purpose.

Trace a real request with these steps:

  1. Pick one real request and follow it from the gateway to the model server, noting every hop it crosses.
  2. At each hop, record whether the traffic is encrypted.
  3. At each hop, record whether the destination authenticates the caller or accepts whatever reaches it.
  4. Compare what you found with the design you believe is in place, especially at the boundary into the model server.
  5. For any hop that quietly degraded to plain TLS, decide deliberately whether to restore it or accept it.

A Worked Example: The Sidecar That Quietly Stopped Enforcing Anything

Say a service mesh sidecar was configured to enforce mutual authentication when the model serving cluster launched, and that configuration lived in a template nobody revisited since. A later infrastructure upgrade replaces the sidecar with a newer version whose defaults changed, and the enforcement quietly stops applying because nobody explicitly re-specified it in the new configuration. Traffic keeps flowing normally, which is exactly the problem: nothing about the failure is visible until someone traces a request by hand and notices the authentication step never actually happens anymore.

Executive Capability Standard

What Good Looks Like

Every hop to the model server is encrypted, with mutual authentication at the boundary that matters most, and certificate rotation is automated with expiration alerts rather than tracked manually.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Trace a real request through your system and check at each hop whether it is encrypted and whether the destination authenticates the caller.
2. Do Manually:Set up mutual authentication by hand at your highest-risk boundary, such as the entry point into your model server, as a first step.
3. Delegate:Assign an engineer to own certificate lifecycle management, including rotation and expiration alerting.
4. Automate:Automate certificate rotation with alerts well ahead of expiration so it never relies on someone remembering a manual date.
5. Buy:Bring in fractional infrastructure advisory to design your encryption and authentication boundaries if the current setup grew ad hoc.

How to Get Started

Frequently Asked Questions

Do we need mutual TLS on every internal hop, or just at the model server?

Not necessarily every hop. A common middle ground encrypts everywhere but reserves full mutual authentication for the boundary into the model server itself, where an unauthenticated caller reaching it would matter most.

What's the most common real-world failure with mTLS setups?

Certificate rotation failing silently, either causing an outage or, worse, silently falling back to an unauthenticated connection. Automating rotation with an expiration alert well ahead of time prevents most of these failures.

Is it safe to terminate TLS at the gateway and run plain traffic internally?

It can be, if your internal network segment is genuinely isolated and tightly controlled. The risk is that this depends entirely on that isolation holding, which is worth verifying directly rather than assuming.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides