Where to Terminate TLS in Your Model Serving Path
Terminate TLS at the gateway if your internal network is genuinely isolated, and carry mutual TLS to the model server when it handles sensitive input or serves multiple tenants. Every hop toward the model server is a place traffic could be intercepted if it is unencrypted, so the decision turns on where your trusted network boundary sits.
The decision is really about where you draw the boundary of your trusted network, and how confident you actually are in that boundary holding over time, not just on the day it was first configured.
Should You Terminate TLS at the Gateway?
Ending encryption at your API gateway and running plain, unencrypted traffic between the gateway and internal model servers is the simplest setup, and it is a reasonable choice if that internal network segment is genuinely isolated and access to it is tightly controlled. The risk is that this simplicity depends entirely on that internal isolation holding, and a misconfigured security group or an unexpected peering connection can quietly turn what you assumed was a private segment into something reachable from further away than intended.
Carrying mTLS All the Way to the Model Server
Mutual TLS all the way to the model server means every hop authenticates both directions, so even a service that finds itself on the same network segment cannot simply connect without a valid certificate. This closes the gap that a gateway-only termination leaves open if internal isolation ever fails, at the cost of managing and rotating certificates for every service in the path. For a model server handling sensitive input or serving multiple tenants, that cost is usually worth it.
A Middle Path: Encrypt Everywhere, Authenticate Where It Matters Most
A reasonable middle ground for many teams is encrypting every hop but reserving full mutual authentication for the boundary between your network and the model server itself, where the consequences of a bypass are highest. This avoids the full certificate management burden of mTLS everywhere while still closing the specific gap, an unauthenticated internal caller reaching the model server, that matters most. It also gives you a natural place to start if you are retrofitting encryption onto an existing setup: get the highest-risk boundary right first, then decide deliberately whether the rest of the path needs the same treatment.
Why Does Certificate Rotation Usually Break?
The design decision above matters less than whether certificate rotation actually happens reliably. An expired certificate silently failing over to plain connections, or causing an outage because nothing rotated it in time, is a more common real-world failure than the encryption boundary being wrong in the first place. Automate rotation and alert on an approaching expiration well before it happens, rather than relying on someone remembering a manual renewal date. Test the alert itself periodically too, by checking that it would actually fire with enough lead time to act, not only that it exists somewhere in a configuration file.
For example, list every certificate in the request path with its expiry date and the mechanism that renews it. Any certificate with no automated renewal, or whose expiry alert has never been tested, is your first fix. When a certificate is replaced, confirm the new one is actually presented and that the destination still rejects callers without a valid one, since a rotation can succeed while enforcement quietly changes. Keep the list with your runbook so a different on-call engineer can follow it without reconstructing the setup.
Checking What You Actually Have Today
Trace a real request through your system and confirm, at each hop, whether it is encrypted, and whether the destination is authenticating the caller or just accepting whatever reaches it. Many teams discover during this exercise that an assumed mTLS boundary quietly degraded to plain TLS at some point during a refactor, without anyone deciding that on purpose.
Trace a real request with these steps:
- Pick one real request and follow it from the gateway to the model server, noting every hop it crosses.
- At each hop, record whether the traffic is encrypted.
- At each hop, record whether the destination authenticates the caller or accepts whatever reaches it.
- Compare what you found with the design you believe is in place, especially at the boundary into the model server.
- For any hop that quietly degraded to plain TLS, decide deliberately whether to restore it or accept it.
A Worked Example: The Sidecar That Quietly Stopped Enforcing Anything
Say a service mesh sidecar was configured to enforce mutual authentication when the model serving cluster launched, and that configuration lived in a template nobody revisited since. A later infrastructure upgrade replaces the sidecar with a newer version whose defaults changed, and the enforcement quietly stops applying because nobody explicitly re-specified it in the new configuration. Traffic keeps flowing normally, which is exactly the problem: nothing about the failure is visible until someone traces a request by hand and notices the authentication step never actually happens anymore.
What Good Looks Like
Every hop to the model server is encrypted, with mutual authentication at the boundary that matters most, and certificate rotation is automated with expiration alerts rather than tracked manually.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Do we need mutual TLS on every internal hop, or just at the model server?
Not necessarily every hop. A common middle ground encrypts everywhere but reserves full mutual authentication for the boundary into the model server itself, where an unauthenticated caller reaching it would matter most.
What's the most common real-world failure with mTLS setups?
Certificate rotation failing silently, either causing an outage or, worse, silently falling back to an unauthenticated connection. Automating rotation with an expiration alert well ahead of time prevents most of these failures.
Is it safe to terminate TLS at the gateway and run plain traffic internally?
It can be, if your internal network segment is genuinely isolated and tightly controlled. The risk is that this depends entirely on that isolation holding, which is worth verifying directly rather than assuming.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
When You Actually Need Mutual TLS Between Services
A practical way to decide whether mutual TLS between your internal services is worth the operational cost, or whether standard TLS is enough.
A Checklist for mTLS Setups That Look Right and Aren't
A checklist for mutual TLS in a service mesh, and the specific pitfalls, expired certs, weak fallbacks, and skipped validation, that let a setup look secure.
Mutual TLS in a Service Mesh: The Questions Engineers Actually Ask
Plain answers to the questions engineering teams actually run into when rolling out mutual TLS in a service mesh, from cert rotation to debugging failures.
The Real Cost of Rolling Your Own Service-to-Service TLS
What hand-rolled certificate management for service-to-service encryption actually requires to maintain, and where an automated approach earns its cost.
The mTLS Rollout Checklist That Prevents a 2 AM Outage
Mutual TLS fails loud, not quiet, when a certificate expires. Here is a pre-launch checklist that catches the mistakes that cause an outage later.
When Your RAG Pipeline Actually Needs mTLS, Not Just TLS
A decision guide for where TLS is enough and where a production RAG pipeline's service-to-service traffic actually needs mutual TLS instead.