Model Context Protocol & Agentic ArchitecturePlaybook3 min readUpdated September 2026

A Checklist for mTLS Setups That Look Right and Aren't

Mutual TLS gets adopted because it sounds like a strong, complete guarantee: both sides prove who they are before any data moves. In practice, plenty of mTLS setups fail quietly at one specific point in the chain while looking correct everywhere else, which is worse than not having it at all, because it creates false confidence.

Each check below targets one specific failure mode, and each one is worth testing directly rather than inferring from the configuration alone, since a setting that looks correct in a config file doesn't always behave the way the file implies once real traffic hits it.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Do both sides of an mTLS connection validate the certificate?

mTLS means both the client and the server present and validate certificates, but it's common to find one side configured to accept a connection without properly validating the peer's certificate chain, often left over from an early debugging setup that was never tightened. Confirm, for each service pair, that validation failure on either side actually rejects the connection rather than logging a warning and proceeding anyway.

Does your mesh fall back to plain TLS or plaintext?

Some service mesh configurations quietly fall back to standard TLS, or in the worst case plaintext, if mTLS negotiation fails, rather than refusing the connection outright. That fallback defeats the purpose of requiring mTLS in the first place, since it means an attacker who can disrupt the mTLS handshake gets a working, less-protected connection instead of a rejected one. Confirm your mesh is configured to fail closed, not fail open, on a handshake failure.

Check certificate expiration and rotation, not just issuance

A certificate that's valid today but expires in three weeks, with no automated rotation in place, is an outage waiting for a specific date. Confirm rotation is automated and tested, not just configured, since a rotation job that's been silently failing for months looks identical to a working one right up until the old certificate actually expires and every connection depending on it starts failing at once.

Check that internal service-to-service traffic isn't quietly exempted

It's common for mTLS to be enforced for traffic crossing a clear external boundary while internal, same-cluster traffic between services is left on a simpler, unauthenticated path, on the reasoning that internal traffic is already inside a trusted network. That reasoning is exactly what mTLS in a service mesh is meant to remove: internal network trust alone doesn't protect against a compromised container on the same cluster impersonating another service.

Check that certificate revocation actually works

A compromised certificate needs to be revocable, with the revocation actually enforced quickly by every service checking it, not just issued into a revocation list nobody consults. Test this specifically: revoke a test certificate and confirm connections using it are rejected within the time window you expect, rather than assuming the mechanism works because it's configured.

Check that new services are onboarded into the mesh by default

A new service spun up outside the standard deployment template can end up excluded from mTLS enforcement entirely, simply because whoever built it wasn't aware of the requirement or used an older template that predates it. Make mTLS participation the default for any new service in the mesh, enforced by your platform tooling rather than left to individual teams to remember, so coverage doesn't quietly erode every time a new service is added.

Document the failure modes you've actually tested, not just the design

Keep a short, current record of which of these checks have been tested against your live setup and when, separate from the architecture diagram describing how it's supposed to work. The diagram tells you the intent. The test record tells you what's actually been verified, and the two drift apart more often than teams expect, particularly after a mesh upgrade or a change to the certificate authority.

Retest after every mesh or certificate authority change

Each of the checks above can pass today and silently regress after an upgrade to your service mesh software, a change in certificate authority, or a new default introduced by a platform update. Treat a mesh upgrade the same way you'd treat a change to a production database's permissions: something that specifically triggers a re-run of your validation checks, rather than an update you apply and assume left everything else exactly as it was.

Re-run these checks after any mesh upgrade or certificate authority change:

  • Both sides reject a connection when peer certificate validation fails, instead of logging a warning and continuing.
  • The mesh refuses the connection when mTLS negotiation fails, with no fallback to plain TLS or plaintext.
  • Certificate rotation is automated and tested, not just configured.
  • Internal service-to-service traffic is not quietly exempted from enforcement.
  • A revoked test certificate is rejected within the time window you expect.
  • A newly created service joins the mesh with mTLS enforced by default.
Executive Capability Standard

What Good Looks Like

A working mTLS setup validates certificates on both sides with no silent fallback to a weaker connection, automates and tests certificate rotation before expiration, covers internal service-to-service traffic and not just external boundaries, and has a tested, working revocation path.

Building The Capability (5-Stage Skill Ladder)

1. Learn:read your current service mesh's mTLS configuration and confirm whether it fails open or closed on a handshake failure
2. Do Manually:manually test certificate validation and revocation on one service pair to confirm the behavior matches what the configuration claims
3. Delegate:assign ownership of certificate rotation and mesh-wide mTLS configuration to a specific team, separate from individual service owners
4. Automate:automate certificate issuance and rotation with alerting on any certificate approaching expiration without a completed rotation
5. Buy:compliance platforms like Vanta can track encryption-in-transit as an evidenced control, which is useful once an auditor wants proof the mTLS setup is both configured and actually enforced

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Vanta

Vanta can track encryption-in-transit as an evidenced control, which is worth pairing with the checks above since a control that's configured but not verified working isn't the same as one that's actually enforced.

Visit Vanta→

Frequently Asked Questions

Is mTLS necessary if all our traffic is already inside a private network?

Network-level isolation and mTLS answer different questions: isolation controls what can route to what, mTLS proves who's actually on the other end of an allowed connection. A private network alone doesn't stop a compromised workload inside it from impersonating another service.

How do we catch a fail-open misconfiguration before it causes an incident?

Test it deliberately: present an invalid or expired certificate to a service expecting mTLS and confirm the connection is rejected, not just logged. A fail-open configuration typically looks identical to a correct one until something actually tries to exploit the gap.

How often should mTLS certificates rotate?

Shorter-lived certificates with frequent, automated rotation reduce the exposure window if one is ever compromised, and remove the risk of a manual process being forgotten. Confirm your mesh's default rotation cadence rather than assuming it matches what you'd choose deliberately.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides