The Real Cost of Rolling Your Own Service-to-Service TLS
Encrypting traffic between services sounds like a one-time setup: generate certificates, configure both sides, done. The setup is the easy part. What's expensive is everything after it, rotating certificates before they expire, revoking one when a service is compromised, and doing both without an outage, and that ongoing cost is where most hand-rolled mutual TLS setups eventually break down.
The question worth asking isn't whether to encrypt service-to-service traffic. It's who maintains the machinery that keeps it working.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
What manual certificate management actually requires
A workable in house setup needs a certificate authority you control, an issuance process for every service that needs one, and a rotation schedule that runs well before expiry, not the week of. Each of those is a small piece of ongoing infrastructure with its own failure modes: a forgotten rotation, a certificate issued with the wrong hostname, a revocation that doesn't propagate to every service checking it. None of these are hard individually. Maintaining all of them correctly, indefinitely, without dedicated ownership, is where teams actually struggle. Write down, explicitly, who owns each of these steps and what happens if that person is unavailable when a rotation is due. A process that only works because one specific engineer remembers to run it isn't a process, it's a single point of failure wearing a runbook's clothing, and it fails the same way any single point of failure does: quietly, until the person is out and the date arrives anyway.
The failure mode that's invisible until an expiry hits
A certificate quietly approaching expiry produces no symptoms at all, right up until the moment it expires and every connection depending on it starts failing simultaneously, often across multiple services at once if they were issued in a batch. This is a uniquely bad failure mode: it's completely silent beforehand and then total and immediate. Alert on certificate expiry well ahead of the deadline, thirty and seven days out at minimum, not just at the moment of failure.
Where automated rotation earns its cost
A service mesh or managed certificate platform that handles issuance and rotation automatically removes the specific failure mode above entirely: certificates rotate on a schedule before they're ever close to expiry, without a human remembering to trigger it. The tradeoff is a new dependency, and a new thing to understand when something goes wrong with certificate issuance itself. For teams running more than a handful of services, that tradeoff usually favors automation, since the manual alternative scales linearly with headcount.
Tie your rotation window to what an incident would actually cost
How much buffer to build into a rotation schedule depends on how expensive an outage from a missed rotation would be. A 99.95% uptime target allows roughly 4.38 hours of downtime a year across every incident combined1, which is a useful anchor: a rotation process with any chance of a multi-hour gap between expiry and renewal is putting a meaningful share of that annual budget at risk over a single preventable event.
A worked example: what one missed rotation actually costs
Say a batch of internal service certificates was issued together and nobody flagged the shared expiry date. When it hits, every one of those services loses the ability to authenticate to each other at once, not gradually. If that takes two hours to fully diagnose and fix across a dozen affected services, that single incident alone can consume a large share of a tight annual downtime budget, all from one date that a thirty day expiry alert would have caught with time to spare.
A workable split for most teams
Track certificate inventory and expiry dates somewhere visible, a tool like ClickUp works fine for the calendar and ownership, regardless of whether issuance itself is manual or automated. Automate rotation once you're running more than a handful of services, since the manual alternative doesn't scale with headcount. Compliance platforms such as Vanta can often collect encryption-in-transit evidence from your setup when a SOC 2 auditor asks how service-to-service traffic is protected; confirm what it supports for your stack in the vendor's current documentation.
Put certificate management into practice this way:
- Track certificate inventory and expiry dates somewhere visible, with named owners, whether issuance is manual or automated.
- Automate rotation once you run more than a handful of services, since manual rotation doesn't scale with headcount.
- Rotate well before expiry, not the week of, so a missed step never becomes a shared outage.
- Avoid issuing a batch of certificates that all share one expiry date that nobody has flagged.
What Good Looks Like
Good service-to-service encryption means certificates tracked with clear ownership, expiry alerts well ahead of the deadline, and a rotation process, manual or automated, sized to how many services actually depend on it.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
Is mutual TLS overkill for a small number of internal services?
Not necessarily overkill, but the operational cost needs to be proportional. A handful of services can reasonably run manual certificate management with a solid expiry alert and a documented rotation process. The math changes once the service count grows past what one person can track reliably by hand.
What's the biggest risk with hand-rolled certificate management?
A silent, ongoing maintenance burden that produces no symptoms until an expiry is missed, at which point every dependent connection fails at once. The risk isn't the initial setup, it's the accumulating chance of a missed rotation the longer the manual process runs unattended.
How far ahead should expiry alerts fire?
Thirty and seven days out at minimum, with a clear owner for each alert, not just a notification nobody's assigned to act on. A single alert the day of expiry is functionally the same as no alert, since there's no time left to fix anything before the failure hits.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
When You Actually Need Mutual TLS Between Services
A practical way to decide whether mutual TLS between your internal services is worth the operational cost, or whether standard TLS is enough.
Mutual TLS in a Service Mesh: The Questions Engineers Actually Ask
Plain answers to the questions engineering teams actually run into when rolling out mutual TLS in a service mesh, from cert rotation to debugging failures.
Rolling Out Mutual TLS Without Breaking Every Service
A staged approach to adding mutual TLS between services that catches certificate and trust issues before they take down production traffic.
When Mutual TLS Is Worth the Operational Cost
How mutual TLS differs from standard TLS, where it genuinely earns its operational cost inside a service mesh, and where a simpler auth approach is enough.
The mTLS Rollout Checklist That Prevents a 2 AM Outage
Mutual TLS fails loud, not quiet, when a certificate expires. Here is a pre-launch checklist that catches the mistakes that cause an outage later.
When Your RAG Pipeline Actually Needs mTLS, Not Just TLS
A decision guide for where TLS is enough and where a production RAG pipeline's service-to-service traffic actually needs mutual TLS instead.