Cloud FinOps & Infrastructure ScalingPlaybook3 min readUpdated September 2026

The Real Cost of Rolling Your Own Service-to-Service TLS

Encrypting traffic between services sounds like a one-time setup: generate certificates, configure both sides, done. The setup is the easy part. What's expensive is everything after it, rotating certificates before they expire, revoking one when a service is compromised, and doing both without an outage, and that ongoing cost is where most hand-rolled mutual TLS setups eventually break down.

The question worth asking isn't whether to encrypt service-to-service traffic. It's who maintains the machinery that keeps it working.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

What manual certificate management actually requires

A workable in house setup needs a certificate authority you control, an issuance process for every service that needs one, and a rotation schedule that runs well before expiry, not the week of. Each of those is a small piece of ongoing infrastructure with its own failure modes: a forgotten rotation, a certificate issued with the wrong hostname, a revocation that doesn't propagate to every service checking it. None of these are hard individually. Maintaining all of them correctly, indefinitely, without dedicated ownership, is where teams actually struggle. Write down, explicitly, who owns each of these steps and what happens if that person is unavailable when a rotation is due. A process that only works because one specific engineer remembers to run it isn't a process, it's a single point of failure wearing a runbook's clothing, and it fails the same way any single point of failure does: quietly, until the person is out and the date arrives anyway.

The failure mode that's invisible until an expiry hits

A certificate quietly approaching expiry produces no symptoms at all, right up until the moment it expires and every connection depending on it starts failing simultaneously, often across multiple services at once if they were issued in a batch. This is a uniquely bad failure mode: it's completely silent beforehand and then total and immediate. Alert on certificate expiry well ahead of the deadline, thirty and seven days out at minimum, not just at the moment of failure.

Where automated rotation earns its cost

A service mesh or managed certificate platform that handles issuance and rotation automatically removes the specific failure mode above entirely: certificates rotate on a schedule before they're ever close to expiry, without a human remembering to trigger it. The tradeoff is a new dependency, and a new thing to understand when something goes wrong with certificate issuance itself. For teams running more than a handful of services, that tradeoff usually favors automation, since the manual alternative scales linearly with headcount.

Tie your rotation window to what an incident would actually cost

How much buffer to build into a rotation schedule depends on how expensive an outage from a missed rotation would be. A 99.95% uptime target allows roughly 4.38 hours of downtime a year across every incident combined1, which is a useful anchor: a rotation process with any chance of a multi-hour gap between expiry and renewal is putting a meaningful share of that annual budget at risk over a single preventable event.

A worked example: what one missed rotation actually costs

Say a batch of internal service certificates was issued together and nobody flagged the shared expiry date. When it hits, every one of those services loses the ability to authenticate to each other at once, not gradually. If that takes two hours to fully diagnose and fix across a dozen affected services, that single incident alone can consume a large share of a tight annual downtime budget, all from one date that a thirty day expiry alert would have caught with time to spare.

A workable split for most teams

Track certificate inventory and expiry dates somewhere visible, a tool like ClickUp works fine for the calendar and ownership, regardless of whether issuance itself is manual or automated. Automate rotation once you're running more than a handful of services, since the manual alternative doesn't scale with headcount. Compliance platforms such as Vanta can often collect encryption-in-transit evidence from your setup when a SOC 2 auditor asks how service-to-service traffic is protected; confirm what it supports for your stack in the vendor's current documentation.

Put certificate management into practice this way:

  • Track certificate inventory and expiry dates somewhere visible, with named owners, whether issuance is manual or automated.
  • Automate rotation once you run more than a handful of services, since manual rotation doesn't scale with headcount.
  • Rotate well before expiry, not the week of, so a missed step never becomes a shared outage.
  • Avoid issuing a batch of certificates that all share one expiry date that nobody has flagged.
Executive Capability Standard

What Good Looks Like

Good service-to-service encryption means certificates tracked with clear ownership, expiry alerts well ahead of the deadline, and a rotation process, manual or automated, sized to how many services actually depend on it.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Pull a full inventory of your current service-to-service certificates and check how far each one is from expiry right now.
2. Do Manually:Manually set up expiry alerts for every certificate in that inventory if they don't already exist.
3. Delegate:Assign a specific owner for certificate rotation and revocation so it doesn't depend on institutional memory.
4. Automate:Move to automated issuance and rotation once manual tracking is covering more services than one person can reliably watch.
5. Buy:Bring in a managed platform for this once the engineering time spent maintaining it exceeds what the platform would cost.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Frequently Asked Questions

Is mutual TLS overkill for a small number of internal services?

Not necessarily overkill, but the operational cost needs to be proportional. A handful of services can reasonably run manual certificate management with a solid expiry alert and a documented rotation process. The math changes once the service count grows past what one person can track reliably by hand.

What's the biggest risk with hand-rolled certificate management?

A silent, ongoing maintenance burden that produces no symptoms until an expiry is missed, at which point every dependent connection fails at once. The risk isn't the initial setup, it's the accumulating chance of a missed rotation the longer the manual process runs unattended.

How far ahead should expiry alerts fire?

Thirty and seven days out at minimum, with a clear owner for each alert, not just a notification nobody's assigned to act on. A single alert the day of expiry is functionally the same as no alert, since there's no time left to fix anything before the failure hits.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides