Engineering Leadership & Technical HiringPlaybook3 min readUpdated September 2026

Serverless Cold Starts: What's Actually Fixable and What Isn't

"Serverless scales automatically" is true and says nothing about the latency a real user experiences on the request that has to spin up a fresh execution environment first. Cold starts are a measured, specific cost, runtime initialization, dependency loading, sometimes a VPC network attachment, and the fix depends on which of those is actually driving your number, not a single blanket setting.

Here's how to decide which lever to pull first.

Measure Your Actual Cold Start Time Before Assuming a Fix

Most serverless platforms expose whether a given invocation was a cold or warm start as part of their logs or metrics; if you're not separating cold start latency from your overall P99, you're likely averaging a rare-but-severe cost into a number that looks fine most of the time and hides how bad the worst case actually is for the users who hit it. Start by isolating cold start frequency and duration specifically before deciding which mitigation to invest in.

Runtime and Language Choice Sets Your Floor

Interpreted or JIT-compiled runtimes generally have a heavier initialization cost than compiled, statically-linked ones; a function written in a compiled language with a small, self-contained binary often starts meaningfully faster cold than one written in a runtime that has to initialize a larger managed runtime environment first. This isn't a reason to rewrite an existing function purely for cold start performance, but it's worth weighing for a new, latency-sensitive function where cold start time is a hard requirement from day one.

Bundle Size and Dependency Loading Are Often the Bigger Lever

A function that imports a large dependency tree, even one it only uses a small part of, pays the cost of loading all of it during initialization. Tree-shaking unused code, lazy-loading dependencies only the specific invocation path actually needs, and trimming unused packages from the deployment bundle regularly cuts cold start time more than a runtime or language change would, and it's usually a smaller, more contained piece of work to ship. Check bundle size before reaching for a more disruptive fix.

Provisioned Concurrency Trades Cost for Guaranteed Warm Paths

Keeping a set number of execution environments warm and ready removes cold starts entirely for the traffic that fits within that provisioned capacity, at the cost of paying for that capacity whether or not it's actually being used at any given moment. This is worth it specifically for your latency-sensitive, customer-facing paths, a checkout flow, an API a paying customer's own integration depends on synchronously, and usually not worth it for background or batch processing paths where an occasional cold start doesn't meaningfully affect anyone's experience.

VPC Attachment Adds Its Own Cold Start Cost, Separate From Runtime Init

A function that needs network access into a VPC, to reach a private database or internal service, historically added a meaningful cold start penalty from establishing the network interface, on top of whatever the runtime itself costs to initialize. Newer networking models on most major serverless platforms have substantially reduced this specific cost compared to older implementations, but it's worth verifying current behavior for your specific platform and runtime rather than assuming it's still the dominant cost it used to be, since this is an area that changes with platform updates.

When a Cold Path Means Serverless Isn't the Right Fit

Some latency budgets are tight enough, a synchronous call inside another service's own request path with a strict timeout, that no combination of bundle trimming and provisioned concurrency reliably gets a serverless function under the ceiling every single time, since even provisioned concurrency doesn't eliminate cold starts entirely during a scale-up event that outpaces the provisioned capacity. For that narrow category of genuinely latency-critical, always-on paths, a long-running service that never cold-starts at all is sometimes the more honest answer than continuing to tune around a fundamentally different execution model.

Building a Cold Start Budget Into Your SLOs From the Start

Rather than treating cold starts as an unpredictable tail-latency nuisance to chase after the fact, decide up front what fraction of requests you're willing to accept as cold, and size provisioned concurrency and traffic patterns around that explicit budget. A service with a known, accepted five percent cold rate that's been deliberately chosen is in a much stronger position than one that's never measured its actual rate and gets surprised by a support ticket referencing a slow page load nobody can immediately explain.

Pull the levers in this order:

  1. Separate cold-start invocations from warm ones in your logs, so you know your real cold start time.
  2. Trim bundle size by tree-shaking unused code and lazy-loading dependencies a given path doesn't need.
  3. Consider a lighter runtime only after you've checked dependency loading.
  4. Use provisioned concurrency only for latency-sensitive, customer-facing functions, not background jobs.
  5. Check whether VPC attachment adds its own cold start cost on your platform.
  6. Set an explicit cold start budget in your SLOs.
Executive Capability Standard

What Good Looks Like

Your P99 latency on a cold execution path should be measured directly, not assumed away because serverless is supposed to scale.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Read your platform's own cold start metrics before assuming provisioned concurrency is automatically the right fix.
2. Do Manually:Manually trigger a cold invocation and time it end to end so you have a real baseline number.
3. Delegate:Give one engineer ownership of function bundle size so it doesn't quietly creep back up after each new dependency added.
4. Automate:Keep provisioned concurrency on your latency-sensitive paths and let everything else scale to zero as usual.
5. Buy:Consider a platform-native cold start optimization feature, where your provider offers one, before building a custom warm-up solution yourself.

How to Get Started

Frequently Asked Questions

Is provisioned concurrency worth it for every function, or just some?

Just the latency-sensitive, customer-facing ones. Background jobs, cron tasks, and asynchronous processing rarely need it, since an occasional few hundred milliseconds of extra cold start latency there doesn't affect a customer waiting on a synchronous response.

Does switching languages actually fix a serious cold start problem?

It can lower the floor, but check bundle size and dependency loading first. Those are usually a bigger and more contained lever than a full language rewrite, which is a much larger undertaking with risks unrelated to cold starts.

How do we know if VPC attachment is actually contributing to our cold start time?

Compare cold start duration for a version of the function with VPC access against one without, if that's feasible. Otherwise check your platform's current documentation and recent logs, since this cost has changed significantly across platform versions in recent years.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides