Distributed Systems & Enterprise ResiliencePlaybook3 min readUpdated September 2026

Cutting Serverless Cold Starts Without Giving Up on Serverless

Cold starts are the price serverless functions pay for not keeping anything running between requests. The fix isn't always more infrastructure. Sometimes it's trimming what the function does before it can serve a request, and sometimes the honest answer is that the function in question doesn't need fixing at all.

Where the cold start time actually goes

A cold start covers runtime initialization, loading your dependencies, and, for functions attached to a VPC, the extra time to attach a network interface before the function can reach anything inside that VPC. Language runtimes differ meaningfully here: a compiled binary generally initializes faster than a runtime that has to parse and load a large dependency tree on every cold invocation.

Measure each piece separately before assuming which one dominates for your function. Teams often optimize dependency loading when the VPC attachment was actually the larger cost, or the reverse.

Your platform's own invocation logs usually break the total duration into an init phase and an execution phase, which is enough to start with. If you need more detail than that, adding a timestamp at the very top of your handler, before any imports run, and comparing it to the platform's reported start time isolates dependency loading specifically.

Provisioned concurrency versus keeping functions warm yourself

Provisioned concurrency keeps a set number of instances initialized and ready, and you pay for that idle capacity whether or not a request arrives to use it. A scheduled ping to keep a function warm is the cheaper-looking alternative, but it's fragile: it doesn't guarantee the exact instance your next real request will land on, especially once you're running more than one concurrent instance.

That gap between how it behaves in testing and how it behaves under real concurrent traffic is exactly why it can look like it's working right up until the moment it doesn't.

Trimming the function before reaching for infrastructure

A smaller deployment package initializes faster on every cold start, not just the first one. Lazy-loading a heavy SDK only when the code path that needs it actually runs, instead of importing it at the top of the file unconditionally, moves that cost off the common path.

Avoiding a VPC attachment unless the function genuinely needs to reach something inside one removes an entire category of cold start time outright, rather than shaving milliseconds off it.

None of this requires touching infrastructure at all, which is why it's worth doing first: a smaller package and a lazy-loaded dependency ship in the same pull request as any other code change, while provisioned concurrency is a standing cost that shows up on next month's bill whether or not it was the right fix.

Work through these trims before paying for warm capacity:

  • Shrink the deployment package by removing unused dependencies, since a smaller package initializes faster on every cold start.
  • Lazy-load heavy SDKs inside the code path that needs them, instead of importing them at the top of the file.
  • Skip the VPC attachment unless the function must reach something inside one, which removes the network interface setup time.
  • Compare the init and execution phases in your platform's invocation logs to see which one dominates before changing anything.

When cold starts genuinely don't matter

A background job, a batch process, or a low-traffic internal tool that nobody is watching a spinner for doesn't need the same treatment as a customer-facing API endpoint. Spending engineering time and ongoing spend to shave cold start time off a function nobody notices the latency of is effort better spent elsewhere.

List your functions by who's actually waiting on the response before deciding where cold start work pays off. Not every function is on the critical path for a person watching a screen.

The mistake: over-provisioning for a spike that happens once a day

Provisioned concurrency sized for peak traffic makes sense when that peak is sustained. It's expensive and often unnecessary when the real traffic pattern is a single daily spike, such as a nightly batch job's callback, with near-zero traffic the rest of the day.

Look at your actual traffic shape over a full day before buying always-warm capacity sized for the busiest hour. A function that's idle twenty-three hours a day rarely justifies paying to keep it warm all twenty-four.

Deploys reset the warm pool, so time them deliberately

A new deploy typically replaces the warm instances your provisioned concurrency built up, which means the fleet a deploy leaves behind can be entirely cold right when the deploy finishes and real traffic starts arriving at it again.

Shift traffic to a new version gradually rather than all at once immediately after a deploy, and let provisioned concurrency ramp on the new version before it's taking full load. Skipping this step turns every deploy into a self-inflicted cold start event for whoever hits the app right after it ships.

Executive Capability Standard

What Good Looks Like

A mature approach to cold starts matches the fix to the function: trimmed packages and lazy loading everywhere, provisioned concurrency only where sustained latency actually matters.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Measure actual cold start time for your customer-facing functions and identify which ones are attached to a VPC unnecessarily.
2. Do Manually:Manually trim a function's deployment package and move a heavy SDK import behind a lazy load to see the cold start improvement before automating it elsewhere.
3. Delegate:Give one engineer ownership of a checklist for new functions covering package size, VPC attachment, and dependency loading.
4. Automate:Automate provisioned concurrency scaling tied to your actual traffic pattern rather than a fixed always-on number.
5. Buy:Bring in serverless-focused engineering help if cold starts are costing customer-facing latency across many functions and nobody owns the pattern yet.

How to Get Started

Frequently Asked Questions

Does provisioned concurrency eliminate cold starts entirely?

It eliminates them for the instances it keeps warm, up to the concurrency level you've paid for. Traffic that exceeds that level still triggers a cold start for the overflow instances, so it reduces the problem rather than removing it completely unless you provision for your actual peak.

Which language runtimes cold start fastest?

Compiled languages that don't need to parse and load a large runtime or dependency tree on startup generally cold start faster than interpreted languages with heavier initialization. The gap narrows once you've trimmed a heavier runtime's deployment package and lazy-loaded its dependencies, so it's not purely a language choice.

Is a scheduled warmer ping a real fix?

It's a partial one at best. It keeps one instance warm but doesn't guarantee your next real request lands on that specific instance once you're running any meaningful concurrency, so it can look effective in light testing and still fail to help under real traffic.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides