Cutting Serverless Cold Start Time Without Just Throwing Money at It
A cold start happens when your serverless platform has to initialize a fresh execution environment before it can run your function, and for a latency-sensitive endpoint, that initialization time is the difference between a snappy response and a user noticing the delay. There are several real ways to reduce it, but they carry different costs, and reaching for the most expensive fix first often isn't necessary.
Here's what actually moves the number, roughly in order of cost.
Free: Trim What Runs During Initialization
Before spending anything, check what your function actually does before it can handle its first request. A function that imports a large dependency tree, initializes multiple SDK clients it doesn't need for every code path, or loads configuration from a slow external call during initialization is paying for all of that on every cold start. Lazy-load anything not needed for the specific request path being handled, and trim unused dependencies from your package. This is usually the most effective fix per hour spent, because it reduces the cold start cost itself rather than working around it.
Whatever you change, measure first. Record cold start duration separately from warm invocation time, because averaging the two hides the problem and makes every fix look less effective than it is. Change one thing at a time so you know which fix moved the number. For example, if trimming initialization brings a checkout function to a delay users no longer notice, stop there and skip the paid options. If a function is rarely called and nobody is waiting on it, a slow cold start may simply be acceptable, and the honest answer is to leave it alone.
Low Cost: Choose a Runtime With a Faster Startup Profile
Different language runtimes have meaningfully different cold start characteristics. Compiled languages with smaller runtimes generally initialize faster than runtimes that need to start a larger interpreter or virtual machine and load more of the standard library before your code even runs. If you're building a new latency-sensitive function and have flexibility in language choice, this is worth weighing alongside your team's existing expertise, though switching an entire codebase's language purely for cold start improvement is rarely worth it on its own.
Moderate Cost: Reduce Package and Dependency Size
A smaller deployment package generally loads faster, since more of it fits in memory faster and there's simply less code to parse and initialize. Bundling and tree-shaking your code, removing unused dependencies, and splitting a large function into more focused, smaller functions with narrower dependency trees can meaningfully shrink cold start time. This takes real engineering effort but has no ongoing dollar cost, which makes it worth doing before reaching for a paid mitigation.
When is provisioned concurrency worth the cost?
Provisioned concurrency keeps a set number of execution environments warm and ready at all times, eliminating cold starts entirely for traffic within that provisioned capacity, at the cost of paying for that capacity whether or not it's actively serving requests. This is the right tool specifically for endpoints where cold start latency is unacceptable and traffic is predictable enough to provision accurately, a checkout flow, an authentication endpoint. It's the wrong tool for a rarely called internal batch function, where you'd be paying continuously for warmth that a rare cold start wouldn't meaningfully hurt.
Which cold start fix fits your case?
Start with the free fix, trimming initialization work, on every function regardless of how urgent the problem feels, since it costs nothing and often recovers a meaningful chunk of the delay on its own. Move to package size reduction next if cold starts are still a measurable problem after that. Reserve provisioned concurrency for the specific handful of endpoints where latency genuinely matters to the user and traffic is predictable enough to provision without either wasting capacity or under-provisioning during a spike.
Work through the fixes in order of cost:
- Trim initialization work on every function, lazy-loading anything the specific request path does not need.
- Weigh a runtime with a faster startup profile for new latency-sensitive functions, without rewriting a codebase for this reason alone.
- Reduce package and dependency size through bundling, tree-shaking, and splitting large functions into focused ones.
- Reserve provisioned concurrency for predictable, user-facing endpoints such as checkout or authentication, sized against real peak traffic.
A Worked Example of Sizing Provisioned Concurrency
Say your checkout function typically handles ten concurrent requests during normal traffic and spikes to thirty during a promotion. Provisioning ten warm environments handles normal traffic with no cold starts, but the twenty additional requests during a spike still hit cold starts unless you provision closer to your actual peak. Provisioning for thirty removes that gap but means paying for twenty idle environments during every normal hour of the day. A middle path, provisioning fifteen to twenty and accepting occasional cold starts only during the rarest spikes, is often the more honest tradeoff than provisioning for a worst case that happens a few times a year.
What Good Looks Like
Good cold start management means initialization code is trimmed and audited before any paid mitigation is applied, and provisioned concurrency is reserved for the specific endpoints where latency is user-visible and traffic is predictable.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Does provisioned concurrency completely eliminate cold starts?
It eliminates them for traffic within the provisioned capacity. A traffic spike that exceeds what you've provisioned will still trigger cold starts for the excess requests, so provisioned concurrency needs to be sized against your real peak traffic, not just your average, to fully solve the problem for that endpoint.
How much does trimming initialization code actually help compared to provisioned concurrency?
It varies by function, but it's often the single most effective change available, since it reduces the actual cold start duration rather than paying to avoid experiencing it. Many functions carry unnecessary work during initialization that got added incrementally and never got audited, and removing it directly reduces every cold start's length.
Should we just provision concurrency for every function to be safe?
No, that turns serverless's pay-for-what-you-use benefit into something closer to running dedicated instances continuously, and for a lot of functions, the cold start cost genuinely doesn't matter enough to justify that. Reserve it for endpoints where latency is user-visible and traffic is predictable enough to provision accurately.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Cutting Serverless Cold Starts Without Overpaying
A function that answers in 80 milliseconds warm takes four seconds cold. What actually drives cold start time, and when provisioned concurrency is worth it.
The Serverless Cold Start Fixes That Actually Move the Number
A runbook for reducing serverless cold start latency: runtime choice, package size, provisioned concurrency, and the fixes that don't actually help.
A Runbook for Cutting Serverless Cold Start Latency
A step-by-step runbook for reducing serverless cold start latency, from trimming your deployment package to deciding when to pay for provisioned capacity.
Serverless Cold Starts: What's Actually Fixable and What Isn't
Provisioned concurrency, bundle size, and runtime choice: a decision guide for which cold start fixes actually move your P99 latency.
Cutting Serverless Cold Starts Without Abandoning Serverless
Why serverless cold starts happen, which patterns make them worse, and the mitigation options that don't quietly turn serverless into servers.
Cutting Serverless Cold Starts Without Giving Up on Serverless
Where cold start time actually goes, when provisioned concurrency is worth paying for, and which functions don't need the fix at all.