Cutting Serverless Cold Starts Without Abandoning Serverless
A cold start is the delay a serverless function pays the first time it runs after being idle, while the platform provisions a fresh execution environment, loads your code, and initializes your runtime before it can process the actual request.
The mitigation options range from nearly free to expensive, and picking the expensive one before trying the free ones is a common way to erase the cost benefit serverless was supposed to bring in the first place.
What actually makes a cold start slow
Runtime choice matters more than most teams expect: a compiled or lighter-weight runtime typically initializes faster than one that has to start a full virtual machine or interpreter with a large standard library. Package size matters too, since a function bundling a huge dependency tree has more code to load before it can run, and a lot of that code may not even be used on the request path.
Anything the function does in its initialization code before handling the first request, like establishing a database connection pool sized for steady-state load, adds directly to cold start time.
The free fixes: trim before you pay
Reduce package size by removing unused dependencies and using a bundler that tree-shakes anything not actually imported. Move expensive initialization, like large SDK client setup, to lazy-load only when the specific code path that needs it actually runs, instead of unconditionally on every cold start regardless of which handler gets invoked.
For a runtime with multiple flavors, check whether a lighter one gets you most of the way there before reaching for anything that costs money.
Provisioned concurrency: fast but not free
Provisioned concurrency keeps a set number of execution environments warm and ready, eliminating cold starts entirely for the traffic that fits within it, at the cost of paying for that capacity whether it's actively serving requests or not.
It's a reasonable fix for a predictable, latency-sensitive path, like a checkout API where a slow first response directly costs revenue. It's a poor fit for a rarely-called, latency-tolerant function, like an overnight batch job, where you'd be paying continuously for warm capacity a cold start would barely be noticed on.
A decision checklist before reaching for provisioned concurrency
- Have you already trimmed package size and lazy-loaded expensive initialization? Provisioned concurrency on an unoptimized function is paying to mask a problem you could have fixed for free.
- Is this specific function on a latency-sensitive path where a user or another system is waiting synchronously, or is it something that runs in the background where a cold start's delay genuinely doesn't matter?
- Is the traffic pattern predictable enough that provisioned concurrency's fixed capacity matches real demand, or would you be paying for warm capacity that sits idle most of the time?
When cold starts are a sign serverless isn't the right fit anymore
If a workload has grown into a steady, predictable, always-on traffic pattern, the case for serverless in the first place, paying only for what you use, starts to weaken regardless of how well you've mitigated cold starts. At that point, a small always-on container or a traditional server for that specific workload can end up both faster and cheaper than a heavily provisioned serverless function fighting its own architecture.
A worked example: a function that got slower after a dependency update
Say a routine dependency update quietly pulls in a much larger transitive dependency, and cold start times creep up noticeably over the following weeks without anyone connecting the two events. Nobody changed the function's own code, so the instinct is to look everywhere except the dependency tree.
Comparing package size before and after the update, not just eyeballing the diff for logic changes, is what actually surfaces this kind of regression. Treat package size as a metric worth tracking over time on latency-sensitive functions, the same way you'd track response latency itself, rather than something you only check when a cold start problem has already become visible to users.
Warmup pings: a cheap partial fix with real limits
Pinging a function on a schedule to keep an instance warm is a common, low-effort mitigation, and it genuinely helps for low-traffic functions with predictable idle gaps. Its limits show up fast under real concurrency: a warmup ping keeps one instance warm, but a burst of simultaneous requests still needs additional cold instances to scale beyond whatever the ping is holding ready. Treat warmup pings as a stopgap for a specific known gap in traffic, not a substitute for provisioned concurrency once a path genuinely needs guaranteed low latency under load.
What Good Looks Like
A good cold start strategy trims package size and lazy-loads expensive initialization first, reserves provisioned concurrency for genuinely latency-sensitive paths, and reconsiders serverless entirely once a workload's traffic pattern has become steady and predictable.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
What's the fastest free fix for a serverless cold start?
Trim your package size by removing unused dependencies and lazy-loading expensive initialization, like large SDK client setup, so it only runs when the specific code path that needs it actually executes. This alone often cuts cold start time meaningfully before you need to pay for anything.
Is provisioned concurrency worth it for every function?
No. It's worth the added cost for a predictable, latency-sensitive path where a user is waiting synchronously, like a checkout API. For a rarely-called or latency-tolerant function, like a background batch job, you'd be paying continuously for warm capacity a cold start would barely be noticed on.
How do we know if serverless is still the right fit for a workload?
If the workload has grown into steady, predictable, always-on traffic, the pay-only-for-what-you-use case for serverless weakens regardless of how well you've mitigated cold starts. At that point, compare the real cost of heavily provisioned serverless capacity against a small always-on container for that specific workload.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Cutting Serverless Cold Starts Without Overpaying
A function that answers in 80 milliseconds warm takes four seconds cold. What actually drives cold start time, and when provisioned concurrency is worth it.
The Serverless Cold Start Fixes That Actually Move the Number
A runbook for reducing serverless cold start latency: runtime choice, package size, provisioned concurrency, and the fixes that don't actually help.
Cutting Serverless Cold Start Time Without Just Throwing Money at It
Practical ways to reduce serverless cold start latency, from runtime and package size to provisioned concurrency, and when each one is actually worth the cost.
A Runbook for Cutting Serverless Cold Start Latency
A step-by-step runbook for reducing serverless cold start latency, from trimming your deployment package to deciding when to pay for provisioned capacity.
Serverless Cold Starts: What's Actually Fixable and What Isn't
Provisioned concurrency, bundle size, and runtime choice: a decision guide for which cold start fixes actually move your P99 latency.
Cutting Serverless Cold Starts Without Giving Up on Serverless
Where cold start time actually goes, when provisioned concurrency is worth paying for, and which functions don't need the fix at all.