Cutting Serverless Cold Starts Without Overpaying
A function that normally responds in about 80 milliseconds takes four seconds on the first request after a quiet stretch, because the platform has to provision a fresh execution environment from scratch before your handler ever runs. A customer hits that exact request right as traffic starts picking up, and the slow first load costs you the impression.
Fixing this doesn't always mean paying to keep everything warm. Most of the win is available for free by trimming what actually happens before the handler runs.
What Actually Happens During a Cold Start
The platform provisions a new execution environment, loads the runtime, and runs your initialization code, imports, connection setup, dependency loading, before your handler function ever executes. The size of your deployment package and the amount of work done at that import-time stage directly drive how long all of this takes, more than most people expect.
Trimming What Runs Before the Handler
Move expensive imports and connections that aren't needed on every single code path into lazy initialization inside the handler itself, so they only run when actually needed. A smaller deployment package with fewer unused dependencies loads faster on a cold start regardless of which language runtime you're using, and this fix costs nothing beyond the engineering time to make the change.
Provisioned Concurrency vs Just Eating the Latency
Keeping a set number of execution environments warm removes cold starts entirely for that portion of traffic, at the direct cost of paying for that reserved capacity whether it's actively used or not. Reserve it for the specific endpoints where a slow first request genuinely costs you something, a customer-facing checkout step, not the entire application by default.
For example, a team runs two functions: a checkout step called constantly and a nightly report generator called once a day. The checkout function rarely cold starts because traffic keeps it warm, so reserved capacity buys little there. The report function cold starts every run, but nobody is waiting on it, so paying to keep it warm is wasted spend. A common mistake is reserving capacity for the function that feels important instead of the one where a customer actually notices a slow first request. Decide by asking who waits on the response and how often the function sits idle.
Language and Runtime Choices Matter More Than People Expect
A runtime with a heavier startup cost, a large framework doing significant class loading or module resolution at boot, will cold start slower than a lighter one doing equivalent work, all else being equal. This isn't a reason to rewrite an existing service that's working fine; it's worth weighing specifically when you're building a new function where first-request latency genuinely matters to the customer experience.
A Checklist Before You Reach for Provisioned Concurrency
- Confirm cold starts are actually happening often enough to matter by checking your real invocation frequency, rather than assuming based on how the app feels anecdotally.
- Trim initialization code first, since it costs nothing beyond engineering time and often captures most of the available improvement.
- Scope provisioned concurrency to the specific functions where latency is directly customer-facing, not the whole deployment by default.
Watching a Traffic-Shape Change Undo Your Fix
A function that rarely cold starts today because it gets called constantly can start cold starting frequently again after a traffic pattern shifts, a feature gets deprecated and calls that endpoint less, or a batch job that used to keep it warm as a side effect gets rescheduled. Cold start behavior isn't a one-time fix; it's a function of current call frequency, which changes as the product does.
Revisit cold start metrics on functions you've previously tuned whenever their surrounding traffic pattern changes meaningfully, rather than assuming a fix made months ago is still doing its job. A dashboard that shows cold start rate per function over time catches this drift far sooner than waiting for a customer complaint to surface it.
Treat provisioned concurrency spend the same way: review it against current invocation frequency periodically rather than setting it once and forgetting about it. A function that no longer needs reserved capacity because its traffic pattern changed is a quiet, recurring cost with nothing to show for it until someone actually checks.
Put a specific engineer's name on that periodic review, the same way you'd assign an owner to any other recurring infrastructure cost line, rather than leaving it to whoever happens to notice the bill looks a little larger than expected that particular month.
What Good Looks Like
Good cold start management means initialization code is trimmed to only what's genuinely needed before the handler runs, and provisioned concurrency, when used, is scoped to the specific endpoints where latency actually costs you something.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How much does trimming initialization code actually help?
Often a substantial share of total cold start time, since a smaller deployment package and lazy-loaded dependencies both reduce work that happens before your handler ever runs. It's also free beyond the engineering time to make the change, which is why it's worth doing before paying for provisioned concurrency.
Is provisioned concurrency worth the extra cost?
For endpoints where a slow first request directly affects a customer, often yes. For infrequently called internal or batch functions where nobody is waiting on the response in real time, the reserved capacity cost usually isn't worth it. Scope it function by function rather than turning it on for everything.
Does the choice of runtime language actually matter for cold starts?
Yes, all else being equal a runtime with heavier startup overhead will cold start slower. It's rarely worth rewriting an existing working service just for this, but it's a real factor to weigh when building a new function that's specifically sensitive to first-request latency.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Serverless Cold Starts: What's Actually Fixable and What Isn't
Provisioned concurrency, bundle size, and runtime choice: a decision guide for which cold start fixes actually move your P99 latency.
The Serverless Cold Start Fixes That Actually Move the Number
A runbook for reducing serverless cold start latency: runtime choice, package size, provisioned concurrency, and the fixes that don't actually help.
Cutting Serverless Cold Start Time Without Rewriting Everything
A step by step approach to reducing serverless cold start latency, from runtime and package size to provisioned concurrency, and when each is worth it.
Cutting Serverless Cold Start Time Without Just Throwing Money at It
Practical ways to reduce serverless cold start latency, from runtime and package size to provisioned concurrency, and when each one is actually worth the cost.
A Runbook for Cutting Serverless Cold Start Latency
A step-by-step runbook for reducing serverless cold start latency, from trimming your deployment package to deciding when to pay for provisioned capacity.
Cutting Serverless Cold Starts Without Abandoning Serverless
Why serverless cold starts happen, which patterns make them worse, and the mitigation options that don't quietly turn serverless into servers.