Developer Productivity & Platform EngineeringPlaybook3 min readUpdated September 2026

Cutting Serverless Cold Starts Without Overpaying

A function that normally responds in about 80 milliseconds takes four seconds on the first request after a quiet stretch, because the platform has to provision a fresh execution environment from scratch before your handler ever runs. A customer hits that exact request right as traffic starts picking up, and the slow first load costs you the impression.

Fixing this doesn't always mean paying to keep everything warm. Most of the win is available for free by trimming what actually happens before the handler runs.

What Actually Happens During a Cold Start

The platform provisions a new execution environment, loads the runtime, and runs your initialization code, imports, connection setup, dependency loading, before your handler function ever executes. The size of your deployment package and the amount of work done at that import-time stage directly drive how long all of this takes, more than most people expect.

Trimming What Runs Before the Handler

Move expensive imports and connections that aren't needed on every single code path into lazy initialization inside the handler itself, so they only run when actually needed. A smaller deployment package with fewer unused dependencies loads faster on a cold start regardless of which language runtime you're using, and this fix costs nothing beyond the engineering time to make the change.

Provisioned Concurrency vs Just Eating the Latency

Keeping a set number of execution environments warm removes cold starts entirely for that portion of traffic, at the direct cost of paying for that reserved capacity whether it's actively used or not. Reserve it for the specific endpoints where a slow first request genuinely costs you something, a customer-facing checkout step, not the entire application by default.

For example, a team runs two functions: a checkout step called constantly and a nightly report generator called once a day. The checkout function rarely cold starts because traffic keeps it warm, so reserved capacity buys little there. The report function cold starts every run, but nobody is waiting on it, so paying to keep it warm is wasted spend. A common mistake is reserving capacity for the function that feels important instead of the one where a customer actually notices a slow first request. Decide by asking who waits on the response and how often the function sits idle.

Language and Runtime Choices Matter More Than People Expect

A runtime with a heavier startup cost, a large framework doing significant class loading or module resolution at boot, will cold start slower than a lighter one doing equivalent work, all else being equal. This isn't a reason to rewrite an existing service that's working fine; it's worth weighing specifically when you're building a new function where first-request latency genuinely matters to the customer experience.

A Checklist Before You Reach for Provisioned Concurrency

  • Confirm cold starts are actually happening often enough to matter by checking your real invocation frequency, rather than assuming based on how the app feels anecdotally.
  • Trim initialization code first, since it costs nothing beyond engineering time and often captures most of the available improvement.
  • Scope provisioned concurrency to the specific functions where latency is directly customer-facing, not the whole deployment by default.

Watching a Traffic-Shape Change Undo Your Fix

A function that rarely cold starts today because it gets called constantly can start cold starting frequently again after a traffic pattern shifts, a feature gets deprecated and calls that endpoint less, or a batch job that used to keep it warm as a side effect gets rescheduled. Cold start behavior isn't a one-time fix; it's a function of current call frequency, which changes as the product does.

Revisit cold start metrics on functions you've previously tuned whenever their surrounding traffic pattern changes meaningfully, rather than assuming a fix made months ago is still doing its job. A dashboard that shows cold start rate per function over time catches this drift far sooner than waiting for a customer complaint to surface it.

Treat provisioned concurrency spend the same way: review it against current invocation frequency periodically rather than setting it once and forgetting about it. A function that no longer needs reserved capacity because its traffic pattern changed is a quiet, recurring cost with nothing to show for it until someone actually checks.

Put a specific engineer's name on that periodic review, the same way you'd assign an owner to any other recurring infrastructure cost line, rather than leaving it to whoever happens to notice the bill looks a little larger than expected that particular month.

Executive Capability Standard

What Good Looks Like

Good cold start management means initialization code is trimmed to only what's genuinely needed before the handler runs, and provisioned concurrency, when used, is scoped to the specific endpoints where latency actually costs you something.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Measure your actual cold start frequency and duration for your highest-traffic functions before assuming it's a real problem.
2. Do Manually:Manually move one function's non-essential imports to lazy initialization and measure the difference in cold start time.
3. Delegate:Assign one engineer to own deployment package size for functions where cold start latency is customer-facing.
4. Automate:Automate a build-time check that flags deployment package size growth past a threshold before it ships.
5. Buy:Bring in a platform engineer to right-size provisioned concurrency if the current setup is applied uniformly instead of scoped to what actually needs it.

How to Get Started

Frequently Asked Questions

How much does trimming initialization code actually help?

Often a substantial share of total cold start time, since a smaller deployment package and lazy-loaded dependencies both reduce work that happens before your handler ever runs. It's also free beyond the engineering time to make the change, which is why it's worth doing before paying for provisioned concurrency.

Is provisioned concurrency worth the extra cost?

For endpoints where a slow first request directly affects a customer, often yes. For infrequently called internal or batch functions where nobody is waiting on the response in real time, the reserved capacity cost usually isn't worth it. Scope it function by function rather than turning it on for everything.

Does the choice of runtime language actually matter for cold starts?

Yes, all else being equal a runtime with heavier startup overhead will cold start slower. It's rarely worth rewriting an existing working service just for this, but it's a real factor to weigh when building a new function that's specifically sensitive to first-request latency.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides