Production RAG & Vector Data ArchitecturePlaybook3 min readUpdated September 2026

A Runbook for Cutting Serverless Cold Start Latency

A cold start happens when a serverless platform has to initialize a fresh execution environment for a function before it can run, and that initialization time shows up directly as latency on whatever request triggered it. For a background job, a cold start might not matter at all. For a user-facing API endpoint, it's the difference between a snappy response and one that makes a customer wonder if something's broken.

This runbook walks through the fixes in the order they're usually worth trying, from cheapest to most expensive.

Step 1: How do you trim your deployment package?

A larger deployment package takes longer to load into a fresh execution environment, and unused dependencies bundled into it are pure cold start cost with no benefit. Audit your function's dependencies and remove anything not actually needed at runtime, and check whether your build process is bundling development-only tooling into the production package by mistake, which is a surprisingly common and easy-to-fix source of unnecessary bulk. This step alone is often the highest return on effort of anything in this runbook, since it costs nothing but attention and never makes anything worse.

Step 2: move expensive initialization out of the hot path

Code that runs at the top level of your function file, outside the handler itself, executes once per cold start and gets reused across warm invocations, which makes it the right place for expensive setup like establishing a database connection. Code that mistakenly does that same expensive setup inside the handler on every single invocation pays the cost repeatedly instead of once, which is a common and costly mistake worth specifically checking for in your own functions.

Step 3: choose a runtime with a faster startup profile if you have flexibility

Different language runtimes have meaningfully different baseline cold start characteristics, and a function currently written in a slower-starting runtime might be worth rewriting in a faster one if cold start latency is a genuine, measured problem for that specific function. This is a bigger lift than the previous two steps, so reserve it for functions where the latency actually matters to users, not as a blanket policy applied everywhere regardless of impact.

Step 4: keep functions warm deliberately, for the ones that matter most

A scheduled trigger that invokes a function on a regular interval keeps at least one execution environment warm, avoiding cold starts for traffic that arrives between those scheduled pings. This adds a small, ongoing cost and some complexity, so apply it selectively to the specific functions where cold start latency has a real, measured business cost, rather than to every function in your deployment regardless of whether anyone would notice the difference. Remember that a single warm environment only helps traffic that doesn't exceed it concurrently, so a sudden burst can still trigger fresh cold starts on top of the one you've kept warm.

Before adding a keep-warm trigger, decide which functions have earned one. A useful rule is to warm only functions that sit on a user-facing path and see intermittent traffic with long gaps, because those are the ones where a cold start lands on a real person. For example, a checkout function that receives a handful of requests overnight is a good candidate, while a nightly batch job is not. Measure cold versus warm latency for the candidate first, apply the trigger, then measure again. If the gap didn't shrink for real traffic, remove the trigger, since an unhelpful keep-warm job adds cost and complexity without delivering anything.

Step 5: When should you pay for provisioned capacity?

Most serverless platforms offer a provisioned or reserved capacity option that keeps a set number of execution environments permanently warm, eliminating cold starts entirely for that allocation, at a real ongoing cost regardless of actual traffic. This genuinely solves the problem, and it's also the most expensive fix on this list, which is exactly why it belongs at the end: reach for it once the cheaper steps have been tried and measured, for the specific functions where the remaining latency still isn't acceptable.

In order from cheapest to most expensive, the fixes are:

  1. Trim the deployment package by removing unused dependencies and development-only tooling that got bundled into production.
  2. Move expensive setup, such as opening a database connection, outside the handler so it runs once per cold start instead of on every invocation.
  3. Consider a runtime with a faster startup profile, but only for functions where measured cold start latency really matters to users.
  4. Keep selected functions warm with a scheduled trigger, accepting a small ongoing cost and that a burst can still cause fresh cold starts.
  5. Pay for provisioned capacity last, for the specific functions where the remaining latency still isn't acceptable after the cheaper steps.

A worked example: the fix that cost nothing

Say a checkout function pulls in a full-featured logging library that initializes a network connection to an external service at startup, even though the function only uses one simple logging call. Trimming that dependency down to just what's needed removes a meaningful chunk of initialization time from every cold start, at zero ongoing cost and with nothing new to operate or pay for. Teams that reach straight for provisioned capacity without first checking their own package for this kind of avoidable weight often end up paying monthly for a problem a single dependency cleanup would have mostly solved.

Executive Capability Standard

What Good Looks Like

A well-tuned serverless function keeps its deployment package lean, moves expensive setup out of the per-invocation hot path, and reserves paid warm capacity for the specific, measured cases where cheaper fixes aren't enough.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Understand which of your functions actually experience meaningful, user-facing cold start latency, rather than assuming it's a problem everywhere.
2. Do Manually:Manually review your highest-traffic function's dependencies and initialization code for easy trims and hot-path mistakes.
3. Delegate:Assign a standard deployment package review to whoever owns your build pipeline, so unused dependencies don't accumulate silently over time.
4. Automate:Automate a scheduled warming trigger for the specific functions where cold start latency has a measured, real business cost.
5. Buy:Pay for provisioned or reserved capacity on your highest-value user-facing functions once cheaper fixes have been tried and measured.

How to Get Started

Frequently Asked Questions

Do all serverless platforms have the same cold start characteristics?

No, they vary meaningfully by platform and by runtime within a given platform. Measure your own functions directly rather than relying on a general reputation, since the actual number that matters is what your specific function experiences under your specific configuration, not a generic industry figure.

Is a cold start ever completely unavoidable?

For traffic patterns with genuine cold, unpredictable spikes after long idle periods, some cold start exposure is difficult to fully eliminate without paying for permanently warm capacity. The goal for most teams is reducing frequency and duration to an acceptable level for the specific function, not chasing a theoretical zero.

How do we measure whether cold starts are actually a problem worth fixing?

Look at your latency metrics broken out specifically for cold versus warm invocations, not just an overall average that blends the two together and hides the real gap. If cold start latency on a user-facing path is meaningfully worse than the warm baseline and that path sees real, intermittent traffic gaps, it's worth the runbook above.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides