AI Model Serving & Inference OptimizationPlaybook3 min readUpdated September 2026

Where Serverless Cold Start Time Actually Goes

A serverless function's cold start gets blamed as one undifferentiated delay, when it's actually three things stacked on top of each other: the platform provisioning a container, the runtime initializing, and your own code's setup work running for the first time. Fixing the wrong one wastes effort and doesn't move the number.

Knowing which part is actually slow, and whether the path it's slow on even needs to be fast, decides whether the fix is a configuration change, a code change, or leaving serverless behind for that one path.

Where the time actually goes

Container provisioning is mostly outside your control, the platform allocating and starting the execution environment. Runtime initialization varies by language and is partly outside your control too, though your choice of runtime affects its baseline. Your own code's init phase, importing dependencies, constructing SDK clients, opening database connections, is the one part entirely within your control, and it's often the largest of the three.

Most teams blame the platform for a slow cold start without measuring which of the three phases is actually responsible. Add timing around your own init code before assuming the fix has to come from the platform side.

Provisioned concurrency: paying to skip the wait

Provisioned concurrency keeps a set number of instances warm continuously, so a request hits an already-initialized function instead of waiting through a cold start. It's worth paying for on latency-sensitive, user-facing paths where a slow response is directly visible to someone waiting on it.

It's usually wasted spend on background or async work, a queue consumer, a scheduled job, anything where the extra latency of an occasional cold start doesn't actually cost you anything a user notices. Scope provisioned concurrency to the specific functions on the critical path, not applied uniformly across everything.

Trimming what your init phase actually does

Lazy-load dependencies you don't need on every invocation instead of importing everything up front. Construct database clients, HTTP clients, and SDK objects outside the handler function so they're created once and reused across warm invocations, rather than rebuilt from scratch on every single call.

This matters even on already-warm invocations, not just cold ones, since a client rebuilt inside the handler adds overhead every time regardless of whether the container itself was cold. It's a cheap change with a benefit that compounds across your actual request volume, not just the fraction of requests that hit a cold container.

Cheap fixes to try before paying for warm capacity:

  • Add timing around your own init code first, so you know which of the three phases is actually responsible before changing anything.
  • Lazy-load dependencies that not every invocation needs instead of importing everything up front.
  • Construct database, HTTP, and SDK clients outside the handler so they are created once and reused across warm invocations.
  • Reserve provisioned concurrency for latency-sensitive, user-facing paths, and skip it for queue consumers and scheduled background work.

Runtime choice changes the baseline

Compiled languages generally start faster than runtimes carrying a heavier startup cost, since there's less for the runtime to initialize before your code can run. That's a real difference, and it's also not a reason to rewrite an existing service in a different language purely to shave cold start time, the migration cost usually dwarfs the saving.

Where it matters is a genuinely new, latency-sensitive function: knowing the runtime tradeoff going in is worth factoring into the choice, rather than discovering it after the function is already built and the cold start numbers come back worse than expected.

When a cold start is telling you serverless is the wrong fit

Very spiky, latency-critical traffic on a single path is the scenario where fighting cold starts starts to cost more, in provisioned concurrency spend and engineering time, than the problem is worth. If a path needs consistently low latency at meaningful, sustained volume, a small always-on service can end up simpler and cheaper than keeping a function artificially warm around the clock to avoid a cold start that shouldn't be happening in the first place.

That's not an argument against serverless generally, it's an argument for matching the tool to the specific path's actual traffic shape rather than defaulting to serverless everywhere and patching around the cases where it doesn't fit. The rest of your system can stay serverless without issue; moving one consistently hot, latency-sensitive path off it isn't a step backward, it's using each tool for the traffic pattern it actually handles well.

Before making that call, confirm the traffic pattern is real and not a one-off, a single busy launch day looks the same as a permanently hot path until you've watched it for a few weeks. A short spike is worth absorbing with provisioned concurrency; a sustained one is worth reconsidering the platform for.

Executive Capability Standard

What Good Looks Like

Good here means you know exactly which part of your cold start time is provisioning, runtime init, or your own code, and you've only paid to fix the part that's actually slow on the paths where it matters.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Instrument your function's cold start with timing around your own init code so you can see how much of the delay is actually yours.
2. Do Manually:Trim your init phase by hand: lazy-load what you don't need immediately and move client setup outside the handler.
3. Delegate:Have someone review which functions sit on latency-sensitive user paths versus background work, since that split decides where provisioned concurrency is worth the cost.
4. Automate:Set provisioned concurrency automatically on the functions identified as latency-sensitive, scaled to your typical concurrent traffic.
5. Buy:If one path needs consistently low latency at high volume, a small always-on service may cost less than keeping a function warm around the clock.

How to Get Started

Frequently Asked Questions

Does putting a function in a VPC make cold starts worse?

Historically yes, due to network interface attachment overhead, though cloud providers have narrowed that gap significantly. Check current documented behavior for your specific platform before assuming it's still a major factor in your cold start time.

Is provisioned concurrency worth it for a background job?

Usually not. A background job can typically tolerate the extra latency a cold start adds without anyone noticing, so paying to keep it warm continuously is often spend without a matching benefit. Save it for user-facing, latency-sensitive paths.

Does a bigger memory allocation help with cold starts?

Often yes, since a larger memory allocation typically comes with proportionally more CPU during initialization on most serverless platforms. It varies by workload and platform, so test it against your actual function rather than assuming it helps uniformly.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides