Data Engineering & Real-Time Event StreamsPlaybook3 min readUpdated September 2026

Cutting Serverless Cold Start Time Without Rewriting Everything

To cut serverless cold start time, work through the cheap changes first: shrink the deployment package, tune initialization code, reuse connections, and only then pay for provisioned concurrency. A cold start is the extra latency a function pays on a fresh execution environment, and it hurts most with bursty or infrequent traffic.

Start with the cheapest changes before reaching for the more expensive ones, since a smaller package and a faster runtime can often close much of the gap before you need to pay for provisioned concurrency, though results vary by workload.

Package size is usually the first place to look

A larger deployment package takes longer to download and initialize on a fresh execution environment, and this is often the single biggest lever available without any architectural change. Audit your function's dependencies for anything unused or only needed at build time, and split large, infrequently changed dependencies into a separate layer where your platform supports it, so a code change doesn't force re downloading dependencies that didn't actually change.

Measure the actual before and after cold start time when trimming package size, not just the package size itself. The relationship between package size and cold start time isn't perfectly linear, so it's worth confirming a given trimming effort is actually moving the number that matters.

Runtime choice and initialization code both matter

Different language runtimes have meaningfully different baseline cold start characteristics, and a function written in a runtime with a heavier startup cost will carry that cost on every cold start regardless of how well tuned the rest of the code is. This isn't always worth rewriting an existing function over, but it's worth factoring into the decision for a new, latency sensitive function.

Code that runs at module load time, outside the actual handler function, executes on every cold start and adds directly to that latency. Move anything that isn't strictly required for the first invocation, a connection that can be lazily initialized, an expensive computation that could be cached instead, out of module load time and into the handler where it only runs when actually needed.

Reusing connections across warm invocations

A function that opens a fresh database connection on every single invocation pays that connection cost repeatedly even on warm starts, not just cold ones. Initialize expensive resources, a database connection, an SDK client, outside the handler function but check for their existence before reinitializing, so a warm invocation reuses what a previous invocation already set up rather than repeating the work.

Be careful with this pattern against a connection pool ceiling, since a serverless function scaling out to many concurrent instances can each hold their own connection and collectively exhaust a downstream database's connection limit quickly, which connects directly back to the connection pooling problem covered elsewhere in this guide.

When provisioned concurrency is actually worth paying for

Provisioned concurrency keeps a set number of execution environments warm and ready, eliminating cold starts entirely for traffic within that provisioned capacity, at the direct cost of paying for that capacity whether or not it's actively being used. This is worth it specifically for a latency sensitive path with predictable minimum traffic, not for a function that runs rarely and can tolerate an occasional slower cold invocation.

Size provisioned concurrency against your actual traffic pattern's baseline, not its peak, and let burst traffic above that baseline fall back to on demand cold starts. Provisioning for peak traffic that only happens briefly each day usually costs far more than the latency problem it's solving is actually worth.

Measuring whether any of this actually mattered

It's easy to make several of these changes at once and lose track of which one actually moved the needle. Change one thing at a time where practical, package size, then module load code, then connection reuse, and measure cold start latency after each change against a consistent baseline, so you know which fixes are worth carrying into your other functions and which weren't worth the effort.

Set a concrete target tied to what actually matters for the function's use case, a page load waiting on this function, an event that needs to process within a specific window, rather than chasing the lowest possible cold start time for its own sake. A function well within its actual latency budget doesn't need further optimization just because the number could theoretically go lower.

Work through the fixes in this order, measuring after each:

  1. Trim package size by removing unused or build time only dependencies, and split large, stable ones into a separate layer where supported.
  2. Move unnecessary work out of module load time, and factor runtime startup cost into decisions about new functions.
  3. Initialize database connections and SDK clients outside the handler, and reuse them on warm invocations.
  4. Pay for provisioned concurrency only on latency sensitive paths with predictable minimum traffic.
  5. Record cold start latency against a consistent baseline after every change, so you know which fixes are worth repeating.
Executive Capability Standard

What Good Looks Like

A well tuned serverless setup trims package size and module load time first, reuses expensive resources safely across warm invocations, and reserves provisioned concurrency specifically for latency sensitive paths with predictable baseline traffic.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Measure current cold start time for your most latency sensitive functions and identify unused dependencies contributing to package size.
2. Do Manually:Trim package size and move initialization code out of module load time for your single worst offending function, then measure the improvement.
3. Delegate:Assign an engineer to own cold start performance across your serverless functions and to review new functions against these practices before they ship.
4. Automate:Add package size and cold start time checks to your deployment pipeline so regressions are caught before they reach production.
5. Buy:Bring in a fractional CTO or platform specialist if cold start latency is causing a measurable, customer facing problem on a critical path today.

How to Get Started

Frequently Asked Questions

What's the single most effective change for reducing cold start time?

For most functions, trimming package size and moving unnecessary code out of module load time into the handler are the cheapest and most effective first steps. Measure actual before and after cold start time when making these changes, since the relationship between package size and latency isn't perfectly linear.

Is provisioned concurrency worth the extra cost for every function?

No. It's worth it specifically for a latency sensitive path with predictable minimum traffic, where eliminating cold starts matters for user experience. For a function that runs rarely or can tolerate an occasional slower invocation, the ongoing cost of keeping capacity warm usually isn't justified.

Can reusing database connections across warm invocations cause its own problems?

Yes, if it isn't managed carefully. A function scaling out to many concurrent instances, each holding its own reused connection, can collectively exhaust a downstream database's connection limit. Pair connection reuse with proper pooling on the database side so scaling out doesn't turn into a connection exhaustion incident.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides