Cutting Serverless Cold Start Time Without Rewriting Everything
To cut serverless cold start time, work through the cheap changes first: shrink the deployment package, tune initialization code, reuse connections, and only then pay for provisioned concurrency. A cold start is the extra latency a function pays on a fresh execution environment, and it hurts most with bursty or infrequent traffic.
Start with the cheapest changes before reaching for the more expensive ones, since a smaller package and a faster runtime can often close much of the gap before you need to pay for provisioned concurrency, though results vary by workload.
Package size is usually the first place to look
A larger deployment package takes longer to download and initialize on a fresh execution environment, and this is often the single biggest lever available without any architectural change. Audit your function's dependencies for anything unused or only needed at build time, and split large, infrequently changed dependencies into a separate layer where your platform supports it, so a code change doesn't force re downloading dependencies that didn't actually change.
Measure the actual before and after cold start time when trimming package size, not just the package size itself. The relationship between package size and cold start time isn't perfectly linear, so it's worth confirming a given trimming effort is actually moving the number that matters.
Runtime choice and initialization code both matter
Different language runtimes have meaningfully different baseline cold start characteristics, and a function written in a runtime with a heavier startup cost will carry that cost on every cold start regardless of how well tuned the rest of the code is. This isn't always worth rewriting an existing function over, but it's worth factoring into the decision for a new, latency sensitive function.
Code that runs at module load time, outside the actual handler function, executes on every cold start and adds directly to that latency. Move anything that isn't strictly required for the first invocation, a connection that can be lazily initialized, an expensive computation that could be cached instead, out of module load time and into the handler where it only runs when actually needed.
Reusing connections across warm invocations
A function that opens a fresh database connection on every single invocation pays that connection cost repeatedly even on warm starts, not just cold ones. Initialize expensive resources, a database connection, an SDK client, outside the handler function but check for their existence before reinitializing, so a warm invocation reuses what a previous invocation already set up rather than repeating the work.
Be careful with this pattern against a connection pool ceiling, since a serverless function scaling out to many concurrent instances can each hold their own connection and collectively exhaust a downstream database's connection limit quickly, which connects directly back to the connection pooling problem covered elsewhere in this guide.
When provisioned concurrency is actually worth paying for
Provisioned concurrency keeps a set number of execution environments warm and ready, eliminating cold starts entirely for traffic within that provisioned capacity, at the direct cost of paying for that capacity whether or not it's actively being used. This is worth it specifically for a latency sensitive path with predictable minimum traffic, not for a function that runs rarely and can tolerate an occasional slower cold invocation.
Size provisioned concurrency against your actual traffic pattern's baseline, not its peak, and let burst traffic above that baseline fall back to on demand cold starts. Provisioning for peak traffic that only happens briefly each day usually costs far more than the latency problem it's solving is actually worth.
Measuring whether any of this actually mattered
It's easy to make several of these changes at once and lose track of which one actually moved the needle. Change one thing at a time where practical, package size, then module load code, then connection reuse, and measure cold start latency after each change against a consistent baseline, so you know which fixes are worth carrying into your other functions and which weren't worth the effort.
Set a concrete target tied to what actually matters for the function's use case, a page load waiting on this function, an event that needs to process within a specific window, rather than chasing the lowest possible cold start time for its own sake. A function well within its actual latency budget doesn't need further optimization just because the number could theoretically go lower.
Work through the fixes in this order, measuring after each:
- Trim package size by removing unused or build time only dependencies, and split large, stable ones into a separate layer where supported.
- Move unnecessary work out of module load time, and factor runtime startup cost into decisions about new functions.
- Initialize database connections and SDK clients outside the handler, and reuse them on warm invocations.
- Pay for provisioned concurrency only on latency sensitive paths with predictable minimum traffic.
- Record cold start latency against a consistent baseline after every change, so you know which fixes are worth repeating.
What Good Looks Like
A well tuned serverless setup trims package size and module load time first, reuses expensive resources safely across warm invocations, and reserves provisioned concurrency specifically for latency sensitive paths with predictable baseline traffic.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
What's the single most effective change for reducing cold start time?
For most functions, trimming package size and moving unnecessary code out of module load time into the handler are the cheapest and most effective first steps. Measure actual before and after cold start time when making these changes, since the relationship between package size and latency isn't perfectly linear.
Is provisioned concurrency worth the extra cost for every function?
No. It's worth it specifically for a latency sensitive path with predictable minimum traffic, where eliminating cold starts matters for user experience. For a function that runs rarely or can tolerate an occasional slower invocation, the ongoing cost of keeping capacity warm usually isn't justified.
Can reusing database connections across warm invocations cause its own problems?
Yes, if it isn't managed carefully. A function scaling out to many concurrent instances, each holding its own reused connection, can collectively exhaust a downstream database's connection limit. Pair connection reuse with proper pooling on the database side so scaling out doesn't turn into a connection exhaustion incident.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Cutting Serverless Cold Starts Without Overpaying
A function that answers in 80 milliseconds warm takes four seconds cold. What actually drives cold start time, and when provisioned concurrency is worth it.
Serverless Cold Starts: What's Actually Fixable and What Isn't
Provisioned concurrency, bundle size, and runtime choice: a decision guide for which cold start fixes actually move your P99 latency.
The Serverless Cold Start Fixes That Actually Move the Number
A runbook for reducing serverless cold start latency: runtime choice, package size, provisioned concurrency, and the fixes that don't actually help.
Cutting Serverless Cold Start Time Without Just Throwing Money at It
Practical ways to reduce serverless cold start latency, from runtime and package size to provisioned concurrency, and when each one is actually worth the cost.
A Runbook for Cutting Serverless Cold Start Latency
A step-by-step runbook for reducing serverless cold start latency, from trimming your deployment package to deciding when to pay for provisioned capacity.
Cutting Serverless Cold Starts Without Abandoning Serverless
Why serverless cold starts happen, which patterns make them worse, and the mitigation options that don't quietly turn serverless into servers.