Finding a Memory Leak Before It Finds Your Pager
To find a memory leak, take two heap snapshots under load and diff them in Node.js, or read a pprof profile in Go, before the process gets killed and the evidence disappears. Watching RSS climb proves a leak exists but not which object, closure, or goroutine holds the memory, and a restart wipes the clues.
Here's a worked example of catching that evidence in both Node.js and Go before the process dies.
Node.js: Take Two Heap Snapshots and Diff Them
Run the service with `--inspect` enabled, or trigger a heap snapshot programmatically, at two points in time separated by enough load to make a real leak visible against normal noise, an hour or more under typical traffic. Load both snapshots into a comparison view and sort by the objects that grew the most between the two, not by total size alone, since a large but stable cache looks different from a collection that's growing unbounded. The comparison view, not either snapshot in isolation, is what actually points at the leak.
Node.js: The Usual Suspects Are Closures and Listeners
A closure that captures a reference to something large and outlives its useful purpose, an event listener registered but never removed, an entry added to a Map or WeakMap-that-should-have-been-a-WeakMap, are the most common real-world Node.js leak patterns. An event emitter that accumulates listeners across repeated calls to a function that re-subscribes without ever unsubscribing is a particularly common one in long-running services, and it often only becomes visible after enough repetitions that a normal test suite, which runs each path only a handful of times, never catches it.
Go: pprof and the Difference Between a Slow Leak and a Goroutine Leak
Go's built-in `pprof` heap profile shows memory allocation by call site, which points you toward which function is allocating memory that isn't getting freed. A goroutine leak, one that's started but never returns because it's blocked on a channel nobody ever sends to or closes, shows up differently, as a steadily growing goroutine count in the goroutine profile rather than in heap allocations directly, though it often drags memory up alongside it since each leaked goroutine holds its own stack and any captured references. Check both profiles; they point at different, equally common failure patterns.
Capture Evidence Before the Process Gets Killed, Not After
A heap snapshot or profile taken after an out-of-memory kill is nothing, because the process and its memory state are already gone. Set up automatic capture triggered by a memory threshold crossing, well before your actual limit, so a snapshot exists from a moment when the leak was visible but the process was still alive and responsive enough to produce a clean profile. This single piece of automation is the difference between diagnosing a leak in an afternoon and reconstructing one from vague symptoms after it's already caused several restarts.
To set up automatic capture, work through these steps:
- Pick a memory threshold well below your actual limit, so the capture fires while the process is still alive and responsive.
- Trigger a heap snapshot or profile automatically when memory crosses that threshold, instead of waiting for an out-of-memory kill.
- Test the capture in staging first, since taking a snapshot pauses the process briefly and adds some overhead.
- Compare the captured snapshot against an earlier one to see which objects grew the most between the two.
RSS Growth Isn't Always a Leak, Check the Baseline First
Some memory growth is normal and expected, a cache filling up to its intended size, a connection pool warming up, garbage collection running less aggressively under light load and reclaiming memory later than you'd expect. Compare RSS growth against a period of genuinely idle traffic; a service that grows under load and then plateaus, even at a higher baseline than before, is probably fine, while one that keeps climbing during idle periods with no new work coming in is the pattern that actually indicates a leak worth chasing.
Reproducing a Leak Outside Production, Deliberately
Once you have a suspected culprit from a snapshot diff or a pprof allocation profile, write a small, repeatable script that exercises just that code path in a loop and watch memory under a profiler while it runs. This turns a vague, multi-day production symptom into a fast, local feedback loop you can iterate against, which is a much better position than trying to confirm a fix by deploying to production and watching a dashboard for another few days to see if the trend finally flattens out.
What to Do When You Can't Reproduce It Locally at All
Some leaks only show up under production-scale traffic or a specific, hard-to-replicate data shape, a particular customer's usage pattern, a rare combination of concurrent requests. For these, lean harder on the automated snapshot-on-threshold capture from production itself, and consider running the suspect service temporarily with more verbose allocation tracking enabled on a single instance pulled from the load balancer rotation, trading a bit of that instance's performance for the diagnostic detail you can't get any other way.
What Good Looks Like
You should be able to point to the exact object or goroutine keeping memory from being freed, not just watch RSS climb and restart the process.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How much load should pass between the two heap snapshots we compare?
Enough load that the growth clearly exceeds normal allocation noise, typically an hour or more of production-like traffic. A quick synthetic test is too short, because a slow, real leak can look statistically insignificant in a small window.
Can we safely take a heap snapshot in production without affecting users?
Usually yes, but a snapshot pauses the process briefly and adds some overhead, so test the impact first. Try it in staging, and for latency-sensitive services capture during a lower-traffic window or on a single instance pulled from the load balancer rotation.
What's the fastest way to tell a memory leak from normal GC behavior in Go?
Check whether memory returns to a stable baseline after a garbage collection cycle completes during idle periods. Memory that drops after GC runs is being managed normally; memory that keeps climbing even across multiple GC cycles with no new load is the pattern worth profiling further.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Finding a Memory Leak Before It Pages You
A service's memory climbs for days until it gets killed and restarts, then climbs again. How to profile Node and Go to trace a leak to its real cause.
Chasing Down a Slow Memory Leak in Node or Go Before It Pages You
A worked walkthrough of finding a slow memory leak using heap snapshots in Node and pprof in Go, before it turns into a middle of the night restart loop.
Finding a Memory Leak in Node or Go Before It Takes Down a Pod
How to use heap snapshots and pprof to find a real memory leak in Node.js or Go, and the common causes behind a slow, steady memory climb in production.
Profiling a Memory Leak in Node or Go Before It Pages You
A practical approach to finding a memory leak in Node.js or Go, including the tools to reach for first and the leak patterns specific to each.
Your Service Restarts Itself Every Night. That's Not Normal
How to profile and find a real memory leak in Node or Go, why a scheduled restart hides the symptom without fixing anything, and where to start looking.
Tracking Down a Slow Memory Leak in Node or Go, Step by Step
A worked walkthrough of profiling and fixing a slow memory leak in a Node.js or Go service, from spotting the pattern to confirming the fix actually worked.