Finding a Memory Leak: A Heap Snapshot Walkthrough
A process that slowly climbs in memory usage over hours or days, then gets killed and restarted by an orchestrator before anyone investigates, is one of the easier production problems to ignore because the symptom disappears on its own. It's also one of the more expensive problems to leave alone, since every restart drops in-flight work and the underlying cause keeps happening quietly in the background.
This walks through confirming an actual leak rather than normal, bounded memory growth, then diagnosing it with heap snapshots in Node.js and with pprof in Go, since the tools differ but the diagnostic process is nearly identical in both.
The Symptom: Memory Climbs Until the Process Restarts
The pattern is usually visible on a memory graph before anyone reads a single line of code: a steady upward climb that never plateaus, unlike normal memory usage which grows during a busy period and then falls back down once that work completes. If your graph shows a return to baseline after load drops, that's probably healthy garbage collection doing its job, not a leak.
If it doesn't return to baseline, and instead climbs a bit further after each busy period before falling back only partway, that's the signature of an actual leak: something is being retained across garbage collection cycles that shouldn't be, and the retained amount compounds over time until the process runs out of memory or an orchestrator's restart policy kicks in first.
How do you confirm it's a leak, not normal growth?
Before diving into heap snapshots, rule out a few things that look like leaks but aren't. A cache with no eviction policy will grow until it plateaus at the size of everything that's ever been cached, which can look exactly like a leak on a short observation window but is actually bounded, just at a size nobody planned for. Check whether the growth curve is flattening slowly over a longer window before assuming it's unbounded.
Also confirm you're looking at the right memory metric. Total process memory includes the operating system's own caching behavior around freed memory that hasn't been returned to the OS yet, which can look like growth even when your application's actual live heap is stable. Compare application-level heap metrics specifically, not just total process resident memory, before concluding there's a real leak.
Taking and Comparing Heap Snapshots in Node.js
Take two heap snapshots separated by a meaningful gap, ideally spanning at least one full busy-then-idle cycle, so any object that should have been collected has had the chance to be. Compare the two snapshots specifically for objects whose count grew between them without a corresponding drop, rather than just looking at total heap size, which tells you something grew but not what.
Sort the comparison by the delta in object count or retained size, not by total count, since the object type responsible for a leak is often not the most numerous type overall, just the one that kept growing while everything else stayed flat. Once you've found the growing object type, trace its retainers, what's holding a reference to it, which is usually where the actual bug lives: an event listener that's never removed, an entry added to a map keyed by something like a request ID that's never cleaned up, a closure capturing more than it needs to.
Using pprof to Find the Same Pattern in Go
The equivalent process in Go uses pprof's heap profile, captured at two points in time the same way as the Node.js snapshots above, separated by enough of a gap to distinguish real growth from normal allocation churn. Compare the two profiles for the allocation sites and types whose retained memory grew between captures, rather than relying on a single snapshot, which only shows a moment in time and can't distinguish something that's growing from something that's simply large.
Go's garbage collector handles a lot of allocation patterns cleanly on its own, so a real leak here is more often something explicitly held alive: a slice or map that only grows, a goroutine that never exits and keeps its own local variables reachable indefinitely, a channel that's never drained. Trace the growing allocation back to the code path that creates it the same way you would in Node.js, since the diagnostic logic is the same even though the tooling looks different.
What are the common root causes of a growing object?
Once you've identified what's growing, the actual cause usually falls into a small set of patterns:
- An event listener or callback registered once per request or per connection and never removed, so the count grows with traffic instead of staying bounded.
- A cache, map, or dictionary with no eviction policy or size limit, growing unbounded as new keys are added and old ones are never removed.
- A goroutine or async task that never completes, keeping everything it references reachable for as long as it's alive, which in a genuine leak is forever.
- A closure or callback capturing a larger scope than it actually needs, keeping objects alive indirectly through a reference the code doesn't obviously need to hold.
Fix the root cause, not the symptom. A scheduled restart masks a leak instead of fixing it, and the underlying bug keeps consuming memory and dropping in-flight work on every restart until someone actually traces it back to the retaining reference.
What Good Looks Like
Good memory leak diagnosis means comparing two heap snapshots or profiles separated by a real time gap to see what's actually growing, not guessing from total memory usage alone.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How do we tell a memory leak apart from a cache that's just big?
Watch the growth curve over a longer window. An unbounded leak keeps climbing indefinitely. A cache with no eviction policy grows until it plateaus at the size of everything that's ever been cached, which can look identical to a leak on a short observation window but actually stops growing.
Should we just restart the process on a timer to manage memory growth?
That masks the symptom without fixing the cause. Restarts drop in-flight work, and the underlying leak keeps consuming memory between restarts regardless. Use a scheduled restart as a short-term stopgap while you diagnose the real cause, not as the permanent fix.
What's usually the actual bug once we've found the growing object type?
Most often an event listener or callback that's registered but never removed, a cache or map with no eviction policy, a goroutine or async task that never completes, or a closure capturing a larger scope than it needs. Trace what's holding a reference to the growing object to find which one applies.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Finding a Memory Leak Before It Pages You
A service's memory climbs for days until it gets killed and restarts, then climbs again. How to profile Node and Go to trace a leak to its real cause.
Finding a Memory Leak in Node or Go Before It Pages You
How to profile a memory leak with heap snapshots in Node and pprof in Go, common leak patterns in long-running services, and how to confirm a fix.
Finding a Memory Leak in Node or Go Before It Takes Down a Pod
How to use heap snapshots and pprof to find a real memory leak in Node.js or Go, and the common causes behind a slow, steady memory climb in production.
Profiling a Memory Leak in Node or Go Before It Pages You
A practical approach to finding a memory leak in Node.js or Go, including the tools to reach for first and the leak patterns specific to each.
Finding a Memory Leak Before It Finds Your Pager
A worked walkthrough of diagnosing a memory leak: heap snapshots in Node.js, pprof in Go, and capturing evidence before the process gets killed.
Your Service Restarts Itself Every Night. That's Not Normal
How to profile and find a real memory leak in Node or Go, why a scheduled restart hides the symptom without fixing anything, and where to start looking.