Finding a Memory Leak in Node or Go Before It Takes Down a Pod
A memory leak rarely announces itself. What you actually see is a pod's memory usage climbing steadily over hours or days until it hits its limit and gets killed and restarted, which resets the number and quietly hides the underlying cause until it happens again. Catching the actual leak, not just the restart, means capturing a heap profile while the process is still alive and comparing it over time.
The tools differ between Node.js and Go, but the method is the same: take a snapshot, let time pass, take another, and diff them.
Is it a memory leak or normal growth?
Memory usage that climbs and then plateaus, or climbs and drops back down after a garbage collection cycle, is often just normal allocation behavior, not a leak. A real leak is memory that keeps climbing indefinitely, cycle after cycle, without ever coming back down, because something is holding a reference that should have been released. Plot memory usage over several hours before concluding you have a leak at all; a chart that looks alarming over ten minutes sometimes looks completely normal over three hours once garbage collection has had time to run.
Node.js: Heap Snapshots and What to Look For
Node's built-in inspector can capture a heap snapshot you can load into Chrome DevTools, and the useful technique isn't looking at one snapshot, it's taking two, separated by enough time for the suspected leak to grow noticeably, and using the comparison view to see which object types grew between them. Look specifically for objects you'd expect to be temporary, like event listeners, closures, or cache entries, showing up with a growing count instead of staying flat. A common Node-specific cause is an event listener added on every request without ever being removed, which silently accumulates until the process runs out of memory.
Go: pprof and the Allocation Profile That Actually Matters
Go's pprof tooling can capture a heap profile showing exactly which call sites are responsible for currently live allocations, which is more directly actionable than Node's snapshot comparison since it points at the allocating code path immediately. The profile to focus on is in-use memory, not total allocations, since a function that allocates a lot but releases it all is not your leak; a function whose allocations show up in the in-use profile and keep growing across successive captures is. A common Go-specific cause is a goroutine that never exits, holding references to data that should have been garbage collected once its work was done.
The Usual Suspects Across Both Languages
Regardless of runtime, a few patterns account for most real leaks:
- An unbounded in-memory cache with no eviction policy, growing forever as new keys get added.
- A subscription, listener, or timer that gets created per-request but never cleaned up when the request completes.
- A closure that captures a large object by reference and outlives the scope it was created for.
- A goroutine or async task that's supposed to exit but gets stuck waiting on something that never happens.
Checking your own code against this list before profiling can sometimes save you the profiling step entirely.
How do you confirm a memory leak fix worked?
Once the profile points at a specific allocation site, confirm the fix actually worked the same way you confirmed the problem: capture a new baseline, apply the fix, run under the same load for the same duration, and compare. A leak fix that isn't verified this way risks becoming a leak fix that only partially worked, or one that fixed a symptom while the actual root cause, an unbounded cache instead of the specific key you noticed growing, keeps running quietly in the background.
A Worked Example: The Cache That Never Evicted
Say a Node service caches API responses keyed by a combination of user ID and query parameters, and traffic grows steadily over a week until the pod starts restarting daily. Two heap snapshots taken a few hours apart show one object type growing steadily: the cache's internal map. The cache was built with no maximum size and no time-based expiry, so every unique combination of user and query parameters it had ever seen stayed in memory forever. The fix wasn't a code review catching a typo, it was adding a maximum entry count with least-recently-used eviction, confirmed afterward by watching the same map's growth flatten out under the same production traffic instead of climbing indefinitely.
What Good Looks Like
Good leak hygiene means memory growth gets confirmed against a multi-hour graph before it's called a leak, profiled with the right tool for the runtime, and any fix gets verified with a second profile under the same conditions, not just deployed and assumed to have worked.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Is it safe to take a heap snapshot on a production instance under real load?
Generally yes for a single capture, though it does pause the process briefly and adds some overhead. For an ongoing investigation, it's safer to reproduce the leak on a staging instance under simulated load, where you can take multiple snapshots over time without any risk to real users.
Why does restarting the pod make the memory graph look fine again?
A restart resets the process's entire memory state, so the leak's accumulated growth disappears along with the crash. This is why relying on automatic restarts as your only response to memory growth hides the underlying leak instead of fixing it, and the same leak will simply climb again on the same schedule after every restart.
How long should we wait between two heap snapshots to get a useful comparison?
Long enough for the suspected leak to grow by a noticeable amount relative to normal memory noise, often thirty minutes to a few hours depending on how fast the leak grows and how much traffic the instance is handling. A comparison taken too close together can be dominated by normal allocation variance rather than the actual leak.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Finding a Memory Leak Before It Pages You
A service's memory climbs for days until it gets killed and restarts, then climbs again. How to profile Node and Go to trace a leak to its real cause.
Finding a Memory Leak Before It Finds Your Pager
A worked walkthrough of diagnosing a memory leak: heap snapshots in Node.js, pprof in Go, and capturing evidence before the process gets killed.
Profiling a Memory Leak in Node or Go Before It Pages You
A practical approach to finding a memory leak in Node.js or Go, including the tools to reach for first and the leak patterns specific to each.
Tracking Down a Slow Memory Leak in Node or Go, Step by Step
A worked walkthrough of profiling and fixing a slow memory leak in a Node.js or Go service, from spotting the pattern to confirming the fix actually worked.
Your Service Restarts Itself Every Night. That's Not Normal
How to profile and find a real memory leak in Node or Go, why a scheduled restart hides the symptom without fixing anything, and where to start looking.
Chasing Down a Memory Leak in Node and Go Services
How to tell a real leak from normal garbage collection, take a useful heap snapshot, and stop shipping scheduled restarts as the fix.