Production RAG & Vector Data ArchitecturePlaybook3 min readUpdated September 2026

Tracking Down a Slow Memory Leak in Node or Go, Step by Step

A memory leak rarely announces itself immediately. It shows up as a service that gets slowly slower over days, then eventually crashes or gets killed by an out-of-memory monitor, often at the least convenient time. The good news is that both Node.js and Go have solid built-in tooling for finding exactly what's holding onto memory it shouldn't be, once you know the sequence to follow.

Step 1: How do you confirm it's a leak and not just growth?

Memory usage that climbs and then plateaus at a stable, higher level is often normal, reflecting caches warming up or connection pools filling to their configured size. A real leak keeps climbing indefinitely, with no plateau, over a period of hours or days. Graph memory usage over a long enough window to distinguish the two before spending time chasing something that might just be expected behavior.

Step 2: take heap snapshots over time, not just one

A single heap snapshot tells you what's currently allocated, but not what's actually growing. Take two or more snapshots spaced well apart, under representative load, and compare them: the growing leak shows up as the object type or allocation site whose count or size increased the most between snapshots, standing out clearly against everything that stayed roughly flat.

Step 3: trace the growing allocation back to its retaining reference

Finding what's growing is only half the answer. The other half is finding what's still holding a reference to it, preventing it from ever being freed, such as an array or map that things get added to but never removed from, or an event listener that's registered repeatedly without ever being cleaned up. Most heap profiling tools can show you the retaining path directly, which turns this from a guessing exercise into a straightforward trace back through the code.

Step 4: What are the usual suspects to check first?

Certain patterns cause a disproportionate share of real leaks in both Node and Go: an event listener attached inside a function that runs repeatedly without ever being removed, a cache with no eviction policy that grows without bound, a goroutine in Go that blocks forever waiting on a channel nobody writes to again, or a closure that unintentionally captures and retains a large object. Checking these patterns first often finds the leak faster than a fully generic profiling pass.

Step 5: fix it, then reconfirm with the same snapshot comparison

After applying a fix, don't assume it worked. Rerun the same two-snapshot comparison from step two under the same representative load and confirm the previously growing allocation has actually stopped growing. A leak that seems fixed based on code review alone can still be present if the actual retaining reference wasn't the one that got removed, and the snapshot comparison is what proves the fix worked rather than just looking plausible.

The whole sequence, in short:

  1. Graph memory over hours or days to confirm it climbs without plateauing, rather than settling at a stable, higher level after cache warm-up.
  2. Take two or more heap snapshots well apart under representative load, and compare them to see which allocation grew the most.
  3. Trace the growing allocation back to the retaining reference, such as a collection nothing is removed from or a listener never cleaned up.
  4. Check the usual suspects first: repeated listeners, caches with no eviction, goroutines blocked forever, and closures that retain large objects.
  5. Apply the fix, then rerun the same snapshot comparison under the same load and confirm the growth has actually stopped.

A worked example: the listener nobody removed

Say a Node service sets up a fresh event listener on a shared emitter every time a particular request handler runs, intending it to be short-lived, but never explicitly removes it once that request finishes. Under steady traffic, listeners accumulate on that emitter indefinitely, each one holding onto whatever it captured in its closure, and memory climbs in a pattern that looks, on a graph, exactly like the slow, steady leak described above. A heap snapshot comparison points directly at the growing listener count on that one emitter, and the fix is a single explicit removal call the original code was missing, not a rewrite of anything else.

A worked example: the goroutine that never exits

Say a Go service spawns a goroutine per incoming request to wait on a result from a channel, but a code path exists where nothing ever writes to that channel under a specific error condition. That goroutine blocks forever, and its own stack and any variables it's holding onto never get freed, since the goroutine itself never actually exits. Enough requests hitting that error condition over time produces a slow, steady memory increase that traces back not to a data structure growing, but to leaked goroutines accumulating one at a time, each one small but permanent.

Executive Capability Standard

What Good Looks Like

A properly diagnosed memory leak is confirmed with real growth over time, traced to its specific retaining reference through heap snapshot comparison, and the fix is reconfirmed with the same comparison rather than assumed to have worked.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Learn your language's specific common leak patterns, whether that's event listeners and closures in Node or goroutines and channels in Go.
2. Do Manually:Manually take and compare two heap snapshots from a service you suspect is leaking, under representative load.
3. Delegate:Give one engineer ownership of investigating memory growth alerts so the same profiling sequence gets applied consistently.
4. Automate:Automate memory usage alerting on a sustained upward trend, not just a static threshold, so a slow leak gets flagged before it causes a crash.
5. Buy:Use a managed application performance monitoring tool with built-in memory profiling rather than manually taking heap snapshots for every investigation.

How to Get Started

Frequently Asked Questions

How long should we let memory grow before taking snapshots?

Long enough to distinguish a genuine leak from normal cache warm-up, which often means observing over hours rather than minutes. A short window can make normal, bounded growth look alarming, while a real unbounded leak will keep climbing well past the point where legitimate growth would have plateaued.

Are memory leaks more common in Node.js or in Go?

Neither language is inherently worse; both have well-known leak patterns specific to their own memory model, event listeners and closures in Node, goroutines and channels in Go. What matters more than the language is whether the team knows its own language's specific failure patterns and checks for them proactively.

Can a memory leak be fixed without heap profiling, just from reading the code?

Occasionally, for an obvious case like an unbounded cache with no eviction logic at all. For anything less obvious, code review alone tends to miss the actual retaining reference, since a leak is defined by what's still reachable in memory, which is exactly the thing a heap snapshot shows you directly and code reading has to infer.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides