Distributed Systems & Enterprise ResiliencePlaybook3 min readUpdated September 2026

Chasing Down a Memory Leak in Node and Go Services

A service whose memory graph climbs steadily until it gets restarted looks alarming, but not every climbing graph is a leak. Knowing the difference, and knowing how to find the real ones when they happen, saves you from either ignoring a genuine problem or chasing one that was never there.

How do you tell a memory leak from normal growth?

Memory that grows and then plateaus at a new, higher level is often just the garbage collector's normal heap growth pattern, or a cache that has a real, if generous, ceiling. A true leak keeps climbing with no ceiling, restart after restart, until the process runs out of memory and crashes.

Watch the graph over several full garbage collection cycles, not just a few minutes, before deciding which pattern you're looking at. A single snapshot mid-cycle can look identical for both cases.

A service with a high deployment frequency, one shipping several deployments a day, rarely stays up long enough between restarts for a slow leak to become visible on its own, which is part of why these leaks tend to surface first as a vague, hard-to-reproduce complaint rather than an obvious climbing graph1.

How do you take a heap snapshot in Node?

Run the process with the inspector flag enabled and connect Chrome DevTools to take a heap snapshot, then take a second one several minutes later under normal traffic. Comparing the two, rather than reading either one in isolation, shows you what actually grew between them instead of just what's large in absolute terms.

Sort the comparison by the delta in retained size, not total size. A large object that was already there in the first snapshot isn't the leak; an object type whose count keeps climbing between snapshots is.

The snapshot comparison workflow is:

  1. Start the process with the inspector flag enabled and connect Chrome DevTools to it.
  2. Take a first heap snapshot once the service is running under normal traffic.
  3. Wait several minutes under that same traffic, then take a second snapshot.
  4. Compare the two and sort by growth in retained size, not absolute size, to see what actually grew.
  5. Follow the retainer chain of the fastest-growing object back to the code that holds the reference.

Go's approach: pprof and the allocations that don't get collected

Go's net/http/pprof package exposes a heap profile you can pull with the standard pprof tool, showing which call sites are responsible for currently live allocations. For Go specifically, a goroutine leak, one that never exits and keeps its stack and any references it holds alive, is a more common culprit than a raw memory allocation that never gets freed.

Pull the goroutine profile alongside the heap profile. A steadily climbing goroutine count is often the first sign of the problem, and it points you at exactly which function is spawning goroutines that never return.

A worked example: a listener that never unsubscribes

Say a component subscribes to an event emitter on every render but only unsubscribes when it unmounts cleanly, and some code path skips the unmount. Each subscription holds a reference to the component and everything it closes over, so the heap snapshot comparison shows the emitter's internal listener array growing between snapshots, and following the retainer chain from there leads straight back to the missing unsubscribe call.

The same pattern shows up as a cache without eviction, or a Go goroutine blocked forever reading from a channel nobody closes: something is holding a reference that was supposed to be temporary.

The mistake: shipping a restart-on-schedule cron as the fix

A scheduled restart keeps a leaking process from ever running out of memory, which makes the symptom disappear from the dashboard without fixing anything. It also hides the moment the leak gets worse, since the process never runs long enough under the new, faster growth rate to show it.

Treat a scheduled restart as a stopgap you've deliberately chosen while you profile the real cause, with a ticket attached, not as a permanent operational pattern nobody revisits.

Reproducing it outside production without waiting hours

A leak that takes six hours to become visible under real traffic often takes minutes to show under a synthetic load test that replays the same request pattern at a much higher rate. Compressing the timeline like this turns an incident you can only investigate live into something you can profile deliberately, snapshot by snapshot, without the pressure of a production outage running in the background.

This works best when the load test exercises the same code paths a real leak was traced to, not just raw request volume against an endpoint that turned out to be unrelated.

Executive Capability Standard

What Good Looks Like

Good leak-hunting discipline means the team can tell a real leak from normal GC behavior by watching multiple collection cycles, and can pull a heap snapshot or profile on a live process without guessing.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Learn to read a heap snapshot comparison in Node or a pprof profile in Go well enough to identify what grew between two points in time.
2. Do Manually:Take two snapshots or profiles by hand a few minutes apart on a service you suspect is leaking, and trace the retainer chain for whatever grew.
3. Delegate:Give one engineer ownership of investigating any service whose memory graph climbs without plateauing, rather than letting it default to whoever's on call.
4. Automate:Set up alerting on memory growth rate, not just absolute usage, so a real leak gets flagged before it reaches a crash threshold.
5. Buy:Bring in a performance engineering consultant for a one-time deep profile if a leak has resisted your team's own investigation for more than a few days.

How to Get Started

Frequently Asked Questions

How do I tell a real leak from normal GC behavior?

Watch memory over several full garbage collection cycles rather than a short window. Normal behavior grows and then plateaus at a stable level; a real leak keeps climbing with no ceiling across every cycle, eventually leading to a crash if left unaddressed.

Can I profile a leak in production without restarting the service?

Yes. Node's inspector protocol and Go's pprof endpoint are both designed to be attached to a running process without restarting it. Take snapshots or profiles a few minutes apart under normal traffic rather than trying to reproduce the issue in a separate environment first.

Why do Go services leak goroutines more often than memory?

A goroutine that blocks forever, often waiting on a channel nobody closes, keeps its entire stack and every reference it holds alive indefinitely. That pattern is easy to introduce accidentally and shows up as a steadily climbing goroutine count well before it shows up as an obvious memory problem.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.

Related Guides