AI Model Serving & Inference OptimizationPlaybook3 min readUpdated September 2026

Finding a Memory Leak in Node or Go Before It Pages You

Memory that climbs during a load spike and comes back down afterward is normal garbage collection behavior, not a leak. Memory that climbs across restarts of your measurement window and never comes back down is the real thing, and it usually takes down a service with an out-of-memory kill at the worst possible time if nobody catches it first.

Both Node and Go ship the tools to find a leak without adding anything new to your stack, heap snapshots in one, pprof in the other. The skill is knowing what to look for once you've got them open.

How do you tell a memory leak from normal garbage collection?

Every garbage-collected or manually managed runtime has a normal memory sawtooth, usage rises as work happens and drops when it's reclaimed. A leak looks different: the baseline itself creeps upward over hours or days, so each sawtooth peak and trough sits a little higher than the last one.

Confirming that pattern before profiling anything saves time. Watch resident memory over a period long enough to see several cycles, not a single deploy's worth of data, since a short window can't tell a leak from a normal warm-up as caches fill.

Node: heap snapshots and what to look for in them

Take a heap snapshot under realistic load, let the service run for a while doing real work, then take a second one. Comparing the two in Chrome DevTools shows what grew between them, retained size and object counts per constructor, which points you at the specific type of object that's accumulating instead of leaving you to guess.

Look especially at counts that grow roughly in proportion to requests served, an array or map that gains one entry per request and never removes any is one of the most common patterns, and it's usually invisible until the snapshot comparison puts a number on it.

Go: pprof and the allocations that don't get collected

Go's runtime/pprof package exposes a heap profile you can capture and compare the same way, two profiles taken minutes apart under load, diffed to show what's still allocated in the second that shouldn't have accumulated since the first.

Goroutine leaks are a related but distinct problem worth checking alongside the heap profile: a goroutine blocked forever waiting on a channel nobody writes to or closes never gets collected, and it holds onto whatever it captured in its closure for as long as it exists. A steadily growing goroutine count in your metrics is often the earlier warning sign, showing up before the memory graph makes the leak obvious.

Common leak patterns in long-running services

A handful of patterns account for most leaks in services that run for days or weeks at a time. Event listeners registered on a long-lived object and never removed accumulate one per subscription. Closures that capture a large object by reference keep that object alive far longer than intended. Unbounded in-memory caches or maps that grow with usage and never evict. Timers or intervals started and never cleared when the thing they were tracking goes away.

None of these are exotic bugs, which is exactly why they're common: each one is a small, easy-to-miss omission, a removeListener call that isn't there, a cache with no eviction policy, that only becomes visible once the process has been running long enough for it to matter.

How do you confirm a memory leak is actually fixed?

A graph that looks flatter after a change is a good sign, not proof. Reproduce the original load pattern, run it for a sustained period, and watch the memory growth rate specifically, not just the absolute level at one point in time, comparing before and after under the same conditions.

If the growth rate hasn't measurably dropped, the fix addressed a symptom rather than the actual leak, or there's a second leak still running underneath the one that got fixed. Confirming under load before calling it done is what turns a plausible fix into a verified one.

A reliable way to confirm a fix:

  1. Reproduce the original load pattern that exposed the leak instead of judging the change during a quiet period.
  2. Run it for a sustained period, long enough to show several memory cycles rather than a single deploy's worth of data.
  3. Watch the growth rate of resident memory, not just the absolute level at one point in time.
  4. Compare the before and after growth rates under identical conditions, and treat an unchanged rate as a sign the fix addressed a symptom.
Executive Capability Standard

What Good Looks Like

Good here means a service that runs for weeks without its memory footprint trending upward, and when it does drift, you can point at the specific object or goroutine causing it instead of just restarting the process.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Learn to read a Node heap snapshot or a Go pprof heap profile well enough to tell what's actually growing between two points in time.
2. Do Manually:Take a snapshot or profile by hand under realistic load for a service you suspect is leaking, and compare it against one taken later in the same run.
3. Delegate:Have someone own a sustained load test that runs long enough for a slow leak to show up, since a short test won't catch one.
4. Automate:Add memory usage tracking to your monitoring so a service trending upward gets flagged before it gets killed for running out of memory in production.
5. Buy:For services where a leak is expensive to chase by hand, an APM tool with built-in memory profiling can shorten the time between noticing the trend and finding the cause.

How to Get Started

Frequently Asked Questions

Is rising memory always a leak?

No. Caches, connection pools, and normal garbage collection all produce memory usage that rises before it falls. The distinguishing signal is whether the baseline itself trends upward across several cycles, not whether usage ever goes up at all.

Do I need APM tooling to find a leak, or is the built-in profiler enough?

The built-in profiler, heap snapshots in Node or pprof in Go, is usually enough to find and confirm the cause once you suspect a leak. APM tooling mainly helps you notice the trend in production earlier, before it becomes an incident.

Can a memory leak in one service take down others?

Yes, in two common ways: it can eventually trigger an out-of-memory kill that drops requests mid-flight, and if it shares a host with other processes, its growing footprint can starve those other services of memory before it gets killed itself.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides