Developer Productivity & Platform EngineeringPlaybook3 min readUpdated September 2026

Finding a Memory Leak Before It Pages You

A service's memory climbs steadily for days until it hits the container's limit, gets killed, restarts, and starts climbing again. Nobody profiled it, because a slow climb that takes days to matter is easy to write off as probably fine, right up until the restarts start happening during business hours.

A rising memory graph tells you a leak exists somewhere. It doesn't tell you where. Finding that takes an actual profile, not a guess.

Why Memory Leaks Are a Symptom, Not a Cause

A dashboard showing memory trending upward is a symptom, the same way a fever is a symptom. The actual cause is some specific object, closure, or goroutine that's being retained longer than intended, and finding it requires comparing what's actually accumulating in memory over time, not staring harder at the trend line. Treating the trend line itself as the diagnosis is how teams end up restarting a service on a schedule instead of fixing anything.

Profiling in Node: Heap Snapshots and What They Show

Take two heap snapshots several minutes apart under real load, then diff them to see which object types actually grew between the two. Common Node leak culprits show up clearly this way: event listeners registered but never removed, a cache with no eviction policy quietly growing without bound, or a closure holding a reference to something large well past when it should have been eligible for garbage collection.

Profiling in Go: pprof and the Goroutine Leak Variant

Go's memory leaks are frequently goroutine leaks in disguise: a goroutine blocked forever on a channel that nobody ever writes to or closes keeps alive everything that goroutine references, and the memory grows as a side effect of a stuck goroutine, not a direct memory bug. pprof's goroutine profile, showing a count that climbs steadily over time instead of settling, is the tell, distinct from what a heap profile alone would show you.

A Worked Example: Tracing One Leak to Its Actual Line

A heap diff shows a growing number of retained closures. Each one references a per-request context object that should have been garbage collected the moment its request finished. Tracing the retaining path back reveals the cause: an event listener registered on a shared emitter during every single request, but never removed once that request completed, so each request's context stays alive forever, held in place by a listener nobody ever cleaned up.

Getting a Profile From a Running Production Instance Safely

Most runtimes support capturing a heap snapshot or a pprof profile from a live process without a restart, though collection itself can add real CPU and memory overhead while it runs. Capture it during a period of moderate real traffic rather than your absolute peak, so the profile reflects genuine allocation patterns without risking a latency spike during your busiest window of the day.

A Checklist for Catching This Earlier Next Time

  • Alert on memory trending upward over a period of hours, not only on a single crossed absolute threshold that a normal traffic spike could also trip.
  • Add automatic heap or goroutine snapshotting on a sustained trend breach, so you capture a profile from the moment it's actually happening, not after a restart has already wiped the evidence.
  • Review any new code that adds an event listener, a cache, or a spawned goroutine for a clear removal or exit path before it merges, since that review catches most of these before they ship.

Deciding Between a Bigger Memory Limit and an Actual Fix

Raising the container's memory limit buys time and can be the right short-term move while a fix is being developed, but it doesn't address the underlying retention, it just delays the restart by however much extra headroom you bought. Treat a limit increase as a stopgap with an expiration, not a resolution, and track the actual profiling work as a separate, real task with an owner.

Watch for the same leak simply taking longer to matter after a limit increase rather than disappearing. If the growth rate stayed the same and only the ceiling moved, the underlying retaining code path is still there, quietly waiting for the next deploy cycle or traffic increase to bring the restarts back on the same schedule as before.

Executive Capability Standard

What Good Looks Like

Good memory leak response means a slow upward trend gets caught and profiled before it ever turns into a restart-crash-restart cycle a customer notices, with the actual retaining code path identified, not just patched with a bigger memory limit.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Take a heap snapshot or pprof profile of your highest-memory service under normal load to see what's actually being retained today.
2. Do Manually:Manually diff two snapshots taken minutes apart on your leakiest known service to trace one growing object type back to its source.
3. Delegate:Assign one engineer to own memory trend alerting and the profiling runbook so it doesn't depend on whoever happens to notice first.
4. Automate:Automate snapshot capture on a sustained memory trend breach so a profile exists from the actual moment of the leak, not after a restart erases it.
5. Buy:Bring in a performance specialist to profile a service where the leak has resisted a straightforward heap or goroutine diff.

How to Get Started

Frequently Asked Questions

How do you know if rising memory is actually a leak and not just normal growth?

Watch whether memory returns to a stable baseline after a period of load drops off, or keeps climbing regardless of traffic. Normal caching and connection pooling grow memory but should plateau; a genuine leak keeps climbing steadily even during quiet periods, which is the pattern worth pulling a heap snapshot to investigate.

What's a common memory leak pattern specific to Node?

An event listener registered on a shared emitter, often once per request, that never gets removed once its work is done. Every request's context and anything it closes over stays alive as long as that listener does, which in practice means forever, since nothing ever calls removeListener on it.

How is a typical Go leak usually different from a Node one?

Go leaks are often goroutine leaks rather than a direct object retention problem: a goroutine blocked forever on a channel nobody writes to or closes keeps everything it references alive as a side effect. A growing goroutine count in a pprof profile is usually a clearer early signal than the heap profile alone.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides