Your Service Restarts Itself Every Night. That's Not Normal
A service that needs a nightly restart almost always has an undiagnosed memory leak, and the restart hides the bug rather than fixing it. The leak eventually catches up as an out-of-memory crash during a traffic spike, or as slower responses while garbage collection works harder against a growing heap.
Here's a practical answer to "where do I even start" for both Node and Go, since the tools and the usual culprits differ meaningfully between the two.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Confirm it's actually a leak before profiling anything
Memory that climbs and then plateaus at a higher baseline after a garbage collection cycle is normal behavior, not a leak, since most runtimes intentionally hold onto some reclaimed memory rather than returning it to the OS immediately. A real leak looks like memory that climbs steadily over hours or days without ever plateauing, even during periods of comparable, steady traffic. Graph heap usage over at least a full day, ideally several, before concluding you have a leak at all, since chasing a false positive wastes real profiling time on a problem that isn't there.
How do you find a memory leak in Node?
Take two heap snapshots using the built-in inspector, one shortly after startup and one after the service has been running long enough to show the climbing pattern, then compare object counts between them in a tool that supports snapshot diffing. The object type whose count grew the most between snapshots is almost always your leak's source. The most common Node culprits are event listeners added but never removed, closures captured in a cache that has no eviction policy, and promises held in an array that's appended to but never cleared.
How do you find a memory leak in Go?
Go's built-in pprof tooling gives you both heap profiling and goroutine profiling, and it's worth checking both, since a "memory leak" in Go is often actually a goroutine leak, background goroutines started but never terminated, each holding onto its own stack and any captured variables. Capture a heap profile and a goroutine count at two points in time the same way you would with Node's snapshots, and look specifically for a channel that's written to but never read from, which is the single most common way a Go goroutine ends up blocked forever instead of exiting cleanly.
The four causes that account for most real leaks
- Unbounded caches: a cache with no size limit or expiry policy grows forever under real traffic, even though it looked fine in every test with a handful of requests.
- Event listeners or subscriptions never cleaned up: especially common in long-lived connection handlers where a cleanup path exists but isn't reliably called on every exit path, including error paths.
- Closures capturing more than intended: a closure that captures an entire outer scope, including large objects the function never actually uses, keeps that whole scope alive for as long as the closure itself is referenced.
- Goroutines or async tasks blocked forever: waiting on a channel, queue, or promise that will never resolve, holding their entire stack and captured state indefinitely.
Building a leak test into your release process
The most reliable way to catch a leak before it reaches production is a load test that runs long enough, well past the point where transient allocation noise settles, with heap usage graphed throughout rather than sampled once at the end. A leak that only shows up after several hours of sustained traffic will never surface in a five-minute smoke test, no matter how thorough the test's request coverage otherwise is. Make this a gate before a release that touches connection handling, caching, or background task code specifically, since that's where leaks concentrate.
A common mistake is running the long load test against a staging environment with a tiny dataset, so caches never grow enough to reveal an unbounded entry count. For example, a cache keyed by user ID looks flat when a handful of test users cycle through it, then climbs all week in production. The fix is to make test traffic resemble production key diversity: replay anonymized request patterns, or generate many distinct users, tenants, and identifiers. Graph heap usage and, in Go, goroutine count together, and fail the release gate if either keeps rising after warm-up. That catches the leak that only diversity exposes.
What Good Looks Like
A team that handles memory leaks well confirms a genuine climbing pattern before profiling, uses heap and goroutine snapshots to find the actual source rather than guessing, and runs long-duration load tests as a release gate for code that touches caching or connection handling.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
Is a nightly scheduled restart an acceptable long-term fix for a memory leak?
It buys time, not a fix. The leak is still there, and a restart schedule tuned for normal traffic can fail exactly when you need the service most, during an unexpected spike that pushes memory past the limit before the scheduled restart arrives.
How long should a load test run to reliably catch a memory leak?
Long enough to see whether heap usage plateaus or keeps climbing after normal allocation noise settles, often several hours for a genuinely slow leak. A five- or ten-minute smoke test is enough to catch a crash but rarely enough to catch a gradual leak.
Are memory leaks more common in Node or in Go?
Both are equally capable of leaking, just through different mechanisms. Node leaks tend to concentrate in unbounded caches and uncleaned event listeners; Go leaks are more often goroutines blocked forever on a channel, which is a distinct failure mode worth checking for specifically.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Finding a Memory Leak Before It Pages You
A service's memory climbs for days until it gets killed and restarts, then climbs again. How to profile Node and Go to trace a leak to its real cause.
Finding a Memory Leak in Node or Go Before It Takes Down a Pod
How to use heap snapshots and pprof to find a real memory leak in Node.js or Go, and the common causes behind a slow, steady memory climb in production.
Profiling a Memory Leak in Node or Go Before It Pages You
A practical approach to finding a memory leak in Node.js or Go, including the tools to reach for first and the leak patterns specific to each.
Finding a Memory Leak Before It Finds Your Pager
A worked walkthrough of diagnosing a memory leak: heap snapshots in Node.js, pprof in Go, and capturing evidence before the process gets killed.
Tracking Down a Slow Memory Leak in Node or Go, Step by Step
A worked walkthrough of profiling and fixing a slow memory leak in a Node.js or Go service, from spotting the pattern to confirming the fix actually worked.
Chasing Down a Memory Leak in Node and Go Services
How to tell a real leak from normal garbage collection, take a useful heap snapshot, and stop shipping scheduled restarts as the fix.