Enterprise DevSecOps & Automated CompliancePlaybook3 min readUpdated September 2026

Your Service Restarts Itself Every Night. That's Not Normal

A service that needs a nightly restart almost always has an undiagnosed memory leak, and the restart hides the bug rather than fixing it. The leak eventually catches up as an out-of-memory crash during a traffic spike, or as slower responses while garbage collection works harder against a growing heap.

Here's a practical answer to "where do I even start" for both Node and Go, since the tools and the usual culprits differ meaningfully between the two.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Confirm it's actually a leak before profiling anything

Memory that climbs and then plateaus at a higher baseline after a garbage collection cycle is normal behavior, not a leak, since most runtimes intentionally hold onto some reclaimed memory rather than returning it to the OS immediately. A real leak looks like memory that climbs steadily over hours or days without ever plateauing, even during periods of comparable, steady traffic. Graph heap usage over at least a full day, ideally several, before concluding you have a leak at all, since chasing a false positive wastes real profiling time on a problem that isn't there.

How do you find a memory leak in Node?

Take two heap snapshots using the built-in inspector, one shortly after startup and one after the service has been running long enough to show the climbing pattern, then compare object counts between them in a tool that supports snapshot diffing. The object type whose count grew the most between snapshots is almost always your leak's source. The most common Node culprits are event listeners added but never removed, closures captured in a cache that has no eviction policy, and promises held in an array that's appended to but never cleared.

How do you find a memory leak in Go?

Go's built-in pprof tooling gives you both heap profiling and goroutine profiling, and it's worth checking both, since a "memory leak" in Go is often actually a goroutine leak, background goroutines started but never terminated, each holding onto its own stack and any captured variables. Capture a heap profile and a goroutine count at two points in time the same way you would with Node's snapshots, and look specifically for a channel that's written to but never read from, which is the single most common way a Go goroutine ends up blocked forever instead of exiting cleanly.

The four causes that account for most real leaks

  • Unbounded caches: a cache with no size limit or expiry policy grows forever under real traffic, even though it looked fine in every test with a handful of requests.
  • Event listeners or subscriptions never cleaned up: especially common in long-lived connection handlers where a cleanup path exists but isn't reliably called on every exit path, including error paths.
  • Closures capturing more than intended: a closure that captures an entire outer scope, including large objects the function never actually uses, keeps that whole scope alive for as long as the closure itself is referenced.
  • Goroutines or async tasks blocked forever: waiting on a channel, queue, or promise that will never resolve, holding their entire stack and captured state indefinitely.

Building a leak test into your release process

The most reliable way to catch a leak before it reaches production is a load test that runs long enough, well past the point where transient allocation noise settles, with heap usage graphed throughout rather than sampled once at the end. A leak that only shows up after several hours of sustained traffic will never surface in a five-minute smoke test, no matter how thorough the test's request coverage otherwise is. Make this a gate before a release that touches connection handling, caching, or background task code specifically, since that's where leaks concentrate.

A common mistake is running the long load test against a staging environment with a tiny dataset, so caches never grow enough to reveal an unbounded entry count. For example, a cache keyed by user ID looks flat when a handful of test users cycle through it, then climbs all week in production. The fix is to make test traffic resemble production key diversity: replay anonymized request patterns, or generate many distinct users, tenants, and identifiers. Graph heap usage and, in Go, goroutine count together, and fail the release gate if either keeps rising after warm-up. That catches the leak that only diversity exposes.

Executive Capability Standard

What Good Looks Like

A team that handles memory leaks well confirms a genuine climbing pattern before profiling, uses heap and goroutine snapshots to find the actual source rather than guessing, and runs long-duration load tests as a release gate for code that touches caching or connection handling.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Graph heap usage for your longest-running service over several days to confirm whether it actually shows a climbing pattern or just normal plateauing behavior.
2. Do Manually:Manually take and diff two heap snapshots, or two pprof captures, around your most suspected service the next time memory climbs.
3. Delegate:Assign an engineer to own a long-duration load test with heap graphing as a required gate before releases touching caching or connection code.
4. Automate:Automate heap and goroutine count alerting so a climbing trend pages someone well before it becomes an out-of-memory crash.
5. Buy:Bring in a profiling or observability platform with built-in leak detection if manual snapshot diffing isn't catching leaks fast enough for your release cadence.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Tenable

Industry-leading platform for Enterprise DevSecOps: Memory Profiling in Node and Go.

Visit Tenable→
CrowdStrike

Alternative enterprise solution for scaling Enterprise DevSecOps: Memory Profiling in Node and Go.

Visit CrowdStrike→

Frequently Asked Questions

Is a nightly scheduled restart an acceptable long-term fix for a memory leak?

It buys time, not a fix. The leak is still there, and a restart schedule tuned for normal traffic can fail exactly when you need the service most, during an unexpected spike that pushes memory past the limit before the scheduled restart arrives.

How long should a load test run to reliably catch a memory leak?

Long enough to see whether heap usage plateaus or keeps climbing after normal allocation noise settles, often several hours for a genuinely slow leak. A five- or ten-minute smoke test is enough to catch a crash but rarely enough to catch a gradual leak.

Are memory leaks more common in Node or in Go?

Both are equally capable of leaking, just through different mechanisms. Node leaks tend to concentrate in unbounded caches and uncleaned event listeners; Go leaks are more often goroutines blocked forever on a channel, which is a distinct failure mode worth checking for specifically.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides