The page comes in at 3 a.m. The service restarted again, the orchestrator says it was killed for exceeding its memory limit, and the dashboards show the familiar sawtooth: memory climbs for two days, then drops to zero. Nobody changed anything recently, and the code review history is clean. That pattern is almost always a leak, and it is one of the more frustrating bugs to chase because nothing is visibly wrong until the process dies.
This guide walks through a workflow that works on real services: confirm that you actually have a leak, capture evidence at the right moments, read the evidence, fix the cause, and prove the fix held. It is written for a typical Node.js web service, but the approach applies to any garbage-collected runtime.
Photo by Kier in Sight Archives on Unsplash
Confirm That It Is Actually a Leak
Not every growing memory graph is a leak. Node.js keeps a large heap with generous headroom, and the garbage collector is lazy by design. A service that warms caches, loads a big configuration, or handles a burst of traffic will legitimately climb for a while and then flatten out. Restarting at the first sign of growth teaches you nothing.
A real leak has a specific shape. Memory trends upward across many hours, and it keeps trending upward even at the low points between garbage collections. If you only look at the raw resident set size, you will be fooled by the collector's timing. Look at the heap used after collection instead, which the runtime exposes through process.memoryUsage() and which most monitoring agents can scrape.
Before you go further, write down three numbers: the growth rate in megabytes per hour, the traffic level during that window, and the time since the last deploy. Growth that scales with request volume points at something per-request. Growth that is steady regardless of traffic points at a timer, a connection pool, or a background job. Those two stories lead to very different investigations.
Rule Out the Boring Causes First
Most leaks are not exotic. Before you open a profiler, spend ten minutes on the short list of things that cause the vast majority of them. This list is boring on purpose, because boring causes are the ones you will find.
Unbounded in-memory caches are the most common offender. A plain object or Map used as a cache with no size limit or expiry grows with every distinct key. If the keys include user IDs, URLs, or query strings, the cache is effectively a log of everything the service has ever seen. A bounded alternative such as an LRU cache gives you a ceiling, and the ceiling is the whole point.
Event listeners that never get removed are the second most common. Adding a listener inside a request handler, to a long-lived emitter, without a matching removal means each request leaves a closure behind. Node prints a warning once you pass ten listeners on a single emitter, and that warning in your logs is a free clue. The Node.js documentation on events explains the limit and how to raise it deliberately when it is legitimate.
Timers are the third. A setInterval created per connection and never cleared keeps the connection's whole object graph alive, because the timer callback references it. The same goes for promises that never settle, which pin their closures forever.
Capture Heap Snapshots at the Right Moments
When the boring causes are ruled out, you need evidence of what is actually on the heap. A heap snapshot is a full dump of every live object and what is holding it. The trick is not taking one snapshot. The trick is taking at least three, spaced so that growth is visible between them.
A reliable sequence looks like this. Start the service and let it warm up so one-time allocations are out of the way. Take the first snapshot as a baseline. Apply realistic load for a while, force a garbage collection, and take a second. Apply the same load again, force another collection, and take a third. Anything that grows between snapshots two and three, after the warmup noise has been excluded, is a leak candidate.
In Node you can write a snapshot on demand with the built-in v8.writeHeapSnapshot() call, or start the process with the inspector enabled and trigger one from Chrome DevTools over a remote connection. Be careful on production hosts. Taking a snapshot pauses the process and can temporarily double its memory use, so do it on one instance that has been pulled out of the load balancer rather than on the whole fleet.

Photo by panumas nikhomkhai on Pexels
Read the Snapshot Like a Detective
A raw snapshot can contain millions of objects, and the first time you open one it is overwhelming. You do not need to understand all of it. You need two views and one habit.
The first view is the comparison view. Load the third snapshot and compare it against the second. The tool lists object constructors with the number of objects added and removed, and the size delta. Sort by size delta and look at the top few rows. In a typical leak, one or two constructors dominate, and they are often plain Object, Array, Closure, or string, which tells you less than you would like but gives you a starting point.
The second view is retainers. Pick a suspicious object and read the retainer tree from the bottom up. This tells you why the object is still alive, which is the actual question. Follow the chain until you hit something that looks like it should not live forever: a module-level variable, a global cache, an emitter, or a timer. That link in the chain is your culprit.
The habit is to distrust the shallow size and trust the retained size. A small object that holds the only reference to a large graph is more important than a large object that nothing keeps alive. Sort by retained size whenever you are hunting for the thing that is pinning everything else.
Use Allocation Timelines When Snapshots Are Not Enough
Sometimes the snapshots show growth but the retainer chain is too generic to be useful, because everything ends at a shared utility function. When that happens, switch from snapshots to an allocation timeline. This records where each allocation happened, with a stack trace, over a window of time.
In DevTools this is the allocation instrumentation on timeline option. You start recording, run a controlled burst of load, stop, and look for blue bars that stay blue. A bar that turns gray was collected, and a bar that stays blue is still alive. Clicking a persistent bar shows the allocation stack, which points straight at the line of code that created the object.
The cost is overhead, because the runtime is recording every allocation, so use this on a staging replica or a single isolated instance. For a lighter-weight option, tools from the Clinic.js suite can profile a running Node process and produce a readable report without you hand-reading raw snapshots. They are useful as a first pass when you do not yet know where to look.
Fix the Cause, Not the Symptom
It is tempting to treat a leak with a scheduled restart. Restarting nightly makes the graph look healthy and hides the problem, and it is a legitimate stopgap while you investigate. It is not a fix, and it tends to become permanent. Meanwhile the leak still affects latency, because a heap that keeps growing makes garbage collection pauses longer and more frequent long before the process dies.
Once the retainer chain points at a cause, the fix is usually small. For an unbounded cache, add a size limit and a time-to-live. For a leaked listener, move the registration out of the request path or remove it in a finally block. For a timer, clear it when the owning object closes. For a closure that captures far more than it needs, restructure the code so the callback only references the few values it actually uses.
Weak references are worth knowing about but are rarely the right first tool. A WeakMap keyed by an object lets the entry disappear when the key does, which is ideal for attaching metadata to objects you do not own. Reaching for them to paper over an ownership problem usually just moves the confusion somewhere else.
"A leak is almost always a lifecycle bug in disguise. Somebody created a thing and nobody was assigned to destroy it, so the fix is to name the owner, not to tune the garbage collector." - Dennis Traina, founder of 137Foundry

Photo by Godfrey Atima on Pexels
Prove the Fix Held
A fix you have not verified is a guess with a commit message attached. Repeat the same measurement that convinced you there was a leak. Run the same load pattern, take the same sequence of snapshots, and confirm that the suspect constructor no longer grows between the second and third. Then deploy and watch the post-collection heap on the dashboard for at least as long as the original leak took to become visible.
If the original leak took two days to kill the process, a green graph after two hours proves nothing. Give it the full window, and keep the old graph next to the new one so the difference is obvious to everyone, including the person who reviews the incident later.
A regression test is possible for some leaks and worth writing when it is cheap. A test that creates and destroys a thousand instances of the leaking object, forces a collection with the --expose-gc flag, and asserts that the heap returned close to its starting size will catch the same mistake being reintroduced. It will not catch every leak, but it protects the one you just paid for.
Add Guardrails So the Next Leak Is Boring

Photo by panumas nikhomkhai on Pexels
The best time to find a leak is before it pages anyone. A few inexpensive guardrails change that. Export heap used after garbage collection as a metric and alert on its trend rather than its absolute value, so a slow climb triggers a ticket days before it triggers an outage. Set a sensible memory limit on the process with --max-old-space-size so the failure is predictable and happens below the container's hard limit, which gives the runtime a chance to log something useful on the way out.
Add listener-count and cache-size gauges for the components you know are long-lived. When a cache that should hold five hundred entries reports fifty thousand, you do not need a profiler to know where to look. Review the V8 blog occasionally as well, because changes to how the engine handles garbage collection affect what healthy memory behavior looks like, and a graph that looked alarming on an old runtime version can be perfectly normal on a new one.
Finally, make leak hunting a documented routine instead of tribal knowledge. Write down the snapshot sequence, who is allowed to pull an instance out of rotation, and where the files go afterward. The next engineer who gets the 3 a.m. page will thank you, even if they never know your name.
When to Bring in Outside Help
Some leaks resist the playbook. They live in a native addon, they appear only under a specific customer's traffic shape, or they are tangled up with a framework's internals in a way that makes the retainer chain unreadable. If you have spent two or three focused days and the snapshots keep ending in code you do not own, it can be cheaper to bring in someone who has seen the pattern before.
Our team at 137Foundry works on exactly this kind of production problem, and the web development service page covers how we approach performance and reliability work for existing applications. If the root cause turns out to be architectural, such as a design that holds per-user state in process memory, the data integration service is where moving that state into a proper store usually begins.
A Short Checklist to Keep Nearby
When the graph starts climbing again, work through the same sequence every time. Confirm the trend using post-collection heap, not raw resident memory. Check caches, listeners, and timers before you open a profiler. Take three snapshots with a warmup and forced collections between them. Compare the last two, sort by retained size, and read the retainer chain from the bottom up. Fix the lifecycle bug at its owner, then re-run the exact same measurement to prove it, and leave a metric behind so the next one announces itself early.
None of this is glamorous, and that is the point. Leaks feel mysterious only until you have a routine, and a routine turns a three-day mystery into an afternoon of careful reading. For related background, the about page explains how we think about maintainable systems, and the MDN Web Docs remain a dependable reference for the JavaScript language features, such as WeakMap and FinalizationRegistry, that come up along the way.