• 1 min read

Hunting Down Silent Memory Leaks in High-Throughput Node.js Microservices

A post-mortem on debugging an elusive 12 MB/hour heap leak under 4,500 req/sec load. V8 heap snapshots, weak references, closure traps, and event listener leak detection.

Hunting Down Silent Memory Leaks in High-Throughput Node.js Microservices

Few things in backend engineering are as demoralizing as watching a Kubernetes pod graph show a steady, unyielding 45-degree upward slope in resident memory.

Last quarter, our real-time notification gateway (a Node.js service handling approximately 4,500 WebSocket connections and HTTP webhook dispatches per second) started dying with SIGABRT (Out of Memory) every 72 hours. The container would boot at 120 MB of RSS (Resident Set Size), climb steadily by roughly 12 MB per hour, hit the 1.4 GB container limit, and get unceremoniously terminated by the Linux OOM killer.

The immediate band-aid was automated container restarts. But restarts drop active WebSockets and degrade latency. We needed the root cause. Here is the post-mortem of how we captured the leak in production without taking the cluster offline.


The V8 Garbage Collection Mental Model

Node.js delegates memory management to Google’s V8 engine. V8 divides the heap into two primary generations:

  1. Young Generation (Nursery & Intermediate): Small (16–64 MB). Fast “Scavenge” garbage collection cycles run every few hundred milliseconds. Most objects (request payloads, short-lived promises) die here.
  2. Old Generation: Large. Objects that survive multiple scavenge cycles are promoted here. Garbage collection here uses the Mark-Sweep-Compact algorithm.
V8 HEAP ALLOCATION TIMELINE:
[ Incoming Request Object ] ──> [ Young Generation (Scavenge) ]
                                          │
                               (Survived 2 cycles)
                                          │
                                          ▼
                                [ Old Generation Heap ]
                                          │
               ┌──────────────────────────┴──────────────────────────┐
               ▼                                                     ▼
     [ Legitimate Cache ]                                   [ LEAKING RETAINER ]
     (Has TTL & eviction)                                (Accidental global root / closure)

A leak in Node.js occurs when an object that is no longer needed remains reachable from a Root (the global object, an active timer, or an uncleared event listener closure). If a root holds a reference, V8 can never free it.


Step 1: Generating Heap Snapshots in Production Without Downtime

You cannot attach Chrome DevTools or a debugger directly to a production Kubernetes pod under heavy load without introducing severe latency penalties.

Instead, we used Node.js’s built-in v8 module to trigger a non-blocking heap snapshot when RSS crossed a specified threshold:

// diagnostic-heap.ts
import v8 from 'node:v8';
import fs from 'node:fs';

let snapshotCount = 0;

export function captureHeapSnapshot(tag: string) {
  if (snapshotCount >= 3) return; // Prevent disk flooding
  snapshotCount++;

  const fileName = `/tmp/heap-${tag}-${Date.now()}.heapsnapshot`;
  const stream = v8.getHeapSnapshot();
  const fileStream = fs.createWriteStream(fileName);

  stream.pipe(fileStream);
  fileStream.on('finish', () => {
    console.log(`[Diagnostic] Dumped heap snapshot to ${fileName}`);
  });
}

We captured three snapshots:

  • Baseline Snapshot: 10 minutes after pod initialization (~140 MB).
  • Snapshot 2: 12 hours later (~280 MB).
  • Snapshot 3: 24 hours later (~420 MB).

Step 2: Comparing Snapshots in Chrome DevTools

We pulled the .heapsnapshot files locally and loaded them into Chrome DevTools (Memory tab -> Load).

Using the Comparison view between Snapshot 1 and Snapshot 3, we sorted by # Delta and Size Delta. Immediately, three culprits jumped to the top of the retainer table:

Heap Comparison Delta (Snapshot 1 -> Snapshot 3):
Constructor            # Alloc    # Freed    # Delta    Size Delta
------------------------------------------------------------------
(closure)              482,910    12,100     +470,810   +38.4 MB
EventEmitter           142,000    1,200      +140,800   +18.2 MB
SocketContext          140,800    0          +140,800   +56.3 MB

Over 140,000 instances of SocketContext were alive, even though our active connection gauge reported only 4,500 active sockets.


The Culprit: The Innocent Event Listener Trap

Looking at the Retainers Tree for SocketContext, we traced the retention path up to the global singleton metrics client.

Here was the offending code in our telemetry middleware:

// THE BUGGY CODE:
class MetricCollector {
  private eventEmitter = new EventEmitter();

  trackSocket(socket: WebSocket, context: SocketContext) {
    // ⚠️ FATAL LEAK:
    // This closure captures `context` in its lexical scope!
    // The global `eventEmitter` holds onto this callback FOREVER
    // because `removeListener` was never invoked on socket disconnect!
    this.eventEmitter.on('telemetry_tick', () => {
      if (socket.readyState === WebSocket.OPEN) {
        socket.send(JSON.stringify(context.getStats()));
      }
    });
  }
}

Every time a mobile user reconnected over a flaky cellular connection, a new SocketContext was allocated. The anonymous arrow function registered an event listener on the long-lived singleton MetricCollector. Even when the TCP socket disconnected, the closure kept the entire SocketContext, user profile, and memory buffer pinned to the global root.


The Fix: WeakRefs & Explicit Cleanup

We solved the issue with two surgical changes:

  1. Explicit Teardown on Disconnect:
class MetricCollector {
  trackSocket(socket: WebSocket, context: SocketContext) {
    const onTick = () => {
      if (socket.readyState === WebSocket.OPEN) {
        socket.send(JSON.stringify(context.getStats()));
      }
    };

    this.eventEmitter.on('telemetry_tick', onTick);

    // CRITICAL: Guaranteed cleanup when the socket terminates
    socket.once('close', () => {
      this.eventEmitter.off('telemetry_tick', onTick);
    });
  }
}
  1. Using WeakRef for Loose Coupling: For components where strict lifecycle teardown is impossible, use ES2021 WeakRef and FinalizationRegistry:
// WeakRef allows V8 to collect the target object if no other strong pointers exist
const weakContext = new WeakRef(context);

this.eventEmitter.on('telemetry_tick', () => {
  const ctx = weakContext.deref();
  if (ctx) {
    // Target is still alive
    ctx.recordTick();
  } else {
    // Target was garbage collected; unsubscribe safely!
  }
});

The Impact: Flat 140 MB Memory Line

Following the deployment, memory stabilized at a flat 138 MB RSS, running continuously for 30 days without a single restart.

Cluster Telemetry Post-Fix:
├── Uptime: 30 days uninterrupted
├── Average RSS: 138 MB (previously climbing to 1,400 MB)
├── p99 GC Pause Duration: 3.8 ms (previously spiking to 84 ms)
└── Total Pod Restarts: 0

Never leave event listener registration to chance in Node.js. If you write .on(), your code must always have a verified matching .off().

P
Written by Piush KS · Mobile & Full-Stack Developer

Engineer specializing in Android, iOS, Flutter, and web applications. Focuses on scalable software architectures, performance optimization, and developer productivity tools.

🛠️ Explore free developer tools: format JSON, generate app icons, test code, and pick colors.
Explore All Tools →
Separator