In the previous three lessons we have quoted a lot of numbers: 641 requests per second, then 3,105, then 11,840; 99th percentiles of 204, 54 and 11 milliseconds; an event loop delay that went from 4,176 ms to 11 ms. Those numbers did not come from intuition: they came from measuring first, changing one thing and measuring again. That practice, which we have been applying without naming it, is the content of this lesson.
You will not find a list of tricks here, but a method, a catalog of instruments (load tests, CPU profiles, flamegraphs, heap snapshots) and a catalog of Node's typical bottlenecks with their fix, all anchored to what you already know from the course.
Contents
- The method: measure, locate, fix, measure again
- What to measure: latency, throughput, errors, saturation
- Load testing with
autocannon - CPU profiling and flamegraphs
- Memory profiling and leaks
- Event loop delay as a health metric
- A catalog of typical Node bottlenecks
- Optimizations in the HTTP layer
- The database: the connection pool
- V8 without falling into useless micro-optimization
- The correct order of the levers
- The method: measure, locate, fix, measure again
The cycle has four steps and none of them is optional: measure the complete system under a realistic load and write down the baseline; locate the bottleneck with a profile, not with a hunch; fix only that, one change at a time; and measure again with the same test, comparing against the baseline.
One change per iteration is non-negotiable. If you touch three things and the system improves by 30%, you do not know which of the three worked, nor whether one of them made things worse and the other two made up for it.
Donald Knuth's quote gets repeated a lot and almost always mutilated. The full sentence is: "we should forget about small efficiencies, say about 97% of the time: premature optimization is the root of all evil. Yet we should not pass up our opportunities in that critical 3%." Both extremes are mistakes: rewriting a loop to save microseconds in code that runs once a day, and also ignoring that your main query does a sequential scan over a million-row table. The difference between the two is not philosophical: it is that you measured one and not the other. Practical corollary: optimizing without measuring is guessing, and a programmer's intuition about where the time goes is notoriously bad. In Escena Viva anyone would have bet the bottleneck was generating the PDF; the profile proved that in the catalog 62% of the time went on serializing JSON and on a query with no index.
- What to measure: latency, throughput, errors, saturation
Four signals, known as the golden signals:
| Signal | What it is | How it is measured in Escena Viva |
|---|---|---|
| Latency | Response time | Percentiles per route, separating 2xx from 5xx |
| Throughput | Requests served per second | autocannon, and counters in production |
| Error rate | Proportion of 5xx and timeouts | Structured logging (M6) by error code |
| Saturation | How full the limiting resource is | CPU, memory, pool connections, loop delay |
On latency there is one point worth insisting on, and it is that the average lies:
| Statistic | Value in the catalog | What it means |
|---|---|---|
| Average / p50 | 78 ms / 72 ms | No individual user experiences the average |
| p95 | 161 ms | 1 in every 20 requests is worse |
| p99 | 204 ms | 1 in 100. With 12 requests per page, 1 in every 8 visits sees this |
| p99.9 / maximum | 890 ms | The long tail: what triggers the complaints |
If ten requests take 10 ms and one takes 5,000 ms, the average is 463 ms: a number that describes neither the fast ones nor the slow one. And the p99 arithmetic is frightening once translated into experience: a page that makes 12 API calls has a 1 − 0.99¹² ≈ 11% probability of hitting at least one 99th-percentile request. Optimize the tail, not the average. These are Escena Viva's service objectives, agreed with the business:
| Route | p99 target | Availability | Rationale |
|---|---|---|---|
GET /api/events |
< 150 ms | 99.9% | It is the first screen; if it is slow, nothing sells |
GET /api/events/:id/sessions |
< 200 ms | 99.9% | The step before the purchase |
POST /api/orders |
< 800 ms | 99.95% | Transactional; more latency is tolerated, failures are not |
POST /auth/login |
< 400 ms | 99.9% | bcrypt with cost 12 has a floor of its own |
That last case shows that not everything slow is a problem: bcrypt with cost 12 takes ~250 ms on purpose, because that slowness is the defense against brute force (M8), and optimizing it would be a security regression.
- Load testing with
autocannon
autocannonAn honest load test meets four conditions. Warm-up: the first seconds do not count, because V8 needs to run the code several times before the optimizing compiler (TurboFan) does its job, and the caches are cold. Sufficient duration: 30-60 seconds minimum, since ten seconds can hide a leak or the effect of a major garbage collection. Realistic concurrency: not 5,000 connections if your real peak is 200. And a realistic scenario: do not hammer a single trivial route, reproduce the user's journey.
# Warm-up (10 s that are discarded) and measurement (60 s, 100 connections).
npx autocannon -c 50 -d 10 -w 4 http://localhost:3000/api/events > /dev/null
npx autocannon -c 100 -d 60 -w 4 --latency http://localhost:3000/api/eventsA realistic purchase scenario is defined programmatically:
'use strict';
const autocannon = require('autocannon');
// Festival de Jazz opening-night scenario: browse the catalog, open the
// event page, check the sessions and, in a fraction of cases, buy.
const instance = autocannon({
url: 'http://localhost:3000',
connections: 100,
duration: 60,
warmup: { connections: 10, duration: 10 },
requests: [
{ method: 'GET', path: '/api/events' },
{ method: 'GET', path: '/api/events/evt-003' },
{ method: 'GET', path: '/api/events/evt-003/sessions' },
{ method: 'POST', path: '/api/orders',
headers: { 'content-type': 'application/json', authorization: 'Bearer <token>' },
body: JSON.stringify({ sessionId: 'ses-003-1', quantity: 2 }) },
],
});
autocannon.track(instance, { renderProgressBar: true });
instance.on('done', (r) => console.log(`p99: ${r.latency.p99} ms, non-2xx: ${r.non2xx}`));And this is how the output is read:
| Field | What to look at |
|---|---|
Latency p99 |
The number that governs your service objectives |
Req/Sec |
Throughput; always compare it with the baseline |
non2xx |
If it is not 0, the test is worthless: you are measuring how fast you fail |
errors / timeouts |
Rejected or timed-out connections: you have exceeded capacity |
Bytes/Sec |
If you approach the bandwidth, the bottleneck is the network, not your code |
k6 is the alternative when you need stateful scenarios (login, cookies, chained sessions), thresholds that fail the test in CI, and distributed load: autocannon is perfect for the fast "measure, change, measure" loop, and k6 for the performance regression test in continuous integration (M11).
An important, non-negotiable warning. A load test is, technically, a denial-of-service attack. Run it only against systems that are yours and in test environments: never against production without prior coordination (and there, prefer mirrored traffic or limited synthetic load), and never against a third-party API, neither the M8 currency service, nor the payment gateway, nor a public service. It is illegal in many jurisdictions and it is always professionally disrespectful.
- CPU profiling and flamegraphs
The load test tells you that something is slow; the profile tells you where.
Option 1: V8's built-in profiler. node --prof src/server.js generates an isolate-*.log with program-counter samples; you run the load, stop the server and node --prof-process isolate-*.log > profile.txt turns it into a readable report that starts with the summary that matters most:
[Summary]: ticks total nonlib name
4821 61.9% 72.4% JavaScript
1901 24.4% GC
502 6.4% Shared libraries
[Bottom up (heavy) profile]:
1204 15.5% JSON.stringify
890 11.4% LazyCompile: *serializeCatalog src/repositories/events.js:88
731 9.4% LazyCompile: *calculateAvailableCapacity src/domain/session.js:41Two immediate readings: JSON.stringify takes 15.5% and garbage collection 24.4%, which suggests we are creating too many temporary objects. The * in front of the name means V8 optimized that function; a ~ would mean it ran unoptimized, and that is sometimes the clue to a hidden-classes problem.
Option 2: the Chrome DevTools inspector. With node --inspect src/server.js and chrome://inspect you get a CPU profile with a flame chart, heap snapshots and comparisons between them in a single interface. On a remote server you use --inspect=0.0.0.0:9229 with an SSH tunnel, never exposed to the internet: the inspector lets anyone run arbitrary code in your process.
Option 3: 0x and flamegraphs.
npx 0x --output-dir profiles -- node src/server.js
# Run the load, stop with Ctrl+C and open the generated HTML.How to read a flamegraph, which is where almost everybody gets it wrong:
- The horizontal axis is NOT time, it is the proportion of samples; functions are ordered alphabetically, not chronologically.
- The width of a box is the total time in that function, including whatever it called. It is the only thing worth looking for: wide boxes.
- The height is the stack depth. A very tall, narrow tower is just code with many layers: it is not a problem.
- A wide box with a flat plateau on top means the time is spent in the function itself, not in its calls: there is your expensive function.
- Color means nothing by default (except in
0x, which highlights unoptimized functions).
In Escena Viva's flamegraph before the cache, a wide plateau over getCatalog → mapEvents → JSON.stringify took up nearly two thirds of the width; after lesson 10-03 that plateau disappears and the width is taken by libuv's idle loop, which is exactly what you want to see.
- Memory profiling and leaks
A memory leak in Node does not show up as an error, but as degradation: the process grows, garbage collection works harder and harder (and its pauses are time stolen from your event loop), and eventually the process dies with JavaScript heap out of memory. The diagnostic method with heap snapshots: start the process and let it settle; take snapshot A (DevTools → Memory → Heap snapshot, or v8.writeHeapSnapshot()); apply sustained load for several minutes; force a collection and take snapshot B; and in DevTools choose the Comparison view between B and A.
The decisive column is Delta: how many extra objects of each type there are. If Array or a constructor of yours grows monotonically snapshot after snapshot, there is your leak; then Retainers tells you who is stopping the collector from taking them away.
'use strict';
// v8.writeHeapSnapshot(path) takes the snapshot on demand, in production
// too (carefully: it pauses the process and generates hundreds of MB).
// And this is the cheap, continuous watch over memory usage.
function recordMemoryUsage(logger) {
const usage = process.memoryUsage();
const mb = (bytes) => Math.round(bytes / 1048576);
logger.info({ rssMb: mb(usage.rss), heapUsedMb: mb(usage.heapUsed),
heapTotalMb: mb(usage.heapTotal), externalMb: mb(usage.external) });
}
module.exports = { recordMemoryUsage };The two leaks you will see in real life, and both appeared in Escena Viva. The first, an unbounded cache: before lesson 10-03, the currency cache was an in-memory Map which with fixed pairs did not grow, but which, had the key included the amount (EUR:USD:4500), would have grown forever. Every in-memory cache needs a ceiling with LRU eviction (lru-cache) or it needs to move to Redis, which expires on its own. The second, listeners that are never removed: our SalesManager is an EventEmitter (M2), and if every request registers a listener that nobody removes, the emitter accumulates references to the closures and, with them, to everything they capture.
'use strict';
// BAD: the listener outlives the request and retains 'res' forever, and
// with it the socket, the headers and the body.
const badController = (manager) => (req, res) => {
manager.on('sale-recorded', (sale) => res.write(JSON.stringify(sale)));
};
// GOOD: it is registered, and removed when the request ends. The 'close'
// event covers both a normal ending and the client disconnecting.
const goodController = (manager) => (req, res) => {
const onSale = (sale) => res.write(JSON.stringify(sale));
manager.on('sale-recorded', onSale);
res.on('close', () => manager.off('sale-recorded', onSale));
};
module.exports = { goodController };The MaxListenersExceededWarning: Possible EventEmitter memory leak detected warning is Node literally telling you that you have this leak. Never silence it by raising setMaxListeners: fix the removal.
- Event loop delay as a health metric
If you had to pick one single metric to know whether a Node process is healthy, it would be this one: CPU at 90% can be normal, but a loop delay of 800 ms never is.
'use strict';
const { monitorEventLoopDelay } = require('node:perf_hooks');
// resolution: how often a sample is taken, in ms. 10 ms is a good point.
function createLoopWatcher({ logger, intervalMs = 30_000, p99ThresholdMs = 100 }) {
const histogram = monitorEventLoopDelay({ resolution: 10 });
histogram.enable();
const ms = (value) => Number((value / 1e6).toFixed(2));
const timer = setInterval(() => {
const p99 = ms(histogram.percentile(99));
logger[p99 > p99ThresholdMs ? 'warn' : 'info']({
metric: 'event-loop-delay', p99Ms: p99,
p50Ms: ms(histogram.percentile(50)), maxMs: ms(histogram.max),
});
histogram.reset(); // sliding window, not cumulative
}, intervalMs);
timer.unref();
return { stop: () => { clearInterval(timer); histogram.disable(); } };
}
module.exports = { createLoopWatcher };How to interpret the values:
| p99 delay | Diagnosis |
|---|---|
| < 10 ms | Healthy |
| 10-50 ms | High but manageable load |
| 50-200 ms | There is synchronous work: investigate with a CPU profile |
| > 200 ms | Something blocks the loop. It is the symptom from lesson 10-02 |
| > 1,000 ms | The process is effectively down even if it answers a ping |
It is also the best value for a health check: a process with a jammed loop answers 200 OK to a trivial /health while leaving every user stranded, and only a /health that looks at the loop delay detects that state.
- A catalog of typical Node bottlenecks
| Bottleneck | Symptom | Fix | Reference |
|---|---|---|---|
Synchronous I/O on the hot path (readFileSync, existsSync per request) |
High loop delay, profile with fs in the stack |
The asynchronous version, or read once at startup | M3 |
| A giant JSON serialized on every response | Wide JSON.stringify in the flamegraph, high GC |
Paginate, select fields, cache the already serialized JSON | 10-03, 10-05 |
| N+1 query | Many short, identical queries per request | include/JOIN, or DataLoader (10-06) |
M7 |
| Missing indexes | EXPLAIN ANALYZE shows a Seq Scan |
Create the index and verify it is used | M7 |
Avoidable sequential await |
Latency equal to the sum of the waits | Promise.all on independent operations |
M2 |
| Synchronous logging | Loop delay rises with the log volume | An asynchronous, buffered logger (pino), no console.log on the hot path |
M6, M11 |
| ReDoS (catastrophic backtracking) | One specific request freezes the process | Regex with no nested ambiguity, or prior length/zod validation |
M8 |
| Cacheable data regenerated | High CPU and database load to return the same thing | Cache-aside with a TTL | 10-03 |
| CPU work on the main thread | Loop delay in seconds | Threads or a queue | 10-02 |
Two deserve elaboration. The first, sequential await where Promise.all fitted, is the most common and cheapest performance bug to fix:
'use strict';
// BAD: three 'await' in series over the three repositories.
// Latency = 40 + 35 + 30 = 105 ms.
//
// GOOD: all three are independent. Latency = max(40, 35, 30) = 40 ms.
async function loadEventScreen(eventId, repos) {
const [event, sessions, venue] = await Promise.all([
repos.events.getEventById(eventId),
repos.sessions.listByEvent(eventId),
repos.venues.getByEvent(eventId),
]);
return { event, sessions, venue };
}The condition is independence: if the second query needs the first one's result, the sequential await is mandatory. And for operations where one failure must not take the rest down, use Promise.allSettled. The second is ReDoS: regular expressions with catastrophic backtracking, the only bottleneck on this list that is also a security vulnerability (OWASP, M8), because a single 40-character request can freeze the process for minutes.
'use strict';
// DANGEROUS: (\s*\w+)+ has nested ambiguity. With an input such as
// 'aaaaaaaaaaaaaaaaaaaaaaaaaaaaaa!' the engine tries an exponential number
// of combinations before giving up, and the whole process freezes.
const DANGEROUS_REGEX = /^(\s*\w+)+$/;
// SAFE: no ambiguity, a single quantifier over disjoint classes.
const SAFE_REGEX = /^[\w\s]+$/;
// And on top of that: limit the length BEFORE applying the regex.
const validateEventTitle = (title) =>
typeof title === 'string' && title.length <= 120 && SAFE_REGEX.test(title);
module.exports = { validateEventTitle };Defense in layers: validate length and shape with zod (M6) before any regex, avoid nested quantifiers, and distrust expressions copied from the internet to validate e-mails or URLs.
- Optimizations in the HTTP layer
Before touching your code, check that the transport layer is not giving work away:
| Lever | How | Typical gain |
|---|---|---|
| Compression | compression (already installed), a 1 KB threshold |
60-80% fewer bytes on JSON |
| ETag + 304 | We did it by hand in M4; Express ships it | An empty response on repeated reads |
Cache-Control |
public, max-age=60 on the catalog |
Requests that never reach your server |
Outbound keep-alive |
An agent with keepAlive: true in fetch/HTTP |
Saves the TCP+TLS handshake on every call |
| Mandatory pagination | A maximum limit on the server | Prevents megabyte-sized responses |
Compression is not free: it burns CPU. With small responses (< 1 KB) it costs more than it saves, and on a server with saturated CPU it can be counterproductive; in production it is usually delegated to the reverse proxy (M11). About outbound connections, every call to the M8 currency service without keep-alive pays a TCP handshake and a TLS negotiation: some 40-80 ms of pure protocol.
'use strict';
const { Agent, setGlobalDispatcher } = require('undici');
// Reuses connections for every outbound call made with fetch.
setGlobalDispatcher(new Agent({
connections: 32, // maximum connections per origin
keepAliveTimeout: 30_000,
keepAliveMaxTimeout: 120_000,
}));
- The database: the connection pool
Opening a connection to PostgreSQL costs between 5 and 50 ms (authentication, TLS, a process on the server). That is why Sequelize keeps a pool: a set of open connections that are lent out and returned.
'use strict';
const os = require('node:os');
const { Sequelize } = require('sequelize');
const { configuration } = require('../config/index.js');
// PostgreSQL's global budget: max_connections (100). Split it into
// cluster workers x pool size, plus room for queue consumers,
// migrations and administration.
const WORKERS = Math.max(1, os.availableParallelism() - 1); // 7
const MAX_PER_PROCESS = Math.max(2, Math.floor(80 / (WORKERS + 2)));
const sequelize = new Sequelize(configuration.postgresUrl, {
logging: false,
pool: {
max: MAX_PER_PROCESS, // 8 with 7 workers
min: 2, // warm connections always ready
acquire: 10_000, // maximum wait for a free connection
idle: 10_000, // closes idle ones after this time
},
});
module.exports = { sequelize };How it is sized, in order: find out the server's max_connections (100 by default in PostgreSQL); reserve room for migrations, queue consumers and administration; divide the rest by the number of processes that connect; and remember that bigger is not better, because a pool of 100 against 4 database cores only adds contention (there is a classic formula suggesting cores × 2 + disk spindles as a starting point).
What happens when the pool runs out: requests wait up to acquire and then fail with SequelizeConnectionAcquireTimeoutError, with a characteristic and misleading symptom —very high latency with no high CPU, because everybody is waiting. If you see it, first check whether slow queries are holding connections or whether there are open transactions that nobody closes.
- V8 without falling into useless micro-optimization
V8 optimizes your code through hidden classes: objects created with the same structure share an internal description and access to their properties compiles down to a fixed offset, extremely fast. If you change the shape, V8 creates another one and access becomes polymorphic or megamorphic, much slower.
'use strict';
// BAD: the object's shape changes as it goes (three hidden classes).
function createTicketBad(code, priceCents) {
const ticket = {};
ticket.code = code;
ticket.priceCents = priceCents;
if (priceCents === 0) ticket.isComplimentary = true; // only sometimes
return ticket;
}
// GOOD: a single shape, always the same fields and in the same order.
const createTicket = (code, priceCents) => ({
code, priceCents, isComplimentary: priceCents === 0,
});
// 'delete' degrades the object to dictionary mode and cancels the
// optimization: instead of 'delete ticket.seat', we reassign.
const clearSeat = (ticket) => ({ ...ticket, seat: null });
module.exports = { createTicket, clearSeat };And about memory: the old heap has a limit (typically ~2 GB on 64-bit, or less in small containers) that is adjusted with --max-old-space-size=2048. Two indispensable nuances: raising it does not fix a leak, it only delays the death and lengthens the garbage collection pauses; and in a container you must set it below the container's memory limit, because otherwise the system kills the process (the OOM killer) before V8 even notices (a detail we will pick up again in M11).
The general warning: these micro-optimizations matter in code that runs millions of times. In an HTTP controller that runs 3,000 times per second and spends 95% of its time waiting on the database, reordering properties changes nothing. Apply them only if the profile points at that function.
- The correct order of the levers
When something is slow there is an order by return on investment:
| Order | Lever | Typical gain | Cost |
|---|---|---|---|
| 1 | Algorithm and data structure | 10× to 1000× | Thinking; sometimes rewriting |
| 2 | Query and schema (index, N+1, SELECT of fewer columns) |
10× to 100× | Low; a migration |
| 3 | Cache | 5× to 50× | Medium: invalidation (10-03) |
| 4 | Concurrency (cluster, threads, queues) | 2× to 8× | Medium-high: shared state |
| 5 | Code micro-optimization | 1.05× to 1.5× | High in readability |
| 6 | Hardware | Linear, with a monthly bill | Recurring money |
You work it from top to bottom. Buying a machine twice as big to paper over a query with no index means paying twice as much every month for a problem that a CREATE INDEX would have fixed. And that brings us to Escena Viva's actual path through this module: first concurrency (10-01, 10-02), then caching (10-03)... with the honest caveat that it would have been more efficient to start with the index and the cache. We did it in the syllabus order for teaching reasons; in your project, follow the table. Finally, what to record in production to know that everything is still fine:
| Metric | Alert threshold | Why |
|---|---|---|
| p99 per route | Above the service objective for 5 min | Real perceived degradation |
| 5xx rate | > 0.5% | Something has broken |
| Loop p99 delay | > 100 ms | Synchronous work has crept in |
heapUsed |
Monotonic growth for 30 min | A memory leak |
| Pool connections in use | > 80% of the maximum | Imminent saturation |
| Cache hit ratio | < 90% | Invalidation too aggressive or a short TTL |
| Waiting / failed jobs | Sustained growth | Not enough consumers, or broken ones |
Collecting, storing and alerting on these metrics is production monitoring, the subject of Module 11; here we have learned to produce them and to interpret them.
Common Mistakes and Tips
- Optimizing without a profile. That is guessing. In Escena Viva the suspect was the PDF and the culprit was the JSON.
- Looking at the average. Optimize the p99; it is what people experience and remember.
- Measuring in development with toy data. With 3 events everything is fast; with 30,000 the query with no index shows up.
- A load test with a non-zero
non2xx, or changing several things at once: in the first case you are measuring how fast you fail, in the second you will not know what worked. - Profiling in development mode. Without
NODE_ENV=production, Express does not cache views and there are extra checks: the profile lies. - Raising
--max-old-space-sizeto paper over a leak. It delays the crash and makes the GC pauses worse. - Tip: keep the baseline results in the repository; comparing against last month catches regressions no unit test ever sees.
- Tip: add a load test with a threshold to CI (
k6withthresholds) and performance stops being an opinion and becomes a verifiable requirement.
Exercises
Exercise 1: baseline and objectives
Define a reproducible load test for GET /api/events/evt-003/sessions with 10 s of warm-up and 60 s of measurement at 100 connections. Record the average, p50, p95, p99, req/s and non2xx. Compare it with the 200 ms service objective and decide whether you need to act.
Exercise 2: hunt the bottleneck with a flamegraph
Add a GET /api/reports/occupancy endpoint to Escena Viva that computes the occupancy of the 7 sessions, sorting the ticket array with a deliberately quadratic algorithm. Profile with 0x, locate the wide box, replace it with Array.prototype.sort and measure again.
Exercise 3: diagnose a leak
Introduce a leak by registering a sale-recorded listener on every request without removing it. Fire 30,000 requests while recording heapUsed every 10 s. Take two heap snapshots, compare them and locate the retainer. Fix it and verify.
Solutions
Exercise 1. The scenario defines warmup: { connections: 10, duration: 10 } and connections: 100, duration: 60. A typical result without a cache: average 96 ms, p50 84 ms, p95 210 ms, p99 340 ms, 1,030 req/s, non2xx: 0. The p99 of 340 ms misses the 200 ms objective, so you have to act. The fine-grained reading is more interesting than the verdict: the average (96 ms) looks comfortable and so does the median; only the p99 reveals the problem. The usual cause here is that the sessions are loaded with one query per session (N+1): the fix is an include (M7) plus a cache (10-03), and after applying them the p99 drops to ~35 ms.
Exercise 2. The flamegraph shows a wide, flat plateau over sortByOccupancy, taking up ~70% of the width, with very little depth on top: the unmistakable sign that time is spent in that function and not in its calls. With 7 sessions the quadratic algorithm is irrelevant; the exercise reveals its cost by running it over the 1,811 tickets sold (~3.3 million comparisons per request). Replacing it with tickets.sort((a, b) => b.occupancy - a.occupancy) (O(n log n)) cuts the route's time from ~420 ms to ~6 ms. It is lever number 1 in the table —the algorithm— and no amount of cluster, threads or caching would have given a factor of 70.
Exercise 3. heapUsed grows monotonically: 42 MB, 78 MB, 121 MB, 168 MB... never settling, not even after a forced collection. Around request 10,000 the MaxListenersExceededWarning shows up. In the snapshot comparison, the Delta shows tens of thousands of extra Closure and ServerResponse objects; the Retainers point at SalesManager._events['sale-recorded'], an array growing without end. Each closure retains res, and with it the socket and the headers: that is why a leak of "one listener" costs kilobytes per request. The fix is res.on('close', () => manager.off('sale-recorded', onSale)), and after applying it heapUsed settles into a sawtooth between 45 and 70 MB: up on allocation, down on collection. That stable sawtooth shape is exactly what a healthy process looks like.
Conclusion
We no longer optimize by intuition. We have a method —measure, locate, fix one thing, measure again— and the four signals to watch, knowing that the average lies and that the p99 is what people experience. We know how to design an honest load test with autocannon (realistic warm-up, duration, concurrency and scenario) and why you never run one against somebody else's system. We know how to take a CPU profile with --prof, with the inspector or with 0x, and how to read a flamegraph by looking for wide boxes and not tall towers. We know how to hunt a leak by comparing heap snapshots and looking at the retainers. And we have event loop delay as health metric number one. On top of that, we have walked through Node's bottleneck catalog —synchronous I/O, giant JSON, N+1, missing indexes, sequential await, synchronous logging, ReDoS, cacheable data— with its fix, we have tuned the HTTP layer and the connection pool, and we have put V8 micro-optimizations in their place: at the end of the list, and only if the profile points at them. Above all, we have the order of the levers: algorithm, query, cache, concurrency, hardware.
Escena Viva is now fast and measured, so what remains is not performance but design. Its API grew by accumulation: routes with verbs, status codes chosen by eye, no cursor pagination, no versioning, no deprecation policy, and one doubt we left open in Module 4 and never resolved: what happens when a POST /api/orders is sent twice because the user lost coverage on opening night. In the next lesson, Building RESTful APIs, we move from "it works" to "it is well designed".
Node.js Course: From Beginner to Advanced
Module 1: Introduction to Node.js
- What Is Node.js?
- Installing and Setting Up the Environment
- Your First Node.js Program
- The Node.js REPL
- Modern JavaScript for Node.js
- The Course Project: the Escena Viva Platform
Module 2: Core Concepts
- Node.js Architecture
- The Event Loop
- Callbacks and Asynchronous Programming
- Promises and async/await
- Events and EventEmitter
- CommonJS Modules and require()
- ES Modules and Interoperability
Module 3: File System and I/O
- Reading and Writing Files
- The fs Module in Depth
- Cross-Platform Paths with the path Module
- Working with Streams
- Transform Streams and pipeline
- Buffers and Binary Data
Module 4: HTTP and Web Servers
- Creating a Simple HTTP Server
- Handling Requests and Responses
- Manual Routing
- Serving Static Files
- Receiving Data: Request Bodies and JSON
- Consuming External APIs from Node.js
Module 5: NPM and Package Management
- Introduction to NPM and package.json
- Installing and Using Packages
- Semantic Versioning and package-lock
- npm Scripts and Project Automation
- Creating and Publishing Packages
- Dependency Security and Maintenance
Module 6: The Express.js Framework
- Introduction to Express.js
- Setting Up an Express Application
- Routing in Express
- Middleware
- Essential Third-Party Middleware
- Input Data Validation
- Error Handling
Module 7: Databases and ORMs
- Introduction to Databases
- Using MongoDB with Mongoose
- CRUD Operations
- Relationships, Population and Advanced Queries
- Using SQL Databases with Sequelize
- Migrations, Transactions and Seed Data
Module 8: Authentication and Authorization
- Introduction to Authentication
- User Registration and Password Hashing
- Sessions and Cookies with Passport.js
- Authentication with JWT
- Role-Based Access Control
- API Security Best Practices
Module 9: Testing and Debugging
- Introduction to Testing
- Unit Testing with Mocha and Chai
- Test Doubles with Sinon
- Integration Testing
- Coverage and Test Automation
- Debugging Node.js Applications
Module 10: Advanced Topics
- The Cluster Module
- Worker Threads
- Caching and Job Queues with Redis
- Performance Optimization
- Building RESTful APIs
- GraphQL with Node.js
Module 11: Deployment and DevOps
- Configuration and Environment Variables
- Logging and Monitoring in Production
- Using PM2 for Process Management
- Packaging with Docker
- Deploying to Heroku and Other PaaS
- Continuous Integration and Deployment
