In the previous three lessons we have quoted a lot of numbers: 641 requests per second, then 3,105, then 11,840; 99th percentiles of 204, 54 and 11 milliseconds; an event loop delay that went from 4,176 ms to 11 ms. Those numbers did not come from intuition: they came from measuring first, changing one thing and measuring again. That practice, which we have been applying without naming it, is the content of this lesson.

You will not find a list of tricks here, but a method, a catalog of instruments (load tests, CPU profiles, flamegraphs, heap snapshots) and a catalog of Node's typical bottlenecks with their fix, all anchored to what you already know from the course.

Contents

  1. The method: measure, locate, fix, measure again
  2. What to measure: latency, throughput, errors, saturation
  3. Load testing with autocannon
  4. CPU profiling and flamegraphs
  5. Memory profiling and leaks
  6. Event loop delay as a health metric
  7. A catalog of typical Node bottlenecks
  8. Optimizations in the HTTP layer
  9. The database: the connection pool
  10. V8 without falling into useless micro-optimization
  11. The correct order of the levers

  1. The method: measure, locate, fix, measure again

The cycle has four steps and none of them is optional: measure the complete system under a realistic load and write down the baseline; locate the bottleneck with a profile, not with a hunch; fix only that, one change at a time; and measure again with the same test, comparing against the baseline.

One change per iteration is non-negotiable. If you touch three things and the system improves by 30%, you do not know which of the three worked, nor whether one of them made things worse and the other two made up for it.

Donald Knuth's quote gets repeated a lot and almost always mutilated. The full sentence is: "we should forget about small efficiencies, say about 97% of the time: premature optimization is the root of all evil. Yet we should not pass up our opportunities in that critical 3%." Both extremes are mistakes: rewriting a loop to save microseconds in code that runs once a day, and also ignoring that your main query does a sequential scan over a million-row table. The difference between the two is not philosophical: it is that you measured one and not the other. Practical corollary: optimizing without measuring is guessing, and a programmer's intuition about where the time goes is notoriously bad. In Escena Viva anyone would have bet the bottleneck was generating the PDF; the profile proved that in the catalog 62% of the time went on serializing JSON and on a query with no index.

  1. What to measure: latency, throughput, errors, saturation

Four signals, known as the golden signals:

Signal What it is How it is measured in Escena Viva
Latency Response time Percentiles per route, separating 2xx from 5xx
Throughput Requests served per second autocannon, and counters in production
Error rate Proportion of 5xx and timeouts Structured logging (M6) by error code
Saturation How full the limiting resource is CPU, memory, pool connections, loop delay

On latency there is one point worth insisting on, and it is that the average lies:

Statistic Value in the catalog What it means
Average / p50 78 ms / 72 ms No individual user experiences the average
p95 161 ms 1 in every 20 requests is worse
p99 204 ms 1 in 100. With 12 requests per page, 1 in every 8 visits sees this
p99.9 / maximum 890 ms The long tail: what triggers the complaints

If ten requests take 10 ms and one takes 5,000 ms, the average is 463 ms: a number that describes neither the fast ones nor the slow one. And the p99 arithmetic is frightening once translated into experience: a page that makes 12 API calls has a 1 − 0.99¹² ≈ 11% probability of hitting at least one 99th-percentile request. Optimize the tail, not the average. These are Escena Viva's service objectives, agreed with the business:

Route p99 target Availability Rationale
GET /api/events < 150 ms 99.9% It is the first screen; if it is slow, nothing sells
GET /api/events/:id/sessions < 200 ms 99.9% The step before the purchase
POST /api/orders < 800 ms 99.95% Transactional; more latency is tolerated, failures are not
POST /auth/login < 400 ms 99.9% bcrypt with cost 12 has a floor of its own

That last case shows that not everything slow is a problem: bcrypt with cost 12 takes ~250 ms on purpose, because that slowness is the defense against brute force (M8), and optimizing it would be a security regression.

  1. Load testing with autocannon

An honest load test meets four conditions. Warm-up: the first seconds do not count, because V8 needs to run the code several times before the optimizing compiler (TurboFan) does its job, and the caches are cold. Sufficient duration: 30-60 seconds minimum, since ten seconds can hide a leak or the effect of a major garbage collection. Realistic concurrency: not 5,000 connections if your real peak is 200. And a realistic scenario: do not hammer a single trivial route, reproduce the user's journey.

# Warm-up (10 s that are discarded) and measurement (60 s, 100 connections).
npx autocannon -c 50 -d 10 -w 4 http://localhost:3000/api/events > /dev/null
npx autocannon -c 100 -d 60 -w 4 --latency http://localhost:3000/api/events

A realistic purchase scenario is defined programmatically:

'use strict';

const autocannon = require('autocannon');

// Festival de Jazz opening-night scenario: browse the catalog, open the
// event page, check the sessions and, in a fraction of cases, buy.
const instance = autocannon({
  url: 'http://localhost:3000',
  connections: 100,
  duration: 60,
  warmup: { connections: 10, duration: 10 },
  requests: [
    { method: 'GET', path: '/api/events' },
    { method: 'GET', path: '/api/events/evt-003' },
    { method: 'GET', path: '/api/events/evt-003/sessions' },
    { method: 'POST', path: '/api/orders',
      headers: { 'content-type': 'application/json', authorization: 'Bearer <token>' },
      body: JSON.stringify({ sessionId: 'ses-003-1', quantity: 2 }) },
  ],
});

autocannon.track(instance, { renderProgressBar: true });
instance.on('done', (r) => console.log(`p99: ${r.latency.p99} ms, non-2xx: ${r.non2xx}`));

And this is how the output is read:

Field What to look at
Latency p99 The number that governs your service objectives
Req/Sec Throughput; always compare it with the baseline
non2xx If it is not 0, the test is worthless: you are measuring how fast you fail
errors / timeouts Rejected or timed-out connections: you have exceeded capacity
Bytes/Sec If you approach the bandwidth, the bottleneck is the network, not your code

k6 is the alternative when you need stateful scenarios (login, cookies, chained sessions), thresholds that fail the test in CI, and distributed load: autocannon is perfect for the fast "measure, change, measure" loop, and k6 for the performance regression test in continuous integration (M11).

An important, non-negotiable warning. A load test is, technically, a denial-of-service attack. Run it only against systems that are yours and in test environments: never against production without prior coordination (and there, prefer mirrored traffic or limited synthetic load), and never against a third-party API, neither the M8 currency service, nor the payment gateway, nor a public service. It is illegal in many jurisdictions and it is always professionally disrespectful.

  1. CPU profiling and flamegraphs

The load test tells you that something is slow; the profile tells you where.

Option 1: V8's built-in profiler. node --prof src/server.js generates an isolate-*.log with program-counter samples; you run the load, stop the server and node --prof-process isolate-*.log > profile.txt turns it into a readable report that starts with the summary that matters most:

 [Summary]:  ticks  total  nonlib   name
              4821  61.9%   72.4%  JavaScript
              1901  24.4%          GC
               502   6.4%          Shared libraries
 [Bottom up (heavy) profile]:
  1204   15.5%  JSON.stringify
   890   11.4%    LazyCompile: *serializeCatalog src/repositories/events.js:88
   731    9.4%  LazyCompile: *calculateAvailableCapacity src/domain/session.js:41

Two immediate readings: JSON.stringify takes 15.5% and garbage collection 24.4%, which suggests we are creating too many temporary objects. The * in front of the name means V8 optimized that function; a ~ would mean it ran unoptimized, and that is sometimes the clue to a hidden-classes problem.

Option 2: the Chrome DevTools inspector. With node --inspect src/server.js and chrome://inspect you get a CPU profile with a flame chart, heap snapshots and comparisons between them in a single interface. On a remote server you use --inspect=0.0.0.0:9229 with an SSH tunnel, never exposed to the internet: the inspector lets anyone run arbitrary code in your process.

Option 3: 0x and flamegraphs.

npx 0x --output-dir profiles -- node src/server.js
# Run the load, stop with Ctrl+C and open the generated HTML.

How to read a flamegraph, which is where almost everybody gets it wrong:

  • The horizontal axis is NOT time, it is the proportion of samples; functions are ordered alphabetically, not chronologically.
  • The width of a box is the total time in that function, including whatever it called. It is the only thing worth looking for: wide boxes.
  • The height is the stack depth. A very tall, narrow tower is just code with many layers: it is not a problem.
  • A wide box with a flat plateau on top means the time is spent in the function itself, not in its calls: there is your expensive function.
  • Color means nothing by default (except in 0x, which highlights unoptimized functions).

In Escena Viva's flamegraph before the cache, a wide plateau over getCatalog → mapEvents → JSON.stringify took up nearly two thirds of the width; after lesson 10-03 that plateau disappears and the width is taken by libuv's idle loop, which is exactly what you want to see.

  1. Memory profiling and leaks

A memory leak in Node does not show up as an error, but as degradation: the process grows, garbage collection works harder and harder (and its pauses are time stolen from your event loop), and eventually the process dies with JavaScript heap out of memory. The diagnostic method with heap snapshots: start the process and let it settle; take snapshot A (DevTools → Memory → Heap snapshot, or v8.writeHeapSnapshot()); apply sustained load for several minutes; force a collection and take snapshot B; and in DevTools choose the Comparison view between B and A.

The decisive column is Delta: how many extra objects of each type there are. If Array or a constructor of yours grows monotonically snapshot after snapshot, there is your leak; then Retainers tells you who is stopping the collector from taking them away.

'use strict';

// v8.writeHeapSnapshot(path) takes the snapshot on demand, in production
// too (carefully: it pauses the process and generates hundreds of MB).
// And this is the cheap, continuous watch over memory usage.
function recordMemoryUsage(logger) {
  const usage = process.memoryUsage();
  const mb = (bytes) => Math.round(bytes / 1048576);
  logger.info({ rssMb: mb(usage.rss), heapUsedMb: mb(usage.heapUsed),
    heapTotalMb: mb(usage.heapTotal), externalMb: mb(usage.external) });
}

module.exports = { recordMemoryUsage };

The two leaks you will see in real life, and both appeared in Escena Viva. The first, an unbounded cache: before lesson 10-03, the currency cache was an in-memory Map which with fixed pairs did not grow, but which, had the key included the amount (EUR:USD:4500), would have grown forever. Every in-memory cache needs a ceiling with LRU eviction (lru-cache) or it needs to move to Redis, which expires on its own. The second, listeners that are never removed: our SalesManager is an EventEmitter (M2), and if every request registers a listener that nobody removes, the emitter accumulates references to the closures and, with them, to everything they capture.

'use strict';

// BAD: the listener outlives the request and retains 'res' forever, and
// with it the socket, the headers and the body.
const badController = (manager) => (req, res) => {
  manager.on('sale-recorded', (sale) => res.write(JSON.stringify(sale)));
};

// GOOD: it is registered, and removed when the request ends. The 'close'
// event covers both a normal ending and the client disconnecting.
const goodController = (manager) => (req, res) => {
  const onSale = (sale) => res.write(JSON.stringify(sale));
  manager.on('sale-recorded', onSale);
  res.on('close', () => manager.off('sale-recorded', onSale));
};

module.exports = { goodController };

The MaxListenersExceededWarning: Possible EventEmitter memory leak detected warning is Node literally telling you that you have this leak. Never silence it by raising setMaxListeners: fix the removal.

  1. Event loop delay as a health metric

If you had to pick one single metric to know whether a Node process is healthy, it would be this one: CPU at 90% can be normal, but a loop delay of 800 ms never is.

'use strict';

const { monitorEventLoopDelay } = require('node:perf_hooks');

// resolution: how often a sample is taken, in ms. 10 ms is a good point.
function createLoopWatcher({ logger, intervalMs = 30_000, p99ThresholdMs = 100 }) {
  const histogram = monitorEventLoopDelay({ resolution: 10 });
  histogram.enable();
  const ms = (value) => Number((value / 1e6).toFixed(2));

  const timer = setInterval(() => {
    const p99 = ms(histogram.percentile(99));
    logger[p99 > p99ThresholdMs ? 'warn' : 'info']({
      metric: 'event-loop-delay', p99Ms: p99,
      p50Ms: ms(histogram.percentile(50)), maxMs: ms(histogram.max),
    });
    histogram.reset(); // sliding window, not cumulative
  }, intervalMs);

  timer.unref();
  return { stop: () => { clearInterval(timer); histogram.disable(); } };
}

module.exports = { createLoopWatcher };

How to interpret the values:

p99 delay Diagnosis
< 10 ms Healthy
10-50 ms High but manageable load
50-200 ms There is synchronous work: investigate with a CPU profile
> 200 ms Something blocks the loop. It is the symptom from lesson 10-02
> 1,000 ms The process is effectively down even if it answers a ping

It is also the best value for a health check: a process with a jammed loop answers 200 OK to a trivial /health while leaving every user stranded, and only a /health that looks at the loop delay detects that state.

  1. A catalog of typical Node bottlenecks

Bottleneck Symptom Fix Reference
Synchronous I/O on the hot path (readFileSync, existsSync per request) High loop delay, profile with fs in the stack The asynchronous version, or read once at startup M3
A giant JSON serialized on every response Wide JSON.stringify in the flamegraph, high GC Paginate, select fields, cache the already serialized JSON 10-03, 10-05
N+1 query Many short, identical queries per request include/JOIN, or DataLoader (10-06) M7
Missing indexes EXPLAIN ANALYZE shows a Seq Scan Create the index and verify it is used M7
Avoidable sequential await Latency equal to the sum of the waits Promise.all on independent operations M2
Synchronous logging Loop delay rises with the log volume An asynchronous, buffered logger (pino), no console.log on the hot path M6, M11
ReDoS (catastrophic backtracking) One specific request freezes the process Regex with no nested ambiguity, or prior length/zod validation M8
Cacheable data regenerated High CPU and database load to return the same thing Cache-aside with a TTL 10-03
CPU work on the main thread Loop delay in seconds Threads or a queue 10-02

Two deserve elaboration. The first, sequential await where Promise.all fitted, is the most common and cheapest performance bug to fix:

'use strict';

// BAD: three 'await' in series over the three repositories.
// Latency = 40 + 35 + 30 = 105 ms.
//
// GOOD: all three are independent. Latency = max(40, 35, 30) = 40 ms.
async function loadEventScreen(eventId, repos) {
  const [event, sessions, venue] = await Promise.all([
    repos.events.getEventById(eventId),
    repos.sessions.listByEvent(eventId),
    repos.venues.getByEvent(eventId),
  ]);
  return { event, sessions, venue };
}

The condition is independence: if the second query needs the first one's result, the sequential await is mandatory. And for operations where one failure must not take the rest down, use Promise.allSettled. The second is ReDoS: regular expressions with catastrophic backtracking, the only bottleneck on this list that is also a security vulnerability (OWASP, M8), because a single 40-character request can freeze the process for minutes.

'use strict';

// DANGEROUS: (\s*\w+)+ has nested ambiguity. With an input such as
// 'aaaaaaaaaaaaaaaaaaaaaaaaaaaaaa!' the engine tries an exponential number
// of combinations before giving up, and the whole process freezes.
const DANGEROUS_REGEX = /^(\s*\w+)+$/;
// SAFE: no ambiguity, a single quantifier over disjoint classes.
const SAFE_REGEX = /^[\w\s]+$/;

// And on top of that: limit the length BEFORE applying the regex.
const validateEventTitle = (title) =>
  typeof title === 'string' && title.length <= 120 && SAFE_REGEX.test(title);

module.exports = { validateEventTitle };

Defense in layers: validate length and shape with zod (M6) before any regex, avoid nested quantifiers, and distrust expressions copied from the internet to validate e-mails or URLs.

  1. Optimizations in the HTTP layer

Before touching your code, check that the transport layer is not giving work away:

Lever How Typical gain
Compression compression (already installed), a 1 KB threshold 60-80% fewer bytes on JSON
ETag + 304 We did it by hand in M4; Express ships it An empty response on repeated reads
Cache-Control public, max-age=60 on the catalog Requests that never reach your server
Outbound keep-alive An agent with keepAlive: true in fetch/HTTP Saves the TCP+TLS handshake on every call
Mandatory pagination A maximum limit on the server Prevents megabyte-sized responses

Compression is not free: it burns CPU. With small responses (< 1 KB) it costs more than it saves, and on a server with saturated CPU it can be counterproductive; in production it is usually delegated to the reverse proxy (M11). About outbound connections, every call to the M8 currency service without keep-alive pays a TCP handshake and a TLS negotiation: some 40-80 ms of pure protocol.

'use strict';

const { Agent, setGlobalDispatcher } = require('undici');

// Reuses connections for every outbound call made with fetch.
setGlobalDispatcher(new Agent({
  connections: 32,        // maximum connections per origin
  keepAliveTimeout: 30_000,
  keepAliveMaxTimeout: 120_000,
}));

  1. The database: the connection pool

Opening a connection to PostgreSQL costs between 5 and 50 ms (authentication, TLS, a process on the server). That is why Sequelize keeps a pool: a set of open connections that are lent out and returned.

'use strict';

const os = require('node:os');
const { Sequelize } = require('sequelize');
const { configuration } = require('../config/index.js');

// PostgreSQL's global budget: max_connections (100). Split it into
// cluster workers x pool size, plus room for queue consumers,
// migrations and administration.
const WORKERS = Math.max(1, os.availableParallelism() - 1); // 7
const MAX_PER_PROCESS = Math.max(2, Math.floor(80 / (WORKERS + 2)));

const sequelize = new Sequelize(configuration.postgresUrl, {
  logging: false,
  pool: {
    max: MAX_PER_PROCESS, // 8 with 7 workers
    min: 2,               // warm connections always ready
    acquire: 10_000,      // maximum wait for a free connection
    idle: 10_000,         // closes idle ones after this time
  },
});

module.exports = { sequelize };

How it is sized, in order: find out the server's max_connections (100 by default in PostgreSQL); reserve room for migrations, queue consumers and administration; divide the rest by the number of processes that connect; and remember that bigger is not better, because a pool of 100 against 4 database cores only adds contention (there is a classic formula suggesting cores × 2 + disk spindles as a starting point).

What happens when the pool runs out: requests wait up to acquire and then fail with SequelizeConnectionAcquireTimeoutError, with a characteristic and misleading symptom —very high latency with no high CPU, because everybody is waiting. If you see it, first check whether slow queries are holding connections or whether there are open transactions that nobody closes.

  1. V8 without falling into useless micro-optimization

V8 optimizes your code through hidden classes: objects created with the same structure share an internal description and access to their properties compiles down to a fixed offset, extremely fast. If you change the shape, V8 creates another one and access becomes polymorphic or megamorphic, much slower.

'use strict';

// BAD: the object's shape changes as it goes (three hidden classes).
function createTicketBad(code, priceCents) {
  const ticket = {};
  ticket.code = code;
  ticket.priceCents = priceCents;
  if (priceCents === 0) ticket.isComplimentary = true; // only sometimes
  return ticket;
}

// GOOD: a single shape, always the same fields and in the same order.
const createTicket = (code, priceCents) => ({
  code, priceCents, isComplimentary: priceCents === 0,
});

// 'delete' degrades the object to dictionary mode and cancels the
// optimization: instead of 'delete ticket.seat', we reassign.
const clearSeat = (ticket) => ({ ...ticket, seat: null });

module.exports = { createTicket, clearSeat };

And about memory: the old heap has a limit (typically ~2 GB on 64-bit, or less in small containers) that is adjusted with --max-old-space-size=2048. Two indispensable nuances: raising it does not fix a leak, it only delays the death and lengthens the garbage collection pauses; and in a container you must set it below the container's memory limit, because otherwise the system kills the process (the OOM killer) before V8 even notices (a detail we will pick up again in M11).

The general warning: these micro-optimizations matter in code that runs millions of times. In an HTTP controller that runs 3,000 times per second and spends 95% of its time waiting on the database, reordering properties changes nothing. Apply them only if the profile points at that function.

  1. The correct order of the levers

When something is slow there is an order by return on investment:

Order Lever Typical gain Cost
1 Algorithm and data structure 10× to 1000× Thinking; sometimes rewriting
2 Query and schema (index, N+1, SELECT of fewer columns) 10× to 100× Low; a migration
3 Cache 5× to 50× Medium: invalidation (10-03)
4 Concurrency (cluster, threads, queues) 2× to 8× Medium-high: shared state
5 Code micro-optimization 1.05× to 1.5× High in readability
6 Hardware Linear, with a monthly bill Recurring money

You work it from top to bottom. Buying a machine twice as big to paper over a query with no index means paying twice as much every month for a problem that a CREATE INDEX would have fixed. And that brings us to Escena Viva's actual path through this module: first concurrency (10-01, 10-02), then caching (10-03)... with the honest caveat that it would have been more efficient to start with the index and the cache. We did it in the syllabus order for teaching reasons; in your project, follow the table. Finally, what to record in production to know that everything is still fine:

Metric Alert threshold Why
p99 per route Above the service objective for 5 min Real perceived degradation
5xx rate > 0.5% Something has broken
Loop p99 delay > 100 ms Synchronous work has crept in
heapUsed Monotonic growth for 30 min A memory leak
Pool connections in use > 80% of the maximum Imminent saturation
Cache hit ratio < 90% Invalidation too aggressive or a short TTL
Waiting / failed jobs Sustained growth Not enough consumers, or broken ones

Collecting, storing and alerting on these metrics is production monitoring, the subject of Module 11; here we have learned to produce them and to interpret them.

Common Mistakes and Tips

  • Optimizing without a profile. That is guessing. In Escena Viva the suspect was the PDF and the culprit was the JSON.
  • Looking at the average. Optimize the p99; it is what people experience and remember.
  • Measuring in development with toy data. With 3 events everything is fast; with 30,000 the query with no index shows up.
  • A load test with a non-zero non2xx, or changing several things at once: in the first case you are measuring how fast you fail, in the second you will not know what worked.
  • Profiling in development mode. Without NODE_ENV=production, Express does not cache views and there are extra checks: the profile lies.
  • Raising --max-old-space-size to paper over a leak. It delays the crash and makes the GC pauses worse.
  • Tip: keep the baseline results in the repository; comparing against last month catches regressions no unit test ever sees.
  • Tip: add a load test with a threshold to CI (k6 with thresholds) and performance stops being an opinion and becomes a verifiable requirement.

Exercises

Exercise 1: baseline and objectives

Define a reproducible load test for GET /api/events/evt-003/sessions with 10 s of warm-up and 60 s of measurement at 100 connections. Record the average, p50, p95, p99, req/s and non2xx. Compare it with the 200 ms service objective and decide whether you need to act.

Exercise 2: hunt the bottleneck with a flamegraph

Add a GET /api/reports/occupancy endpoint to Escena Viva that computes the occupancy of the 7 sessions, sorting the ticket array with a deliberately quadratic algorithm. Profile with 0x, locate the wide box, replace it with Array.prototype.sort and measure again.

Exercise 3: diagnose a leak

Introduce a leak by registering a sale-recorded listener on every request without removing it. Fire 30,000 requests while recording heapUsed every 10 s. Take two heap snapshots, compare them and locate the retainer. Fix it and verify.

Solutions

Exercise 1. The scenario defines warmup: { connections: 10, duration: 10 } and connections: 100, duration: 60. A typical result without a cache: average 96 ms, p50 84 ms, p95 210 ms, p99 340 ms, 1,030 req/s, non2xx: 0. The p99 of 340 ms misses the 200 ms objective, so you have to act. The fine-grained reading is more interesting than the verdict: the average (96 ms) looks comfortable and so does the median; only the p99 reveals the problem. The usual cause here is that the sessions are loaded with one query per session (N+1): the fix is an include (M7) plus a cache (10-03), and after applying them the p99 drops to ~35 ms.

Exercise 2. The flamegraph shows a wide, flat plateau over sortByOccupancy, taking up ~70% of the width, with very little depth on top: the unmistakable sign that time is spent in that function and not in its calls. With 7 sessions the quadratic algorithm is irrelevant; the exercise reveals its cost by running it over the 1,811 tickets sold (~3.3 million comparisons per request). Replacing it with tickets.sort((a, b) => b.occupancy - a.occupancy) (O(n log n)) cuts the route's time from ~420 ms to ~6 ms. It is lever number 1 in the table —the algorithm— and no amount of cluster, threads or caching would have given a factor of 70.

Exercise 3. heapUsed grows monotonically: 42 MB, 78 MB, 121 MB, 168 MB... never settling, not even after a forced collection. Around request 10,000 the MaxListenersExceededWarning shows up. In the snapshot comparison, the Delta shows tens of thousands of extra Closure and ServerResponse objects; the Retainers point at SalesManager._events['sale-recorded'], an array growing without end. Each closure retains res, and with it the socket and the headers: that is why a leak of "one listener" costs kilobytes per request. The fix is res.on('close', () => manager.off('sale-recorded', onSale)), and after applying it heapUsed settles into a sawtooth between 45 and 70 MB: up on allocation, down on collection. That stable sawtooth shape is exactly what a healthy process looks like.

Conclusion

We no longer optimize by intuition. We have a method —measure, locate, fix one thing, measure again— and the four signals to watch, knowing that the average lies and that the p99 is what people experience. We know how to design an honest load test with autocannon (realistic warm-up, duration, concurrency and scenario) and why you never run one against somebody else's system. We know how to take a CPU profile with --prof, with the inspector or with 0x, and how to read a flamegraph by looking for wide boxes and not tall towers. We know how to hunt a leak by comparing heap snapshots and looking at the retainers. And we have event loop delay as health metric number one. On top of that, we have walked through Node's bottleneck catalog —synchronous I/O, giant JSON, N+1, missing indexes, sequential await, synchronous logging, ReDoS, cacheable data— with its fix, we have tuned the HTTP layer and the connection pool, and we have put V8 micro-optimizations in their place: at the end of the list, and only if the profile points at them. Above all, we have the order of the levers: algorithm, query, cache, concurrency, hardware.

Escena Viva is now fast and measured, so what remains is not performance but design. Its API grew by accumulation: routes with verbs, status codes chosen by eye, no cursor pagination, no versioning, no deprecation policy, and one doubt we left open in Module 4 and never resolved: what happens when a POST /api/orders is sent twice because the user lost coverage on opening night. In the next lesson, Building RESTful APIs, we move from "it works" to "it is well designed".

Node.js Course: From Beginner to Advanced

Module 1: Introduction to Node.js

Module 2: Core Concepts

Module 3: File System and I/O

Module 4: HTTP and Web Servers

Module 5: NPM and Package Management

Module 6: The Express.js Framework

Module 7: Databases and ORMs

Module 8: Authentication and Authorization

Module 9: Testing and Debugging

Module 10: Advanced Topics

Module 11: Deployment and DevOps

Module 12: Real-World Projects

© Copyright 2026. All rights reserved