The Escena Viva API sells tickets, persists orders inside transactions that prevent overselling, authenticates users, rotates refresh tokens and authorizes operations based on roles and resource ownership. It does an enormous number of things. And nobody has ever checked that it works. Not once.

npm test has been failing on purpose since Module 5 with the message No tests yet — Module 9, and every line of security we wrote in the previous module — the bcrypt decoy hash, the refresh-reuse detection, every cell of the permission matrix — rests today on our confidence that we wrote it correctly. One inverted if in canManageEvent would open the entire event catalog to any organizer, and nothing would warn us.

In this lesson we will not yet write the first runnable test or install anything: that is the next lesson. Here we lay the conceptual foundations, which is what separates a project with many tests from a project with good tests.

Contents

  1. Why we test: changing code without fear
  2. What an automated test is and the AAA structure
  3. Test types and the pyramid
  4. What to test and what not to: the risk criterion
  5. What makes a test good
  6. Brittle tests and false positives
  7. TDD explained honestly
  8. The tooling landscape for Node
  9. The Escena Viva test plan

Why we test: changing code without fear

There is a widespread idea that tests are there "just in case", like an insurance policy you might one day claim on. It is a poor description and it explains why so many teams abandon them after three weeks. The real value of a test suite is something else, and it is measurable day to day: it gives you back the ability to modify your code.

Let's run the acid test on our own project. Answer honestly:

  • Would you dare today to rewrite buyTickets so that several sessions can be grouped into a single order?
  • Would you dare to refactor PERMISSIONS_BY_ROLE to add a box-office role without manually reviewing all 30 combinations?
  • Would you dare to swap the order of two middleware functions in createApplication so that authenticate runs before purchaseLimit?
  • Would you dare to bump Mongoose to a new major version?

If the honest answer is "no, or only very carefully and after clicking around in the browser", the code is already governing you. That paralysis has a cost that never shows up on any invoice: functions get copied instead of reused, if statements pile up because nobody deletes old branches, and technical debt grows because touching what already works is frightening.

A test suite turns the question "did I break something?" into a ten-second command. That is the product we are buying in this module. The secondary benefits are real, but secondary:

Benefit What it really gives you
Catching regressions A bug you already fixed cannot come back unnoticed
Living documentation it('rejects the sale when capacity is insufficient') never goes stale, because if it lies it fails
Design pressure Code that is hard to test is usually badly coupled; testing reveals it early
Deployment confidence CI can block a broken deployment with no human involved
Onboarding Newcomers read the tests to understand the expected behavior

And there is something tests do not do, worth saying up front: they do not prove the program is correct. They prove that the cases you thought of behave the way you expected. That is a great deal, but it is not a mathematical guarantee.

What an automated test is and the AAA structure

An automated test is simply a program that runs another program and checks the result, finishing with success or failure without anyone watching the screen. There is no magic: a test file is ordinary JavaScript.

The universal structure is known as AAA — arrange, act, assert:

  1. Arrange: build the scenario. Data, objects, test doubles, initial state.
  2. Act: execute exactly one operation, the one under test.
  3. Assert: verify that the observable result is the expected one.

Let's look at a first example on a real piece of our domain, Session.sell(). It is not yet runnable code with Mocha — we haven't installed it — but the anatomy is already final:

// Conceptual sketch of a test for Session.sell()
// (the runnable Mocha and Chai syntax arrives in 09-02)

const { Session } = require('../../src/domain/session.js');

function testSellSubtractsFromCapacity() {
  // 1. ARRANGE: a known, controlled session
  const session = new Session({
    id: 'ses-001-1',
    eventId: 'evt-001',
    startsAt: '2026-10-17T20:00:00.000Z',
    capacity: 400,
    sold: 100,
    priceCents: 2500,
  });

  // 2. ACT: a single operation
  session.sell(3);

  // 3. ASSERT: the observable effect
  if (session.available !== 297) {
    throw new Error('Expected 297 available seats');
  }
}

Three important observations about this sketch:

  • The arrangement is explicit and local. We do not read the session from the database, nor do we depend on the seed data. If tomorrow someone changes the capacity of ses-001-1 in scripts/seed.js, this test remains valid. A test that depends on distant data is a test that will fail for reasons unrelated to what it claims to check.
  • Acting is a single line. If you need five calls to act, you are probably testing five things and the failure will not tell you which one.
  • We assert on observable behavior, available, not on the private field #sold. In fact we could not: it is private. That is an advantage, not an obstacle.

An implicit fourth step, cleanup, appears whenever there are resources to release (connections, fake clocks, files). In the pure domain it is not needed, and that is precisely its charm.

Test types and the pyramid

Not all tests cost the same or catch the same things. This table is the map of the whole module:

Type What it exercises Speed Cost to write Brittleness What it catches
Unit A single function or class in isolation Milliseconds Low Low Logic and calculation errors
Integration Several pieces together (routes + middleware + repository + DB) Tenths of a second to seconds Medium Medium Wiring, contracts between layers, middleware order
End to end The complete system with a browser or a real client Seconds to minutes High High Full user flows, deployment problems
Contract The agreed format between two services Fast Medium Low API changes that would break a consumer
Load The system under pressure Minutes High High Bottlenecks, degradation, resource limits

The first three form the testing pyramid: many unit tests at the base, a fair number of integration tests in the middle, very few end-to-end tests at the tip.

graph TD
    E2E["End to end<br/>few · slow · brittle<br/>critical purchase flows"]
    INT["Integration<br/>a fair number · medium<br/>routes, middleware, repositories, DB"]
    UNI["Unit<br/>many · blazing fast · stable<br/>domain, policy, utilities"]
    E2E --> INT
    INT --> UNI
    style E2E fill:#f8d7da,stroke:#842029
    style INT fill:#fff3cd,stroke:#664d03
    style UNI fill:#d1e7dd,stroke:#0f5132

The shape is not an aesthetic whim: it is economics. A unit test of canManageEvent takes less than a millisecond and, when it fails, points you at one specific function. An end-to-end test of the same permission takes ten seconds, may fail because the browser was slow to load, and when it fails it tells you "the button wasn't there", which can mean fifty different things.

The ice cream cone antipattern

When a team inverts the pyramid — a huge number of end-to-end tests, a few integration tests, almost no unit tests — the result is the so-called ice cream cone. It is a disaster you can recognize by its symptoms:

  • The suite takes 40 minutes and people stop running it locally.
  • A one-line change in a template breaks thirty tests.
  • Flaky tests appear: they fail one time in five for no apparent reason.
  • Because they often fail from noise, the team starts retrying until they pass. At that moment the suite has stopped providing information and provides only waiting.

The practical rule: each behavior is tested at the lowest level that can demonstrate it. Move up a level only when the level below cannot see the problem.

What to test and what not to: the risk criterion

Testing everything is impossible, and chasing it produces expensive suites that age badly. The criterion we will use is risk, understood as the probability of failure multiplied by the damage it would cause.

Do test:

  • Business logic with rules. Session.sell() validates quantity, capacity and status. If it gets it wrong, we sell tickets that do not exist.
  • The permission matrix. A mistake here is a silent security hole.
  • Calculations. Totals in cents, the limit of 6 tickets per order, occupancy, the low-capacity threshold.
  • Error handling. That a ResourceNotFound really ends up as a 404 with the shape { error: { code, message, status, details } }.
  • Edge and boundary cases. Selling exactly the last tickets, selling one more, quantity zero, negative quantity.
  • Every bug that has happened in production. Before fixing it, a test that reproduces it. We will come back to this in 09-06.

Don't test (or test very little):

  • Trivial getters. Whether get priceEuros() { return this.priceCents / 100; } deserves a test is debatable; if it rounded, then yes, because there is a decision there.
  • Third-party libraries. It is not your job to check that Express routes or that bcrypt hashes. It is your job to check how you use them.
  • Static configuration. That PERMISSIONS has five keys says nothing; whether a role can or cannot use them does.
  • Implementation details. That a method internally calls another private method is not a behavior; it is a decision you want to be free to change.
  • Generated code. Auto-generated migrations, trivial bootstrap files.

What makes a test good

Four properties, in order of importance:

  1. Fast. If the unit suite takes more than a few seconds, you will stop running it while you code.
  2. Deterministic. The same code with the same input always gives the same result. No Date.now(), no Math.random(), no network, no dependence on ordering. We will tame all of that with Sinon in 09-03.
  3. Isolated. It does not depend on another test having run first, nor does it leave residue for the next one.
  4. With a single reason to fail. When it turns red, it must be obvious which behavior broke.

And a fifth one that is not technical but decides whether the suite survives: the name must describe the expected behavior, not the method invoked. The name is the only thing you will see in the output when it fails at three in the morning.

Bad name Why it is bad Good name
test sell Says nothing about what is expected subtracts the tickets sold from the available capacity
test 1 Says nothing at all throws StateConflict when there are not enough seats left
works Not falsifiable marks the session as sold out when the last ticket is sold
policy Far too broad lets the venue organizer edit their own event
does not fail Absence of an error is not a behavior returns 422 when more than 6 tickets are requested
case 3 of ticket 481 The ticket will disappear prevents an attendee from seeing another attendee's order

Notice that every good name starts with a third-person verb and describes the observable effect. Read together with the describe that contains them they form a sentence: Session sell — throws StateConflict when there are not enough seats left.

Brittle tests and false positives

There are two quality failures that ruin entire suites.

A brittle test is one that breaks when the code changes without the behavior changing. It is almost always born from testing the implementation instead of the behavior:

// BRITTLE: checks how it is built inside
// If tomorrow the manager computes the total with reduce instead of a loop,
// this test fails even though nothing broke for the user.
test('manager calls calculateSubtotal once per line', () => {
  const spy = spyOnPrivateMethod(manager, 'calculateSubtotal');
  manager.recordSale(order);
  check(spy.calls === 2);
});

// ROBUST: checks what you get
test('the order total adds up the price of all its lines', () => {
  const order = manager.recordSale({ lines: [
    { quantity: 2, priceCents: 2500 },
    { quantity: 1, priceCents: 1800 },
  ] });
  check(order.totalCents === 6800);
});

A false positive (or false green) is a test that always passes, even with broken code. It is worse than having no test, because it creates unfounded confidence. The most common causes in Node:

  • An unawaited promise: the it finishes before the assertion runs. We will see this in detail in 09-02.
  • An assertion inside a catch that never executes.
  • Faking the dependency so thoroughly that the test only checks your own assumptions (the central danger of 09-03).

A professional trick we will use several times in this module: when you write a test, break it on purpose. Flip a sign in the production code and check that it turns red. A test you have never seen fail is not a test, it is a hope.

TDD explained honestly

Test-Driven Development reverses the usual order: first the test, then the code.

graph LR
    R["RED<br/>write a test<br/>that fails"] --> V["GREEN<br/>the minimum code<br/>that makes it pass"]
    V --> RF["REFACTOR<br/>improve the design<br/>without touching the test"]
    RF --> R
    style R fill:#f8d7da,stroke:#842029
    style V fill:#d1e7dd,stroke:#0f5132
    style RF fill:#cfe2ff,stroke:#084298

A complete cycle on a real Escena Viva rule: no order may exceed 6 tickets per session.

Red. We write the test before the check exists:

// The rule is not implemented yet: this test MUST fail.
test('rejects an order of more than 6 tickets for the same session', () => {
  const session = createSession({ capacity: 400, sold: 0 });
  let threw = false;
  try {
    session.sell(7);
  } catch (error) {
    threw = true;
  }
  check(threw === true);
});

We run it and it fails. This step is not bureaucracy: it verifies that the test can detect the failure. If it passed while still red, the test would be badly written.

Green. The minimum code that satisfies it:

// src/domain/session.js (excerpt)
const MAX_TICKETS_PER_ORDER = 6;

sell(quantity) {
  if (!Number.isInteger(quantity) || quantity < 1) {
    throw new ValidationError('Quantity must be a positive integer');
  }
  if (quantity > MAX_TICKETS_PER_ORDER) {
    throw new ValidationError(`Maximum ${MAX_TICKETS_PER_ORDER} tickets per order`);
  }
  // ...rest of the capacity logic
}

Refactor. Now we improve things: we extract MAX_TICKETS_PER_ORDER into an exported constant so that the zod schema and the error message use the same number, and we add the error's details with the maximum allowed. The test protects us during the change.

And now the promised honesty:

TDD helps a lot when... TDD gets in the way when...
The business rule is clear and verifiable You are exploring an API you do not know
You are fixing a bug (the test reproduces it first) You are laying out a visual interface
You are designing a pure function, like the permission policy The design will change three times this afternoon
You work on legacy code and want to pin down the current behavior You are writing a single-use script

It is not a religion. In this course we will write some tests before and many after, and in 09-06 we will use TDD strictly for one specific thing where it is beyond dispute: reproducing a bug before fixing it.

The tooling landscape for Node

Tool What it is Strengths Drawbacks
Mocha Test runner Mature, flexible, imposes neither assertions nor doubles, transparent configuration You have to assemble the stack (Chai, Sinon) yourself
Chai Assertion library Very readable expect, clear failure messages, extensible An extra dependency
Sinon Test doubles Spies, stubs, fake clocks, the de facto standard Easy to overuse
supertest HTTP client for tests Tests an Express app without a manual listen HTTP only
c8 / nyc Coverage c8 uses V8's native coverage, no instrumentation A metric that is easy to misread
Jest All in one Runner, assertions, mocks and coverage together; snapshots Opinionated; its transform complicates pure CommonJS/ESM
Vitest Modern all in one Blazing fast, Jest-compatible API, excellent ESM support Born in the Vite world; less common in pure Node APIs
node:test Native runner Zero dependencies, shipped with Node, with --test and built-in coverage Younger API, smaller plugin ecosystem

The native runner deserves an honest paragraph. Since Node 20 it is stable and perfectly usable:

// How it would look with the native runner (for reference, we will not use it)
const { test } = require('node:test');
const assert = require('node:assert/strict');

test('subtracts the tickets sold from the capacity', () => {
  const session = createSession({ capacity: 400, sold: 100 });
  session.sell(3);
  assert.equal(session.available, 297);
});

If you were starting a new project today, node:test is a defensible, dependency-free choice.

Why this course uses Mocha + Chai + Sinon + supertest: because it is the most widespread stack in the Node APIs you will meet in production, because it clearly separates the three responsibilities (run, assert, fake) and that teaches the concepts better, because its error reporting is explicit, and because the ideas transfer to any other stack in an afternoon. Having chosen the stack, we will not switch halfway through the module: mixing runners is a classic source of confusion.

The Escena Viva test plan

This is the route through the module:

Lesson What gets tested Tools
09-02 Pure domain: Session, Event, SalesManager, and the full matrix in policy.js Mocha, Chai
09-03 Units with dependencies: currency-exchange.js, expiring tokens, repositories Sinon
09-04 The real API end to end: purchases, 401/403/404/409/422 rejections, concurrency supertest
09-05 Coverage, thresholds, npm scripts, flaky tests, getting ready for CI c8
09-06 Debugging whatever fails: inspector, VS Code, stack traces, memory leaks Node and DevTools

The folder structure we will create, mirroring src/:

escena-viva/
├── src/
│   ├── domain/
│   ├── authorization/
│   ├── repositories/
│   └── ...
├── test/
│   ├── unit/
│   │   ├── session.test.js
│   │   ├── event.test.js
│   │   ├── policy.test.js
│   │   └── sales-manager.test.js
│   ├── integration/
│   │   ├── events.test.js
│   │   └── orders.test.js
│   └── helpers/
│       ├── factories.js
│       └── authentication.js
└── .mocharc.json

Three conventions we will follow without exception:

  • The .test.js suffix on every test file: it makes Mocha's search pattern trivial and distinguishes tests from helpers.
  • Mirroring the source code: src/domain/session.js is tested in test/unit/session.test.js. Finding the test for a file must never require a search.
  • test/helpers/ contains no tests, only shared utilities: data factories and the authentication helper that will finally deliver on the promise made in Module 8.

Common Mistakes and Tips

  • Writing tests to push a number up. Coverage is an indicator, not a goal. We will cover it thoroughly in 09-05, but internalize it now: a test that runs code without checking anything raises coverage and protects you from nothing.
  • Starting with end-to-end tests because "they test for real". You will end up with the ice cream cone and a suite nobody runs.
  • Testing private methods. If you feel that urge, that logic usually wants to be a public function in another module. #sold is tested through available and soldOut.
  • Tests that depend on ordering. "It works if you run the whole file but fails on its own" is the symptom. Every it must be runnable in isolation.
  • Reusing objects across tests to "avoid repetition". Sharing a mutable object between it blocks is the number one cause of flaky tests. Use factories that return fresh objects.
  • Tip: start with the layer of highest risk and lowest cost. In Escena Viva that is src/authorization/policy.js: pure functions, no dependencies, protecting the most sensitive part of the system. It can be tested in full in fifty lines.
  • Tip: treat test code as production code. It is read more often than it is written, and it has to be maintained.

Exercises

Exercise 1: classify by level

For each Escena Viva behavior, say at which level of the pyramid you would test it and why:

  1. The total of an order of 2 tickets at 2500 cents is 5000.
  2. POST /api/orders without an Authorization header returns 401.
  3. An organizer from org-boveda cannot edit evt-001, which belongs to org-almendra.
  4. The purchaseLimit middleware is applied before the orders controller.
  5. hashPassword uses cost 12.

Exercise 2: rewrite the names

Rewrite these test names following the rules from the lesson:

  1. test buyTickets
  2. policy works fine
  3. error 409
  4. does not blow up with quantity 0

Exercise 3: design a TDD cycle

Write (in pseudocode, without running it) the red-green-refactor cycle for this new rule: a session that has already started accepts no sales. State which test you would write first, what minimum code would make it pass and what you would refactor afterwards.

Solutions

Exercise 1

  1. Unit. It is a pure calculation in the domain; it needs neither HTTP nor a database.
  2. Integration. The 401 comes from the combination of route + authenticate middleware + errorHandler. No unit test of the middleware guarantees that it is actually mounted on that route.
  3. Unit, with canManageEvent, because it is a pure function and we can walk the whole matrix cheaply and exhaustively. Plus one confirming integration test (a real 403 over HTTP), not all thirty.
  4. Integration. Middleware order is only observable by running the real application: that is exactly what unit tests cannot see.
  5. None, or almost none. It is a configuration constant. Testing bcrypt's exact cost is testing the library; what does make sense is a test that verifyPassword accepts the hash produced by hashPassword (behavior, not detail).

Exercise 2

  1. creates an order with pending status and subtracts it from the session capacity
  2. lets the administrator manage any event
  3. throws StateConflict when the available seats are fewer than the ones requested
  4. throws ValidationError when the quantity is zero

Exercise 3

Red: a test named rejects the sale when the session has already started, which creates a session with startsAt in the past, calls sell(1) and expects a StateConflict. It fails because sell never looks at the date.

Green: add a check at the top of sell: if (new Date(this.startsAt) <= new Date()) { throw new StateConflict('The session has already started'); }.

Refactor: extract a get hasStarted() getter that encapsulates the comparison, reusable by the catalog so that past sessions are not listed, and inject the clock instead of calling new Date() directly. That last point is precisely what makes lesson 09-03 possible: code that depends on the real clock cannot be tested deterministically.

Conclusion

We have laid the foundations. We know why we test — to be able to change the code without fear, not just in case — how a test is structured with arrange, act and assert, which levels exist and why the pyramid has the shape it has, what deserves a test according to risk and what does not, which properties make a test good, how to tell a brittle test from a robust one, when TDD really helps and which tools the Node ecosystem offers.

We also have the plan: test/unit/, test/integration/, test/helpers/, *.test.js files, and the Mocha + Chai + Sinon + supertest stack.

What we do not have yet is a single runnable line. In the next lesson, Unit Testing with Mocha and Chai, we install the stack, finally fix that test script that has been failing on purpose for four modules, and write the complete unit suite for the Escena Viva domain, including that table of cases walking the permission matrix cell by cell that we have been promising since Module 8.

Node.js Course: From Beginner to Advanced

Module 1: Introduction to Node.js

Module 2: Core Concepts

Module 3: File System and I/O

Module 4: HTTP and Web Servers

Module 5: NPM and Package Management

Module 6: The Express.js Framework

Module 7: Databases and ORMs

Module 8: Authentication and Authorization

Module 9: Testing and Debugging

Module 10: Advanced Topics

Module 11: Deployment and DevOps

Module 12: Real-World Projects

© Copyright 2026. All rights reserved