The Escena Viva API sells tickets, persists orders inside transactions that prevent overselling, authenticates users, rotates refresh tokens and authorizes operations based on roles and resource ownership. It does an enormous number of things. And nobody has ever checked that it works. Not once.
npm test has been failing on purpose since Module 5 with the message No tests yet — Module 9, and every line of security we wrote in the previous module — the bcrypt decoy hash, the refresh-reuse detection, every cell of the permission matrix — rests today on our confidence that we wrote it correctly. One inverted if in canManageEvent would open the entire event catalog to any organizer, and nothing would warn us.
In this lesson we will not yet write the first runnable test or install anything: that is the next lesson. Here we lay the conceptual foundations, which is what separates a project with many tests from a project with good tests.
Contents
- Why we test: changing code without fear
- What an automated test is and the AAA structure
- Test types and the pyramid
- What to test and what not to: the risk criterion
- What makes a test good
- Brittle tests and false positives
- TDD explained honestly
- The tooling landscape for Node
- The Escena Viva test plan
Why we test: changing code without fear
There is a widespread idea that tests are there "just in case", like an insurance policy you might one day claim on. It is a poor description and it explains why so many teams abandon them after three weeks. The real value of a test suite is something else, and it is measurable day to day: it gives you back the ability to modify your code.
Let's run the acid test on our own project. Answer honestly:
- Would you dare today to rewrite
buyTicketsso that several sessions can be grouped into a single order? - Would you dare to refactor
PERMISSIONS_BY_ROLEto add abox-officerole without manually reviewing all 30 combinations? - Would you dare to swap the order of two middleware functions in
createApplicationso thatauthenticateruns beforepurchaseLimit? - Would you dare to bump Mongoose to a new major version?
If the honest answer is "no, or only very carefully and after clicking around in the browser", the code is already governing you. That paralysis has a cost that never shows up on any invoice: functions get copied instead of reused, if statements pile up because nobody deletes old branches, and technical debt grows because touching what already works is frightening.
A test suite turns the question "did I break something?" into a ten-second command. That is the product we are buying in this module. The secondary benefits are real, but secondary:
| Benefit | What it really gives you |
|---|---|
| Catching regressions | A bug you already fixed cannot come back unnoticed |
| Living documentation | it('rejects the sale when capacity is insufficient') never goes stale, because if it lies it fails |
| Design pressure | Code that is hard to test is usually badly coupled; testing reveals it early |
| Deployment confidence | CI can block a broken deployment with no human involved |
| Onboarding | Newcomers read the tests to understand the expected behavior |
And there is something tests do not do, worth saying up front: they do not prove the program is correct. They prove that the cases you thought of behave the way you expected. That is a great deal, but it is not a mathematical guarantee.
What an automated test is and the AAA structure
An automated test is simply a program that runs another program and checks the result, finishing with success or failure without anyone watching the screen. There is no magic: a test file is ordinary JavaScript.
The universal structure is known as AAA — arrange, act, assert:
- Arrange: build the scenario. Data, objects, test doubles, initial state.
- Act: execute exactly one operation, the one under test.
- Assert: verify that the observable result is the expected one.
Let's look at a first example on a real piece of our domain, Session.sell(). It is not yet runnable code with Mocha — we haven't installed it — but the anatomy is already final:
// Conceptual sketch of a test for Session.sell()
// (the runnable Mocha and Chai syntax arrives in 09-02)
const { Session } = require('../../src/domain/session.js');
function testSellSubtractsFromCapacity() {
// 1. ARRANGE: a known, controlled session
const session = new Session({
id: 'ses-001-1',
eventId: 'evt-001',
startsAt: '2026-10-17T20:00:00.000Z',
capacity: 400,
sold: 100,
priceCents: 2500,
});
// 2. ACT: a single operation
session.sell(3);
// 3. ASSERT: the observable effect
if (session.available !== 297) {
throw new Error('Expected 297 available seats');
}
}Three important observations about this sketch:
- The arrangement is explicit and local. We do not read the session from the database, nor do we depend on the seed data. If tomorrow someone changes the capacity of
ses-001-1inscripts/seed.js, this test remains valid. A test that depends on distant data is a test that will fail for reasons unrelated to what it claims to check. - Acting is a single line. If you need five calls to act, you are probably testing five things and the failure will not tell you which one.
- We assert on observable behavior,
available, not on the private field#sold. In fact we could not: it is private. That is an advantage, not an obstacle.
An implicit fourth step, cleanup, appears whenever there are resources to release (connections, fake clocks, files). In the pure domain it is not needed, and that is precisely its charm.
Test types and the pyramid
Not all tests cost the same or catch the same things. This table is the map of the whole module:
| Type | What it exercises | Speed | Cost to write | Brittleness | What it catches |
|---|---|---|---|---|---|
| Unit | A single function or class in isolation | Milliseconds | Low | Low | Logic and calculation errors |
| Integration | Several pieces together (routes + middleware + repository + DB) | Tenths of a second to seconds | Medium | Medium | Wiring, contracts between layers, middleware order |
| End to end | The complete system with a browser or a real client | Seconds to minutes | High | High | Full user flows, deployment problems |
| Contract | The agreed format between two services | Fast | Medium | Low | API changes that would break a consumer |
| Load | The system under pressure | Minutes | High | High | Bottlenecks, degradation, resource limits |
The first three form the testing pyramid: many unit tests at the base, a fair number of integration tests in the middle, very few end-to-end tests at the tip.
graph TD
E2E["End to end<br/>few · slow · brittle<br/>critical purchase flows"]
INT["Integration<br/>a fair number · medium<br/>routes, middleware, repositories, DB"]
UNI["Unit<br/>many · blazing fast · stable<br/>domain, policy, utilities"]
E2E --> INT
INT --> UNI
style E2E fill:#f8d7da,stroke:#842029
style INT fill:#fff3cd,stroke:#664d03
style UNI fill:#d1e7dd,stroke:#0f5132
The shape is not an aesthetic whim: it is economics. A unit test of canManageEvent takes less than a millisecond and, when it fails, points you at one specific function. An end-to-end test of the same permission takes ten seconds, may fail because the browser was slow to load, and when it fails it tells you "the button wasn't there", which can mean fifty different things.
The ice cream cone antipattern
When a team inverts the pyramid — a huge number of end-to-end tests, a few integration tests, almost no unit tests — the result is the so-called ice cream cone. It is a disaster you can recognize by its symptoms:
- The suite takes 40 minutes and people stop running it locally.
- A one-line change in a template breaks thirty tests.
- Flaky tests appear: they fail one time in five for no apparent reason.
- Because they often fail from noise, the team starts retrying until they pass. At that moment the suite has stopped providing information and provides only waiting.
The practical rule: each behavior is tested at the lowest level that can demonstrate it. Move up a level only when the level below cannot see the problem.
What to test and what not to: the risk criterion
Testing everything is impossible, and chasing it produces expensive suites that age badly. The criterion we will use is risk, understood as the probability of failure multiplied by the damage it would cause.
Do test:
- Business logic with rules.
Session.sell()validates quantity, capacity and status. If it gets it wrong, we sell tickets that do not exist. - The permission matrix. A mistake here is a silent security hole.
- Calculations. Totals in cents, the limit of 6 tickets per order, occupancy, the low-capacity threshold.
- Error handling. That a
ResourceNotFoundreally ends up as a 404 with the shape{ error: { code, message, status, details } }. - Edge and boundary cases. Selling exactly the last tickets, selling one more, quantity zero, negative quantity.
- Every bug that has happened in production. Before fixing it, a test that reproduces it. We will come back to this in 09-06.
Don't test (or test very little):
- Trivial getters. Whether
get priceEuros() { return this.priceCents / 100; }deserves a test is debatable; if it rounded, then yes, because there is a decision there. - Third-party libraries. It is not your job to check that Express routes or that bcrypt hashes. It is your job to check how you use them.
- Static configuration. That
PERMISSIONShas five keys says nothing; whether a role can or cannot use them does. - Implementation details. That a method internally calls another private method is not a behavior; it is a decision you want to be free to change.
- Generated code. Auto-generated migrations, trivial bootstrap files.
What makes a test good
Four properties, in order of importance:
- Fast. If the unit suite takes more than a few seconds, you will stop running it while you code.
- Deterministic. The same code with the same input always gives the same result. No
Date.now(), noMath.random(), no network, no dependence on ordering. We will tame all of that with Sinon in 09-03. - Isolated. It does not depend on another test having run first, nor does it leave residue for the next one.
- With a single reason to fail. When it turns red, it must be obvious which behavior broke.
And a fifth one that is not technical but decides whether the suite survives: the name must describe the expected behavior, not the method invoked. The name is the only thing you will see in the output when it fails at three in the morning.
| Bad name | Why it is bad | Good name |
|---|---|---|
test sell |
Says nothing about what is expected | subtracts the tickets sold from the available capacity |
test 1 |
Says nothing at all | throws StateConflict when there are not enough seats left |
works |
Not falsifiable | marks the session as sold out when the last ticket is sold |
policy |
Far too broad | lets the venue organizer edit their own event |
does not fail |
Absence of an error is not a behavior | returns 422 when more than 6 tickets are requested |
case 3 of ticket 481 |
The ticket will disappear | prevents an attendee from seeing another attendee's order |
Notice that every good name starts with a third-person verb and describes the observable effect. Read together with the describe that contains them they form a sentence: Session sell — throws StateConflict when there are not enough seats left.
Brittle tests and false positives
There are two quality failures that ruin entire suites.
A brittle test is one that breaks when the code changes without the behavior changing. It is almost always born from testing the implementation instead of the behavior:
// BRITTLE: checks how it is built inside
// If tomorrow the manager computes the total with reduce instead of a loop,
// this test fails even though nothing broke for the user.
test('manager calls calculateSubtotal once per line', () => {
const spy = spyOnPrivateMethod(manager, 'calculateSubtotal');
manager.recordSale(order);
check(spy.calls === 2);
});
// ROBUST: checks what you get
test('the order total adds up the price of all its lines', () => {
const order = manager.recordSale({ lines: [
{ quantity: 2, priceCents: 2500 },
{ quantity: 1, priceCents: 1800 },
] });
check(order.totalCents === 6800);
});A false positive (or false green) is a test that always passes, even with broken code. It is worse than having no test, because it creates unfounded confidence. The most common causes in Node:
- An unawaited promise: the
itfinishes before the assertion runs. We will see this in detail in 09-02. - An assertion inside a
catchthat never executes. - Faking the dependency so thoroughly that the test only checks your own assumptions (the central danger of 09-03).
A professional trick we will use several times in this module: when you write a test, break it on purpose. Flip a sign in the production code and check that it turns red. A test you have never seen fail is not a test, it is a hope.
TDD explained honestly
Test-Driven Development reverses the usual order: first the test, then the code.
graph LR
R["RED<br/>write a test<br/>that fails"] --> V["GREEN<br/>the minimum code<br/>that makes it pass"]
V --> RF["REFACTOR<br/>improve the design<br/>without touching the test"]
RF --> R
style R fill:#f8d7da,stroke:#842029
style V fill:#d1e7dd,stroke:#0f5132
style RF fill:#cfe2ff,stroke:#084298
A complete cycle on a real Escena Viva rule: no order may exceed 6 tickets per session.
Red. We write the test before the check exists:
// The rule is not implemented yet: this test MUST fail.
test('rejects an order of more than 6 tickets for the same session', () => {
const session = createSession({ capacity: 400, sold: 0 });
let threw = false;
try {
session.sell(7);
} catch (error) {
threw = true;
}
check(threw === true);
});We run it and it fails. This step is not bureaucracy: it verifies that the test can detect the failure. If it passed while still red, the test would be badly written.
Green. The minimum code that satisfies it:
// src/domain/session.js (excerpt)
const MAX_TICKETS_PER_ORDER = 6;
sell(quantity) {
if (!Number.isInteger(quantity) || quantity < 1) {
throw new ValidationError('Quantity must be a positive integer');
}
if (quantity > MAX_TICKETS_PER_ORDER) {
throw new ValidationError(`Maximum ${MAX_TICKETS_PER_ORDER} tickets per order`);
}
// ...rest of the capacity logic
}Refactor. Now we improve things: we extract MAX_TICKETS_PER_ORDER into an exported constant so that the zod schema and the error message use the same number, and we add the error's details with the maximum allowed. The test protects us during the change.
And now the promised honesty:
| TDD helps a lot when... | TDD gets in the way when... |
|---|---|
| The business rule is clear and verifiable | You are exploring an API you do not know |
| You are fixing a bug (the test reproduces it first) | You are laying out a visual interface |
| You are designing a pure function, like the permission policy | The design will change three times this afternoon |
| You work on legacy code and want to pin down the current behavior | You are writing a single-use script |
It is not a religion. In this course we will write some tests before and many after, and in 09-06 we will use TDD strictly for one specific thing where it is beyond dispute: reproducing a bug before fixing it.
The tooling landscape for Node
| Tool | What it is | Strengths | Drawbacks |
|---|---|---|---|
| Mocha | Test runner | Mature, flexible, imposes neither assertions nor doubles, transparent configuration | You have to assemble the stack (Chai, Sinon) yourself |
| Chai | Assertion library | Very readable expect, clear failure messages, extensible |
An extra dependency |
| Sinon | Test doubles | Spies, stubs, fake clocks, the de facto standard | Easy to overuse |
| supertest | HTTP client for tests | Tests an Express app without a manual listen |
HTTP only |
| c8 / nyc | Coverage | c8 uses V8's native coverage, no instrumentation | A metric that is easy to misread |
| Jest | All in one | Runner, assertions, mocks and coverage together; snapshots | Opinionated; its transform complicates pure CommonJS/ESM |
| Vitest | Modern all in one | Blazing fast, Jest-compatible API, excellent ESM support | Born in the Vite world; less common in pure Node APIs |
node:test |
Native runner | Zero dependencies, shipped with Node, with --test and built-in coverage |
Younger API, smaller plugin ecosystem |
The native runner deserves an honest paragraph. Since Node 20 it is stable and perfectly usable:
// How it would look with the native runner (for reference, we will not use it)
const { test } = require('node:test');
const assert = require('node:assert/strict');
test('subtracts the tickets sold from the capacity', () => {
const session = createSession({ capacity: 400, sold: 100 });
session.sell(3);
assert.equal(session.available, 297);
});If you were starting a new project today, node:test is a defensible, dependency-free choice.
Why this course uses Mocha + Chai + Sinon + supertest: because it is the most widespread stack in the Node APIs you will meet in production, because it clearly separates the three responsibilities (run, assert, fake) and that teaches the concepts better, because its error reporting is explicit, and because the ideas transfer to any other stack in an afternoon. Having chosen the stack, we will not switch halfway through the module: mixing runners is a classic source of confusion.
The Escena Viva test plan
This is the route through the module:
| Lesson | What gets tested | Tools |
|---|---|---|
| 09-02 | Pure domain: Session, Event, SalesManager, and the full matrix in policy.js |
Mocha, Chai |
| 09-03 | Units with dependencies: currency-exchange.js, expiring tokens, repositories |
Sinon |
| 09-04 | The real API end to end: purchases, 401/403/404/409/422 rejections, concurrency | supertest |
| 09-05 | Coverage, thresholds, npm scripts, flaky tests, getting ready for CI | c8 |
| 09-06 | Debugging whatever fails: inspector, VS Code, stack traces, memory leaks | Node and DevTools |
The folder structure we will create, mirroring src/:
escena-viva/
├── src/
│ ├── domain/
│ ├── authorization/
│ ├── repositories/
│ └── ...
├── test/
│ ├── unit/
│ │ ├── session.test.js
│ │ ├── event.test.js
│ │ ├── policy.test.js
│ │ └── sales-manager.test.js
│ ├── integration/
│ │ ├── events.test.js
│ │ └── orders.test.js
│ └── helpers/
│ ├── factories.js
│ └── authentication.js
└── .mocharc.jsonThree conventions we will follow without exception:
- The
.test.jssuffix on every test file: it makes Mocha's search pattern trivial and distinguishes tests from helpers. - Mirroring the source code:
src/domain/session.jsis tested intest/unit/session.test.js. Finding the test for a file must never require a search. test/helpers/contains no tests, only shared utilities: data factories and the authentication helper that will finally deliver on the promise made in Module 8.
Common Mistakes and Tips
- Writing tests to push a number up. Coverage is an indicator, not a goal. We will cover it thoroughly in 09-05, but internalize it now: a test that runs code without checking anything raises coverage and protects you from nothing.
- Starting with end-to-end tests because "they test for real". You will end up with the ice cream cone and a suite nobody runs.
- Testing private methods. If you feel that urge, that logic usually wants to be a public function in another module.
#soldis tested throughavailableandsoldOut. - Tests that depend on ordering. "It works if you run the whole file but fails on its own" is the symptom. Every
itmust be runnable in isolation. - Reusing objects across tests to "avoid repetition". Sharing a mutable object between
itblocks is the number one cause of flaky tests. Use factories that return fresh objects. - Tip: start with the layer of highest risk and lowest cost. In Escena Viva that is
src/authorization/policy.js: pure functions, no dependencies, protecting the most sensitive part of the system. It can be tested in full in fifty lines. - Tip: treat test code as production code. It is read more often than it is written, and it has to be maintained.
Exercises
Exercise 1: classify by level
For each Escena Viva behavior, say at which level of the pyramid you would test it and why:
- The total of an order of 2 tickets at 2500 cents is 5000.
POST /api/orderswithout anAuthorizationheader returns 401.- An organizer from
org-bovedacannot editevt-001, which belongs toorg-almendra. - The
purchaseLimitmiddleware is applied before the orders controller. hashPassworduses cost 12.
Exercise 2: rewrite the names
Rewrite these test names following the rules from the lesson:
test buyTicketspolicy works fineerror 409does not blow up with quantity 0
Exercise 3: design a TDD cycle
Write (in pseudocode, without running it) the red-green-refactor cycle for this new rule: a session that has already started accepts no sales. State which test you would write first, what minimum code would make it pass and what you would refactor afterwards.
Solutions
Exercise 1
- Unit. It is a pure calculation in the domain; it needs neither HTTP nor a database.
- Integration. The 401 comes from the combination of route +
authenticatemiddleware +errorHandler. No unit test of the middleware guarantees that it is actually mounted on that route. - Unit, with
canManageEvent, because it is a pure function and we can walk the whole matrix cheaply and exhaustively. Plus one confirming integration test (a real 403 over HTTP), not all thirty. - Integration. Middleware order is only observable by running the real application: that is exactly what unit tests cannot see.
- None, or almost none. It is a configuration constant. Testing bcrypt's exact cost is testing the library; what does make sense is a test that
verifyPasswordaccepts the hash produced byhashPassword(behavior, not detail).
Exercise 2
creates an order with pending status and subtracts it from the session capacitylets the administrator manage any eventthrows StateConflict when the available seats are fewer than the ones requestedthrows ValidationError when the quantity is zero
Exercise 3
Red: a test named rejects the sale when the session has already started, which creates a session with startsAt in the past, calls sell(1) and expects a StateConflict. It fails because sell never looks at the date.
Green: add a check at the top of sell: if (new Date(this.startsAt) <= new Date()) { throw new StateConflict('The session has already started'); }.
Refactor: extract a get hasStarted() getter that encapsulates the comparison, reusable by the catalog so that past sessions are not listed, and inject the clock instead of calling new Date() directly. That last point is precisely what makes lesson 09-03 possible: code that depends on the real clock cannot be tested deterministically.
Conclusion
We have laid the foundations. We know why we test — to be able to change the code without fear, not just in case — how a test is structured with arrange, act and assert, which levels exist and why the pyramid has the shape it has, what deserves a test according to risk and what does not, which properties make a test good, how to tell a brittle test from a robust one, when TDD really helps and which tools the Node ecosystem offers.
We also have the plan: test/unit/, test/integration/, test/helpers/, *.test.js files, and the Mocha + Chai + Sinon + supertest stack.
What we do not have yet is a single runnable line. In the next lesson, Unit Testing with Mocha and Chai, we install the stack, finally fix that test script that has been failing on purpose for four modules, and write the complete unit suite for the Escena Viva domain, including that table of cases walking the permission matrix cell by cell that we have been promising since Module 8.
Node.js Course: From Beginner to Advanced
Module 1: Introduction to Node.js
- What Is Node.js?
- Installing and Setting Up the Environment
- Your First Node.js Program
- The Node.js REPL
- Modern JavaScript for Node.js
- The Course Project: the Escena Viva Platform
Module 2: Core Concepts
- Node.js Architecture
- The Event Loop
- Callbacks and Asynchronous Programming
- Promises and async/await
- Events and EventEmitter
- CommonJS Modules and require()
- ES Modules and Interoperability
Module 3: File System and I/O
- Reading and Writing Files
- The fs Module in Depth
- Cross-Platform Paths with the path Module
- Working with Streams
- Transform Streams and pipeline
- Buffers and Binary Data
Module 4: HTTP and Web Servers
- Creating a Simple HTTP Server
- Handling Requests and Responses
- Manual Routing
- Serving Static Files
- Receiving Data: Request Bodies and JSON
- Consuming External APIs from Node.js
Module 5: NPM and Package Management
- Introduction to NPM and package.json
- Installing and Using Packages
- Semantic Versioning and package-lock
- npm Scripts and Project Automation
- Creating and Publishing Packages
- Dependency Security and Maintenance
Module 6: The Express.js Framework
- Introduction to Express.js
- Setting Up an Express Application
- Routing in Express
- Middleware
- Essential Third-Party Middleware
- Input Data Validation
- Error Handling
Module 7: Databases and ORMs
- Introduction to Databases
- Using MongoDB with Mongoose
- CRUD Operations
- Relationships, Population and Advanced Queries
- Using SQL Databases with Sequelize
- Migrations, Transactions and Seed Data
Module 8: Authentication and Authorization
- Introduction to Authentication
- User Registration and Password Hashing
- Sessions and Cookies with Passport.js
- Authentication with JWT
- Role-Based Access Control
- API Security Best Practices
Module 9: Testing and Debugging
- Introduction to Testing
- Unit Testing with Mocha and Chai
- Test Doubles with Sinon
- Integration Testing
- Coverage and Test Automation
- Debugging Node.js Applications
Module 10: Advanced Topics
- The Cluster Module
- Worker Threads
- Caching and Job Queues with Redis
- Performance Optimization
- Building RESTful APIs
- GraphQL with Node.js
Module 11: Deployment and DevOps
- Configuration and Environment Variables
- Logging and Monitoring in Production
- Using PM2 for Process Management
- Packaging with Docker
- Deploying to Heroku and Other PaaS
- Continuous Integration and Deployment
