Escena Viva already has a test suite: unit tests for the domain and the permission policy, doubles for the clock and the network, and integration tests with supertest against a real database. npm test is green.

But two questions remain unanswered. The first: which parts of the code has no test ever executed? There may be a branch of the if in canChangeRole that is never walked, a catch that never fires, an exported function nobody calls. Coverage answers that. The second is more practical: what happens when the suite grows? Today it takes one second; with two hundred integration tests it will take five minutes, and a test that fails one time in eight will turn the team's work into a lottery.

Contents

  1. What coverage measures and what it does not
  2. Installing and reading c8
  3. Thresholds and exclusions
  4. Why 100 % is not the goal
  5. Where the risk lives in Escena Viva
  6. Organizing the npm scripts
  7. Fast, reliable tests
  8. Hunting down a flaky test
  9. Getting the project ready for continuous integration
  10. Git hooks

What coverage measures and what it does not

Coverage measures what percentage of the code was executed during the tests, and it breaks down into four metrics that are often confused:

Metric What it counts Example
Lines (lines) Lines of code executed Did line 42 run?
Statements (statements) Statements executed; one line may hold several const a = 1; const b = 2; is two
Branches (branches) The paths of each decision: if/else, ?:, &&, ??, case Was the if tested and the path without it?
Functions (functions) Functions invoked at least once Was canChangeRole ever called?

The one that really matters is branches. Look at this example, where 100 % line coverage coexists with half the branches untested:

function calculateTotal(lines, discount) {
  const gross = lines.reduce((sum, l) => sum + l.quantity * l.priceCents, 0);
  return discount ? Math.floor(gross * 0.9) : gross;
}

// A single test, with no discount:
expect(calculateTotal([{ quantity: 2, priceCents: 2500 }], false)).to.equal(5000);

That test executes every line: 100 % of lines, statements and functions. And 50 % of branches, because the discount branch is never walked. If Math.floor(gross * 0.9) were written wrong — say, gross * 0.09 — the report would still look nearly perfect. The same thing happens in our authorization policy:

function canManageEvent(user, event) {
  if (!user || !event) return false;
  if (user.role === 'administrator') return true;
  return user.role === 'organizer' && user.venueId === event.venueId;
}

A single test with an administrator runs the first two lines and leaves: high line coverage, low branch coverage, because we have not tested the null user, nor the organizer from the right venue, nor the one from the wrong venue, nor the attendee. And that is exactly where the security bug lives.

And what coverage does not measure, which is the most important point of this lesson: it does not measure whether you checked anything (running a line and asserting on its result are different things), it does not measure whether your assertions are correct (expect(true).to.be.true covers just the same), it does not measure the cases that are missing (if sell() never checks capacity, coverage will be 100 % and the oversale will be there), and it says nothing about design quality. Coverage is a hole detector, not a quality certificate: it answers "what did I leave out?", not "is this well tested?".

Installing and reading c8

We will use c8, which builds on V8's native coverage: it neither instruments nor transforms the code and works with CommonJS and ESM out of the box. Its classic alternative is nyc, Istanbul's heir, which instruments the code before running it; it is still perfectly valid and its report format is the same, but c8 is simpler to set up in a modern project.

npm install --save-dev c8
$ npx c8 mocha

  ... 74 passing (3.1s)

------------------------|---------|----------|---------|---------|-------------------
File                    | % Stmts | % Branch | % Funcs | % Lines | Uncovered Line #s
------------------------|---------|----------|---------|---------|-------------------
All files               |   81.42 |    72.09 |   84.61 |   81.42 |
 src/authorization      |  100.00 |   100.00 |  100.00 |  100.00 |
 src/domain             |   97.11 |    91.30 |   96.55 |   97.11 |
  sales-manager.js      |   93.75 |    83.33 |   90.00 |   93.75 | 48-51
 src/services           |   64.28 |    51.72 |   70.00 |   64.28 |
  tokens.js             |   72.41 |    55.00 |   75.00 |   72.41 | 88-97,112
  audit.js              |   31.03 |    16.66 |   33.33 |   31.03 | 12-38,45-52
 scripts                |    0.00 |     0.00 |    0.00 |    0.00 |
  seed.js               |    0.00 |     0.00 |    0.00 |    0.00 | 1-96
------------------------|---------|----------|---------|---------|-------------------

How to read it, from right to left. Uncovered Line #s is the most useful column: those are the lines no test executed, so go straight there. A % Branch far below % Stmts points to functions with many decisions and few scenarios tested: look at tokens.js, with 72 % of statements and 55 % of branches, which translated means we tested the happy path of tokens and almost none of the error paths. src/authorization at 100 % is the fruit of the case table from 09-02, and there is the lesson: pure functions get fully covered with very little effort. And scripts/seed.js at 0 % is not a problem, it is a utility script we are going to exclude.

To investigate properly, the HTML report (c8 --reporter=html mocha, which generates coverage/index.html) paints the colored source: lines that never ran come out in red and partially covered branches are marked with an I (if branch not taken) or an E (else branch not taken). It is the fastest way to see that an entire catch has never fired. Add coverage/ to .gitignore: it is a generated artifact, not code.

Thresholds and exclusions

A report nobody looks at is worthless. Thresholds turn coverage into a failure condition: if it drops below the minimum, the command exits with a non-zero code and CI goes red.

{
  "all": true,
  "include": ["src/**/*.js"],
  "exclude": ["src/config/index.js", "src/server.js", "src/db/sequelize.js"],
  "reporter": ["text", "html", "lcov"],
  "lines": 75,
  "branches": 65,
  "functions": 75,
  "check-coverage": true
}

The all option deserves a paragraph. Without it, c8 only reports on the files the tests loaded: if you create src/services/notifications.js and write no test for it, the file does not appear and the global coverage does not drop. With all: true it appears with a resounding 0 % and the average falls, which is exactly what you want to happen. The rest: include limits the measurement to src/, reporter produces text for the console, HTML for investigating and lcov for the CI server, and check-coverage makes the run fail when the minimums are not met.

On threshold policy there are two rules, and the second is the important one. They start low: set the threshold slightly below the current coverage, because an unreachable threshold gets ignored and an ignored threshold does not exist. And they never go down: when coverage rises to 85 %, raise the threshold to 82, and when someone proposes lowering it because "this urgent release leaves no time for testing", the answer is no. The threshold is a ratchet that only turns one way; the moment it can be lowered, it stops being a constraint and becomes a suggestion. A useful variant is per-folder thresholds, demanding more of critical code: c8 --include='src/domain/**' --include='src/authorization/**' --lines=95 --branches=90 --check-coverage mocha test/unit/**/*.test.js.

Excluding files is legitimate when measuring them provides no information, and a fraud when it is done to dress up the number.

Legitimate exclusions Why
sequelize-cli's migrations/ Generated code that runs once; it is verified by running it
scripts/seed.js A development tool, already exercised by the integration tests
src/config/index.js Reads environment variables; exercised by starting any test
src/server.js Calls listen and listens for signals; testing it would mean killing processes
Configuration files and test/ itself Configuration, not logic

These are illegitimate, and red flags in a code review: excluding src/controllers/ because "it is hard to test" (that is where the API logic lives), adding /* c8 ignore next */ over a branch that can perfectly well be tested just so the threshold passes, excluding the file that just caused a production incident, or excluding whole folders with no comment justifying it. When you use an inline exclusion, write down the reason: /* c8 ignore next 3 -- defensive branch: only happens if the engine changes its error format */.

Why 100 % is not the goal

Chasing 100 % systematically produces worse tests. This is the mechanism:

// Test written to push audit.js coverage from 31 % to 90 %
it('records to the audit log', async () => {
  await audit.recordAction('purchase', { userId: 'usr-001' });
  // ...and that is it. Not a single assertion.
});

This test runs the whole file, raises coverage by twenty points and can never fail unless the code throws. If recordAction stored the event with the wrong user, in the wrong collection or with last year's date, it would stay green. Worse still: now there is a test to maintain, one that takes time, and one that gives the false impression that auditing is tested, so the next person will not write the real test because "there is one already".

The last few points are also the most expensive. Going from 80 % to 90 % is usually useful work — error branches, edge cases, validations; going from 95 % to 100 % usually means inventing artificial scenarios to trigger defensive catch blocks that would only happen if Node itself broke.

If coverage does not tell you whether your tests are good, what does? Mutation testing, an idea as elegant as it is uncomfortable: a tool automatically modifies your production code — changes a >= to a >, an && to an ||, deletes a line, flips a boolean — and runs the suite again. If the tests fail, they have "killed the mutant": they detect that bug. If they pass with the mutated code, the mutant survived and you have code that is executed but not verified. The result is a mutation score, which measures something much closer to what we care about: not how much code runs, but how much of it is actually being watched. In Node the reference tool is Stryker Mutator, and its drawback is the cost, because it runs the suite once per mutant. The reasonable practice is to apply it only to critical code — in Escena Viva, src/domain/ and src/authorization/ — and periodically, not on every commit. We mention it because it is the conceptual antidote to blind faith in the percentage.

Where the risk lives in Escena Viva

Not all code deserves the same coverage. This is our reasoned split:

Layer Target Why
src/authorization/ 95-100 % Pure functions, dirt cheap to test, and a bug here is a security breach
src/domain/ 90-95 % Business rules; a bug sells tickets that do not exist. No dependencies
tokens.js, passwords.js 85-90 % Security, with many error branches you have to force with doubles
src/middleware/ 80 % Heavily exercised by integration; the rare branches deserve unit tests
src/controllers/ 70-80 % Covered mostly by integration; their own logic should be minimal
src/repositories/ 70 % Covered by integration against a real database
src/routes/ 60 % Almost declarative; what matters is that they answer, and integration measures that
scripts/, migrations/ Excluded Single-use tools

Reading this table gives you the central message of the lesson: coverage is not chased uniformly, it is aimed at where failure hurts. 100 % in src/routes/ protects you from nothing; 70 % in src/authorization/ is negligence.

Organizing the npm scripts

With the suite growing, a single npm test is not enough:

{
  "pretest": "npm run lint",
  "test": "mocha",
  "test:unit": "mocha test/unit/**/*.test.js",
  "test:integration": "mocha test/integration/**/*.test.js",
  "test:watch": "mocha --watch --reporter=min test/unit/**/*.test.js",
  "test:coverage": "c8 mocha",
  "test:ci": "c8 --check-coverage mocha --forbid-only",
  "check": "npm run lint && npm run test:coverage"
}

The decisions behind them. pretest runs automatically before test by npm convention, so if the linter fails the tests do not even start and a syntax error is caught in one second instead of thirty. test:unit and test:integration are kept apart because they are different in nature: unit tests take milliseconds and you launch them every couple of minutes, integration tests need a database. test:watch runs only the unit tests and with the min reporter: it is the script you leave open in a terminal while you code, and adding integration tests would make it useless. test:ci is the server variant, with thresholds and .only forbidden. And check is the "can I push this?" you run before a pull request.

Fast, reliable tests

Mocha can spread the files across several processes with mocha --parallel --jobs 4. For unit tests it is a clean win, because they share nothing. For integration tests against a shared database it is a disaster, and it is worth understanding why: each worker runs its own beforeEach with seedTestData(), which wipes the collections, so if worker A wipes the database while B is checking the capacity of ses-001-1, B fails for reasons unrelated to its own code and differently on every run. The possible ways out are not parallelizing integration (our choice: there are few of them and they take seconds), giving one database per worker using process.env.MOCHA_WORKER_ID in the name, or isolating by namespace so that each test uses its own identifiers and never deletes anything global.

A very profitable speed tweak, specific to this project: bcrypt at cost 12 takes around 300 ms on purpose, and in .env.test you can set BCRYPT_COST=4. With two conditions: that src/config/index.js reads it, and that a test exists verifying that the cost in production is 12; otherwise you end up deploying weak hashes.

Other tools: mocha --bail stops at the first failure, useful when you already know something is broken; --sort gives a deterministic alphabetical order; and --retries 2 retries failed tests, with a serious warning attached. Retrying hides flaky tests instead of fixing them, and a flaky test almost always signals a real concurrency or state bug in your code, not in the test; use it, if at all, only in end-to-end tests and with a deadline to investigate. The opposite is highly recommended: randomizing the order to uncover hidden dependencies. Mocha does not ship it, but running the files in reverse order from time to time (mocha $(ls test/unit/*.test.js | sort -r)) is invaluable information and free: if the suite fails in another order, you have shared state.

Hunting down a flaky test

A flaky test passes sometimes and fails other times without the code changing. It is the cancer of test suites: it erodes trust until the team re-runs CI out of habit, and at that moment the suite reports nothing at all.

Cause Symptom Diagnosis Fix
Real time Fails when the day, month or year changes; or on slow machines Date.now() or setTimeout appears in the tested code Sinon's fake clock (09-03)
Execution order Passes alone, fails in the suite (or vice versa) Run the file alone and the suite in reverse order Clean state in beforeEach, never in before
Shared state The second consecutive run fails Module variables, caches, data that survives A reset function, sinon.restore(), clean the database
Unawaited promises A failure in a later test, or a stray unhandledRejection Look for it without async, forgotten await Always await everything
Fixed ports or resources EADDRINUSE; fails only in CI A listen(3000) somewhere An ephemeral port (supertest already does this)

The hunting method, step by step:

# 1. Reproduce: run the suspect test fifty times
for i in $(seq 1 50); do
  npx mocha test/integration/orders.test.js --grep "does not oversell" || echo "FAILED on $i";
done

# 2. Does it fail alone or only in the suite?  3. Does it depend on order?
npx mocha test/integration/orders.test.js
npx mocha $(ls test/**/*.test.js | sort -r)

And the management rule, as important as the technique: a flaky test is either fixed or deleted, never ignored. If you cannot fix it now, mark it with .skip and a comment stating the reason and the date. A disabled, documented test is honest; one that fails at random is noise that teaches the team to ignore the color red.

Getting the project ready for continuous integration

In Module 11 we will set up the complete GitHub Actions workflow. Here we leave the project in a state where any server can run the tests without a human involved. There are six requirements.

1. A reproducible install. In CI you use npm ci, not npm install: it installs exactly what package-lock.json says, wiping node_modules first. That requires the lock file to be versioned and up to date.

2. Configuration through environment variables, not local files. .env.test is in .gitignore, so in CI the variables are injected; that is why Module 6 left src/config/index.js as the single place that reads process.env, and switching environments touches no code. Document the required variables in a .env.test.example that is versioned, with NODE_ENV, the database URIs with the _test suffix, JWT_SECRET=change-in-every-environment, BCRYPT_COST=4, LOG_LEVEL=silent and TZ=UTC, and no real secrets.

3. An ephemeral database. The server brings up Mongo and PostgreSQL as job services, empty, and our seed in beforeEach gets them ready. The requirement is that the suite must not assume preexisting data: if some test depends on a record you inserted by hand on your machine, it will fail in CI and not locally, which is the worst possible combination.

4. Exiting with a non-zero code, which is what a CI server reads as failure. mocha exits with the number of failed tests and c8 --check-coverage exits with 1 when the threshold is not met; check it with npm run test:ci; echo "exit code: $?".

5. No interaction and no processes left alive. No prompt, no waiting for a keypress; and with "exit": false in .mocharc.json, an unclosed connection hangs the job until the server kills it on time. That is why the previous lesson's afterAll with disconnect() is not cosmetic.

6. Time determinism. CI servers run in UTC, so any test that assumes your time zone will fail there: set TZ=UTC in .env.test and work with ISO dates, as we have been doing.

With this, the Module 11 workflow boils down to npm ci, bringing up the services and npm run test:ci. The hard work is this, not the YAML.

Git hooks

Thresholds and scripts are only useful if somebody runs them. A git hook runs a command before completing an operation, so the check stops depending on each person's discipline; the most useful one is pre-commit, which can block a commit if something fails.

npm install --save-dev husky lint-staged
npx husky init          # creates .husky/pre-commit

In .husky/pre-commit we put npx lint-staged followed by npm run test:unit, and in package.json the configuration of what to do with each file type:

{
  "lint-staged": {
    "*.js": ["eslint --fix", "prettier --write"],
    "*.{json,md}": ["prettier --write"]
  }
}

lint-staged is the key to the balance: it only processes the files in the staging area, not the whole project. Fixing the formatting of three files takes half a second; doing it for two hundred takes fifteen.

Moment What runs Time budget
pre-commit lint-staged + unit tests Under 10 seconds
pre-push The full suite including integration Under 2 minutes
CI (Module 11) Everything, plus coverage and a dependency audit Whatever it takes

And the warning: do not turn every commit into a five-minute wait. If the hook is slow, people learn to type git commit --no-verify and the hook ceases to exist. A fast hook that always runs protects infinitely more than an exhaustive one everybody dodges.

That idea deserves to be closed with the least technical and most decisive consequence in the module: a slow suite does not fail, it gets abandoned, and the sequence is always the same. The suite takes eight minutes and people stop running it locally, trusting CI. Because the feedback loop goes from seconds to minutes, commits become bigger and less frequent. When CI goes red, the change is already huge and locating the cause is expensive. A flaky test appears and, since re-running costs eight minutes, someone adds --retries. A real bug hides behind a retry and reaches production. The team's conclusion: "tests are useless". And with that suite, it is true. That is why the pyramid has the shape it has, why we separated test:unit from test:integration and why we lower the bcrypt cost in tests: a two-second suite gets run a hundred times a day; a two-minute one, twice.

Common Mistakes and Tips

  • Confusing coverage with quality. 95 % with assertion-free tests protects less than 70 % with demanding ones.
  • Not turning on all: true. Files with no tests at all vanish from the report and the average lies outrageously.
  • Lowering a threshold to get the release out. The ratchet breaks once and never climbs back.
  • Excluding what is hard to test. That is precisely what has to be tested; the difficulty usually comes from coupling (09-03).
  • Parallelizing integration against a shared database. Random failures that are impossible to reproduce.
  • Overusing --retries. It turns a real concurrency bug into tolerated noise.
  • Slow pre-commit hooks. They end up dodged with --no-verify.
  • Tip: when a file's branch coverage is far below its line coverage, open its HTML report; the red branches are usually exactly the untested catch blocks and error cases.
  • Tip: review the coverage of what changes in each pull request, not the global figure. The project average moves very slowly and hides the fact that the new module came in at 20 %.

Exercises

Exercise 1: find the untested branch

Run npm run test:coverage and open the HTML report. Locate the function in src/services/tokens.js with the lowest branch coverage, work out which scenario is missing and write the test that covers it. Hint: it is nearly always an error path that requires a Sinon double.

Exercise 2: the ratchet threshold

Configure .c8rc.json with the current thresholds minus three points. Then write three new tests for src/domain/sales-manager.js covering lines 48-51 from the sample report, measure again and raise the thresholds to the new level minus three. Document the date and the value so that the ratchet is visible.

Exercise 3: manufacture and hunt a flaky test

Write a flaky test on purpose: one that checks that the ticket code of the created order starts with EV-2026- using the system's real year. Then explain when it will fail, prove it with a fake clock set to 2027-01-01 and fix it so that it does not depend on the calendar.

Solutions

Exercise 1. The function with the worst branch coverage is usually rotateRefreshToken, because the happy path is tested and the error paths are not. The missing scenario is the reuse of an already rotated token:

it('revokes the whole family when an already rotated refresh is presented', async () => {
  const repository = {
    findToken: sinon.stub().resolves({ family: 'fam-001', usedAt: '2026-10-01T10:00:00.000Z' }),
    revokeFamily: sinon.stub().resolves(),
  };

  let caught = null;
  try {
    await rotateRefreshToken('old-token', { repository });
    expect.fail('Expected AuthenticationError for reuse');
  } catch (error) { caught = error; }

  expect(caught.appCode).to.equal('REFRESH_REUSED');
  expect(repository.revokeFamily.calledOnceWith('fam-001')).to.be.true;
});

The assertion about revokeFamily is the essential one: rejecting the request is not enough, the whole chain has to be invalidated because the token may have been stolen.

Exercise 2. .c8rc.json ends up with "lines": 78, "branches": 68, "functions": 78 if the current coverage is 81/72/84. The three tests for lines 48-51 of sales-manager.js — the path that does not emit low-capacity when the session was already below the threshold, the one that ignores sales of quantity zero and the one for the already sold-out session — are exactly the kind of defensive branch coverage brings to light. After covering them, the new value might be 84/76/85 and the thresholds would move to 81/73/82. And the ratchet is documented in the README, not in anybody's memory.

Exercise 3. It will fail on January 1, 2027 at 00:00 without anyone having touched the code, and also on a CI server with a shifted time zone if the run lands right at the turn of the year. The proof consists of wrapping the purchase with sinon.useFakeTimers(new Date('2027-01-01T00:00:01.000Z')) and checking that the assertion to.match(/^EV-2026-/) fails because EV-2027- arrives. The fix has two valid variants:

// Variant 1: assert the FORMAT, not the specific year
expect(body.tickets[0].code).to.match(/^EV-\d{4}-\d{6}$/);

// Variant 2: if the year matters, pin it with the fake clock
const clock = sinon.useFakeTimers(new Date('2026-11-05T12:00:00.000Z'));
// ...the purchase, and afterwards
expect(body.tickets[0].code).to.match(/^EV-2026-/);
clock.restore();

The general lesson: a test must never depend on a value that changes by itself. Either you assert the stable property (the format), or you control the source of variability (the clock).

Conclusion

We no longer just have tests: we know what they protect and what they do not. We understand the four coverage metrics and why branches is the one that matters, we have c8 wired up with a text and HTML report, thresholds in .c8rc.json that work as a ratchet, justified exclusions and a per-layer target split that aims the effort where failure hurts: the authorization policy and the domain, not the routes.

We also know that 100 % is a trap, that a test with no assertions raises the metric while protecting nothing, and that the serious answer to "are my tests any good?" is mutation testing. We have the npm scripts organized with pretest, test:watch for the short cycle and test:ci for the server. We know how to hunt a flaky test with a method, not with retries. And the project is ready for a continuous integration server to run it on its own: npm ci, environment variables, an ephemeral database and a non-zero exit code. The complete GitHub Actions workflow awaits us in Module 11.

One last piece remains. Tests tell you that something is broken; they hardly ever tell you why. When an integration test returns 500 instead of 201, when the process does not finish and you have no idea which resource stayed open, when the server's memory grows relentlessly over the week, you need another tool.

In the next lesson, Debugging Node.js Applications, we close the module: from a well-used console.log to the V8 inspector, conditional breakpoints, debugging your own Mocha tests from VS Code, reading asynchronous stack traces, diagnosing EADDRINUSE, E11000, unhandledRejection and the process that will not die, finding a memory leak by comparing heap snapshots, and a debugging method that ends where this module began: writing a test that reproduces the bug before fixing it.

Node.js Course: From Beginner to Advanced

Module 1: Introduction to Node.js

Module 2: Core Concepts

Module 3: File System and I/O

Module 4: HTTP and Web Servers

Module 5: NPM and Package Management

Module 6: The Express.js Framework

Module 7: Databases and ORMs

Module 8: Authentication and Authorization

Module 9: Testing and Debugging

Module 10: Advanced Topics

Module 11: Deployment and DevOps

Module 12: Real-World Projects

© Copyright 2026. All rights reserved