With configuration settled, Escena Viva now starts in any environment with the right values and dies immediately if a secret is missing. But a bigger problem remains: when the Festival de Jazz de Primavera opens its early sales and something fails at three in the morning, how do you find out? And how do you work out what exactly happened during Lucía's purchase, among the eleven thousand requests that arrived that minute? This lesson turns Escena Viva into an observable application: one that produces structured logs a machine can query, metrics somebody can watch on a dashboard, health probes the orchestrator can interrogate and alerts that fire when it is worth it. In Module 10 we learned to produce metrics — the event loop watcher, memory usage — but they stayed inside the process. Today we send them somewhere.

Contents

  1. Why console.log stops being enough
  2. pino: structured JSON logging
  3. The child logger and per-request traceability
  4. What gets logged and what NEVER gets logged
  5. Errors: the full stack trace stays server-side
  6. The three pillars of observability
  7. Metrics with prom-client
  8. Distributed tracing: when it pays off
  9. Health probes: live versus ready
  10. Alerts that are worth having
  11. Retention, cost and where the logs go

  1. Why console.log stops being enough

console.log is perfect while you develop. In production it has four serious defects. It has no levels, so you cannot turn the noise down in production and turn it up while investigating an incident: you either print everything or nothing. It has no structure: console.log('Purchase by ' + email) produces a sentence, and answering "how many purchases did Lucía fail?" means writing a brittle regular expression over gigabytes of text. It has no context: you do not know which request, which user or which worker that line came from, and with four clustered workers the lines of four simultaneous requests interleave. And it can be synchronous: when standard output is a file or a TTY, process.stdout.write blocks on Linux, so every console.log stalls the event loop — on a hot path that shows up in the very p99 we cared so much about in Module 10.

A structured log, by contrast, is one line of JSON per event:

{"level":30,"time":"2026-08-15T21:14:02.415Z","pid":4821,"requestId":"a3f1c2d4","route":"/api/purchases","status":201,"durationMs":83,"event":"purchase.completed","eventId":"evt-003","tickets":2,"totalCents":7000,"msg":"Purchase completed"}

The difference is not cosmetic: with that you can ask "give me every purchase of evt-003 that took longer than 500 ms and returned a 5xx, grouped by worker". With a sentence in prose, you cannot.

  1. pino: structured JSON logging

pino is the de facto standard logger in Node for one concrete reason: it is fast. It serializes to JSON with its own serializer and writes asynchronously through a transport on a separate thread, so logging stops competing with serving requests. You install it with npm install pino and npm install --save-dev pino-pretty. pino-pretty goes into devDependencies on purpose: pretty formatting is a development luxury; in production we write raw JSON and let the environment handle the rest.

// src/logging/logger.js
'use strict';

const pino = require('pino');
const { configuration } = require('../config/index.js');
const { redact } = require('./redact.js');

const options = {
  level: configuration.logLevel,
  // Fixed fields on EVERY line.
  base: { service: 'escena-viva', environment: configuration.nodeEnv, pid: process.pid },
  timestamp: pino.stdTimeFunctions.isoTime,   // ISO: readable in any viewer.
  serializers: {
    err: pino.stdSerializers.err,
    request: (req) => ({ method: req.method, route: req.path, ip: req.ip }),
    // Reuses the recursive redaction from M8: nothing sensitive leaves here.
    data: (value) => redact(value),
  },
  // Second layer of protection, in case somebody forgets the serializer.
  redact: {
    paths: ['req.headers.authorization', 'req.headers.cookie', 'password',
      '*.password', 'token', '*.token'],
    censor: '[REDACTED]',
  },
};

// pino-pretty ONLY in development: in production, raw JSON to stdout.
const transport =
  configuration.nodeEnv === 'development'
    ? pino.transport({ target: 'pino-pretty', options: { colorize: true } })
    : undefined;

module.exports = { logger: pino(options, transport) };

Three decisions deserve an explanation. base adds fixed fields to every line: service is indispensable when several services write to the same aggregator, and pid tells apart the cluster workers from Module 10. serializers are applied by field name, so writing logger.info({ data: body }) runs the object through redact before it goes out, reusing src/logging/redact.js from Module 8 without duplicating logic. And redact is the safety net: even if somebody forgets the serializer, those paths are replaced with [REDACTED]. Two layers, because one alone fails.

Level Value Use in Escena Viva In production?
trace 10 Every SQL query, every cache hit No, only while debugging
debug 20 Internal decisions: which capacity policy was applied Not by default
info 30 Business events: purchase completed, ticket issued Yes
warn 40 Something odd but recoverable: a Redis retry, a limit reached Yes
error 50 Failed request, queue job exhausted after retries Yes
fatal 60 The process cannot continue: the DB does not answer at startup Yes

The practical rule: info answers "what is the system doing?", error answers "what broke?", and the rest is for when you already know there is a problem. The LOG_LEVEL variable from the previous lesson lets you raise one instance to debug during an incident, without a deployment.

Replacing morgan and the stray console.error calls

Module 6 left morgan in src/middleware/http-logger.js. Morgan writes classic web-server text, which is now the odd one out in a JSON stream, so we replace it with our own middleware that closes the loop on every request:

// src/middleware/request-logger.js (rewritten on top of pino)
'use strict';

function createRequestLogger() {
  return function requestLogger(req, res, next) {
    const start = process.hrtime.bigint();
    // 'finish' fires once the response has been fully sent.
    res.on('finish', () => {
      const durationMs = Number(process.hrtime.bigint() - start) / 1e6;
      const level = res.statusCode >= 500 ? 'error' : res.statusCode >= 400 ? 'warn' : 'info';
      // The route PATTERN, not the concrete URL: otherwise you cannot group.
      const route = req.route ? req.baseUrl + req.route.path : req.path;
      req.log[level](
        { method: req.method, route, status: res.statusCode, durationMs },
        'request completed'
      );
    });
    next();
  };
}

module.exports = { createRequestLogger };

Details that matter: the route is logged as a pattern (/api/events/:eventId), because logging the URL with the identifier inside makes grouping impossible; the level depends on the status code, so a 500 shows up when you filter by error; and process.hrtime.bigint() gives nanosecond precision, unlike Date.now(). As for the stray console.error calls around the code, they all go: a grep -rn "console\." src/ in the CI pipeline (lesson 11-06) keeps them from coming back.

  1. The child logger and per-request traceability

Here comes the piece we have been promising since Module 6, when we wrote src/middleware/request-id.js to give every request a crypto.randomUUID(). Until now it did little; from here on it becomes the thread that stitches together every log line of a request, thanks to pino's child loggers: logger.child({ field: value }) returns a logger that adds those fields to all of its lines without paying for repeated serialization.

// src/middleware/request-id.js (extended)
'use strict';

const crypto = require('node:crypto');

function createRequestId({ logger }) {
  return function requestId(req, res, next) {
    // We honor the edge identifier if the proxy already set one.
    const incoming = req.get('x-request-id');
    req.requestId = incoming && incoming.length <= 64 ? incoming : crypto.randomUUID();
    res.setHeader('X-Request-Id', req.requestId);
    // Child logger: EVERYTHING logged from here on carries the identifier.
    req.log = logger.child({ requestId: req.requestId });
    next();
  };
}

module.exports = { createRequestId };

In src/app.js the order matters: requestId goes before any middleware that wants to log anything, and after trust proxy so that req.ip is correct. Controllers and services receive req.log and pass it downwards; in the BullMQ queue consumer from Module 10 the requestId travels as a job field, so the asynchronous work stays correlated with the request that originated it. Following a failed purchase. Marc tries to buy two tickets for the Festival de Jazz and sees a 500 with X-Request-Id: 9c4a.... He writes to support with that code and a single query is enough — grep '"requestId":"9c4a' logs.jsonl | jq -c '{time, level, event, msg}':

{"time":"...02.101Z","level":30,"event":"request.received","msg":"POST /api/purchases"}
{"time":"...02.118Z","level":30,"event":"capacity.checked","msg":"Available: 1189"}
{"time":"...02.140Z","level":30,"event":"reservation.created","msg":"Provisional reservation"}
{"time":"...02.402Z","level":50,"event":"payment.failed","msg":"Gateway returned 502"}
{"time":"...02.410Z","level":30,"event":"reservation.released","msg":"Reservation released"}
{"time":"...02.415Z","level":50,"event":"request.completed","msg":"request completed"}

Six lines and the whole story: capacity was fine, the reservation was created, the payment gateway failed with a 502 and — crucially — the reservation was released properly, so those two seats were not left locked. Without the requestId those six lines would be mixed in with the lines of another four hundred simultaneous requests. That is the value of traceability, and it is what justifies all the work in this lesson.

  1. What gets logged and what NEVER gets logged

Layer What to log Level
HTTP middleware Method, route pattern, status, duration info / warn / error
Authentication Successful or failed login, identifier and role info / warn
Controller Business event and its identifiers (evt-003, ticket code) info
Domain Nothing. The domain is pure: it returns results, it does not log —
Repository Slow queries (above a threshold), connection errors warn / error
Queue Job started, completed, failed, attempt number info / error
Startup and shutdown Port, environment, version, signal received, connections closed info

And now the non-negotiable list. Never log passwords, not even failed ones (they reveal patterns and typos of real passwords); full JWTs, refresh tokens, session cookies or Authorization headers; card numbers, CVVs or IBANs, neither complete nor "just the last four, it's harmless"; unnecessary personal data such as the full email address, the postal address or the phone number, instead of the user's identifier; nor entire request bodies "just in case", only the fields you actually care about. Module 8 already gave us src/logging/redact.js with redact and REDACTED_FIELDS, which walks objects recursively replacing sensitive fields. Wiring it in as a pino serializer applies that protection automatically to everything passing through the data field, and the combination of serializer plus redact covers both what you write on purpose and what slips through. Remember the legal frame too: in the European Union, a log containing personal data is processing of personal data, with its retention obligations and its right to erasure. Logging less is not just technical hygiene, it is less legal risk.

  1. Errors: the full stack trace stays server-side

The Module 6 error handler already distinguishes operational errors (expected: insufficient capacity, ticket not found) from bugs (unexpected). That distinction governs logging too:

// src/middleware/errors.js (excerpt from the final handler)
function errorHandler(error, req, res, next) {
  const status = STATUS_BY_CODE[error.appCode] ?? 500;

  if (status >= 500) {
    req.log.error({ err: error, status }, 'unhandled error');       // Full stack trace.
  } else {
    req.log.warn({ code: error.appCode, status }, error.message);   // No trace: it is noise.
  }
  // NEVER the stack trace to the client. Only the identifier to correlate.
  const code = error.appCode ?? 'INTERNAL_ERROR';
  const message = status >= 500 ? 'Internal server error' : error.message;
  res.status(status).json({ error: { code, message, requestId: req.requestId } });
}

Stack traces reveal filesystem paths, internal module names and sometimes values. The client receives only the requestId; with it, support finds the full trace in the aggregator, and the user gets something actionable without anything leaking. What is left is to close the two cases that escape Express, in src/server.js:

for (const event of ['uncaughtException', 'unhandledRejection']) {
  process.on(event, (reason) => {
    logger.fatal({ err: reason, event }, 'unhandled failure; terminating');
    process.exitCode = 1;
    shutdownGracefully();
  });
}

Yes, the process terminates. An uncaughtException leaves it in an undefined state: we log, shut down gracefully (Module 6) and let the supervisor — PM2 in 11-03, Docker in 11-04, the platform in 11-05 — start a healthy instance.

  1. The three pillars of observability

Pillar What it is Which question it answers Cost
Logs Discrete events with context "What exactly happened to this request?" High: grows with traffic
Metrics Numeric values aggregated over time "How is the system doing overall?" Low: constant size
Traces A request's journey across several services "Where did the time go?" Medium: usually sampled

The real sequence of an incident explains it better than any definition: an alert fires because a metric (p99 latency) has spiked; the dashboard shows it only affects /api/purchases; a trace reveals that 80% of the time is spent in a PostgreSQL query; and the logs for that request give you the specific query and the event (evt-003, naturally) that triggers it. Metrics to detect, traces to locate, logs to understand. All three, not one.

  1. Metrics with prom-client

Prometheus is the de facto standard: a server that periodically scrapes an HTTP endpoint on your application and stores time series; Grafana draws them, and prom-client is the library that exposes that endpoint from Node.

// src/observability/metrics.js
'use strict';

const promClient = require('prom-client');
const { createLoopWatcher } = require('./loop.js');

const register = new promClient.Registry();
const httpLabels = ['method', 'route', 'status'];

// Default process metrics: CPU, heap, descriptors, garbage collection... for free.
promClient.collectDefaultMetrics({ register, prefix: 'escenaviva_' });

const requestsTotal = new promClient.Counter({
  name: 'escenaviva_requests_total', help: 'HTTP requests served',
  labelNames: httpLabels, registers: [register],
});
const requestDuration = new promClient.Histogram({
  name: 'escenaviva_request_duration_seconds', help: 'Request duration',
  labelNames: httpLabels, registers: [register],
  // Buckets chosen from the M10 load data, not at random.
  buckets: [0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5],
});
const ticketsSold = new promClient.Counter({
  name: 'escenaviva_tickets_sold_total', help: 'Tickets sold',
  labelNames: ['eventId', 'venue'], registers: [register],
});
const queueSize = new promClient.Gauge({
  name: 'escenaviva_queue_pending', help: 'Pending jobs in BullMQ',
  labelNames: ['queue'], registers: [register],
});
const loopDelay = new promClient.Gauge({
  name: 'escenaviva_loop_delay_ms', help: 'Event loop lag',
  registers: [register],
});

// We reuse the M10 watcher: the metric was already produced, now it is published.
const watcher = createLoopWatcher({
  intervalMs: 1000, onMeasure: (delayMs) => loopDelay.set(delayMs),
});

module.exports = {
  register, requestsTotal, requestDuration, ticketsSold, queueSize, watcher,
};

We use all three metric types: the Counter only goes up and resets when the process restarts (total requests, tickets sold); the Gauge goes up and down (queued jobs, loop lag); and the Histogram spreads observations across buckets, which is where percentiles come from (request duration). There is a crucial detail about the histogram: percentiles cannot be averaged, so if each worker computed its own p99, the mean of those p99s would not be the system's p99. That is why the raw buckets are exported and Prometheus computes the percentile with histogram_quantile over all of them. A critical warning about labels: every combination of values creates a time series. Labeling by eventId is fine (three events); labeling by user id, by full URL or by requestId is a cardinality explosion that takes Prometheus down. It is the number one mistake with metrics.

// src/routes/metrics.js — 404 instead of 401: we do not even confirm it exists.
router.get('/metrics', async (req, res) => {
  const expected = configuration.observability.metricsToken;
  if (!expected || req.get('authorization') !== `Bearer ${expected}`) return res.status(404).end();
  res.set('Content-Type', register.contentType);
  res.end(await register.metrics());
});

It is protected because /metrics reveals internal topology, business volume and versions, and it is excluded from rate limiting because Prometheus scrapes every 15 seconds. With the Module 10 cluster there is a subtlety: each worker has its own in-memory registry and Prometheus scrapes a single port, so you either use prom-client's AggregatorRegistry in the primary or — more simply — expose one metrics port per worker and let them all be discovered. Neither option is magic; what matters is knowing the problem exists. Prometheus stores the series and queries them with PromQL; its minimal configuration is a list of targets with their interval and their credentials. Grafana connects to Prometheus and draws. The minimal Escena Viva dashboard has four charts: requests per second by status, p50/p95/p99 latency, event loop lag and pending queue jobs. If you can only watch one screen during the Festival de Jazz opening night, make it that one.

  1. Distributed tracing: when it pays off

A trace follows a request across several services, chaining spans (segments with a start, an end and a parent). OpenTelemetry is the open standard: it automatically instruments Express, PostgreSQL, Redis and outbound HTTP, and exports to Jaeger, Tempo or whichever service you use. The instrumentation is loaded before any other module (node --require ./src/observability/tracing.js src/server.js) because it needs to wrap them as they load. Now the honest part: Escena Viva does not need it yet. It is one API, one relational database, one document database, Redis and a queue consumer. With the requestId propagated to the logs and to the queue jobs you get 90% of the benefit at almost zero cost. Traces pay off when there are many services, when each request crosses five or more hops, or when the problem is "it is slow" and nobody knows which hop is to blame. When Escena Viva splits into catalog, sales and notification services, you instrument it; before then, it is complexity with no return.

  1. Health probes: live versus ready

Every supervisor (PM2, Docker, the PaaS, Kubernetes) needs to ask the application how it is doing. And there are two distinct questions that get confused constantly:

Probe Question If it fails Checks
/health/live Is the process alive and responding? The process is restarted Only that the event loop responds
/health/ready Can it serve traffic right now? Traffic is withdrawn, no restart Dependencies: DB, Redis

The classic mistake — and it is a classic because it takes whole systems down — is making the liveness probe query the database. Picture it: PostgreSQL saturates during the Festival de Jazz opening, the probe fails on all four instances, the supervisor restarts them all at once and, on startup, all four open their pools at once against an already saturated database. It fails again: cascading restarts, and a latency incident becomes a total outage. The rule: liveness only looks at the process; readiness looks at the dependencies.

// src/routes/health.js
'use strict';

function createHealthRoutes({ checkPostgres, checkMongo, checkRedis, logger }) {
  const router = express.Router();
  const startedAt = Date.now();
  // Graceful shutdown sets this to false BEFORE it starts closing.
  const state = { acceptingTraffic: true };
  // Timeout: a hung probe is worse than one that fails.
  const withLimit = (promise, ms) =>
    Promise.race([promise, new Promise((_, r) => setTimeout(() => r(new Error('timed out')), ms))]);

  // LIVENESS: if this answers, the event loop is running. Nothing else.
  router.get('/health/live', (req, res) => {
    res.json({ status: 'alive', uptimeSeconds: Math.floor((Date.now() - startedAt) / 1000) });
  });

  // READINESS: checks dependencies. allSettled reports on ALL of them,
  // not just the first one that fails.
  router.get('/health/ready', async (req, res) => {
    if (!state.acceptingTraffic) return res.status(503).json({ status: 'shutting down' });
    const names = ['postgres', 'mongo', 'redis'];
    const results = await Promise.allSettled([
      withLimit(checkPostgres(), 1000),
      withLimit(checkMongo(), 1000),
      withLimit(checkRedis(), 500),
    ]);
    const detail = Object.fromEntries(
      results.map((r, i) => [names[i], r.status === 'fulfilled'])
    );
    const ready = Object.values(detail).every(Boolean);
    if (!ready) logger.warn({ detail }, 'readiness probe failed');
    res.status(ready ? 200 : 503).json({ status: ready ? 'ready' : 'not ready', detail });
  });

  return { router, state };
}

module.exports = { createHealthRoutes };

Three deliberate decisions: the checks carry a timeout, because a probe that hangs is worse than one that fails; Promise.allSettled reports on every dependency; and state.acceptingTraffic is set to false on SIGTERM, before closing anything, so the load balancer stops sending traffic during the drain window. That detail is what turns the graceful shutdown from Module 6 into an error-free deployment. Both routes are excluded from rate limiting and require no authentication, but they do not expose internals either: no database versions, no connection strings.

  1. Alerts that are worth having

The golden rule of reliability engineering: alert on symptoms, not on causes. A symptom is something the user notices; a cause is a hypothesis about why. If you alert on "CPU at 90%", you will be woken on nights when the CPU was high and everything was perfect, and you will not be woken on the night everything fell over for a different reason.

Alert Symptom or cause Worth it?
p99 of /api/purchases > 2 s for 5 min Symptom Yes
5xx rate > 1% for 5 min Symptom Yes
Ticket queue > 500 jobs and growing for 10 min Symptom Yes
No tickets sold in 30 min during selling hours Symptom Yes
CPU > 80% or memory > 70% Cause Not as an alert; yes as a chart
Loop lag > 200 ms for 5 min Borderline Yes: it correlates very well with real pain

Two nuances separate a useful alerting system from one everybody ignores. Every alert carries a duration: "over 2 s" fires on an irrelevant spike, "over 2 s sustained for 5 minutes" fires when there is a real problem. And every alert that fires must require a human action now; if it does not, it is not an alert, it is a chart. An alert nobody attends to is worse than no alert at all, because it teaches the team to ignore notifications, and the day a real one fires it will be ignored too. This has a name — alert fatigue — and it has caused more serious incidents than any bug. Every alert should link to a runbook: what to look at, what to check, whom to escalate to; if you cannot write one, the alert probably is not worth having.

  1. Retention, cost and where the logs go

Logs cost money: at info level, an API serving 500 requests per second generates on the order of 100 GB a month from the request-completion line alone, and managed services charge per gigabyte ingested and per day retained.

Type Reasonable retention Why
Application logs (info+) 14-30 days Covers the investigation of a recent incident
Error logs 90 days To spot patterns that repeat
Audit trail (Module 8) 1-7 years Legal requirement; it goes to the DB, not the aggregator
Metrics 13 months To compare this Festival de Jazz with last year's

And so we reach the last point, which closes the circle with lesson 11-01. The twelve factors, factor XI: treat logs as event streams. The application does not write files, does not rotate them and does not know where they go: it writes to stdout and its responsibility ends there. Who collects that output depends on the environment — PM2 redirects it to files (11-03), Docker captures it with its logging driver (11-04), the PaaS forwards it to its aggregator (11-05), Kubernetes picks it up with an agent — and writing files from the application forces you to solve rotation, permissions and disk space, on top of losing the logs when an ephemeral container dies. In code: pino(options) writes to stdout and that is correct; pino(pino.destination('/var/log/escena-viva.log')) makes the application choose the destination, which is exactly what we do not want. One line of difference, and a mountain of operational work you avoid.

Common Mistakes and Tips

  • Leaving pino-pretty enabled in production. It is slow and it produces text the aggregator cannot index.
  • Logging the full URL as a groupable field. /api/events/evt-003 instead of /api/events/:eventId prevents grouping and, in metrics, explodes cardinality.
  • A liveness probe that queries the database. Cascading restarts guaranteed under load.
  • Returning the stack trace to the client. It leaks paths and internal structure. Return only the requestId.
  • Logging entire request bodies. Huge volume plus the risk of personal data and secrets.
  • Tip: include the requestId in the error messages the user sees. Letting support say "read me the code on your screen" turns an hour-long investigation into a ten-second query.
  • Tip: log an event at startup with the version, the commit and the environment. Knowing which version was running during an incident is the first question you will ask.

Exercises

Exercise 1 — Slow query logging

Add a wrapper to src/repositories/ that measures the duration of every query and logs at warn those exceeding 200 ms, with the repository name, the method and the duration, using req.log when it exists so the requestId is preserved.

Exercise 2 — Business metric and probe tests

Instrument the purchase use case so it increments ticketsSold with the eventId and venue labels, and write the PromQL query that answers: tickets sold per minute for the Festival de Jazz de Primavera. Then write an integration test with supertest verifying that /health/live returns 200 even when the dependency checks fail, that /health/ready returns 503 when checkRedis rejects, and that it also returns 503 when state.acceptingTraffic is false.

Solutions

Exercise 1. A generic wrapper saves you from touching every method:

// src/repositories/instrument.js
'use strict';

const { logger } = require('../logging/logger.js');
const THRESHOLD_MS = 200;

function instrumentRepository(name, repository) {
  return new Proxy(repository, {
    get(target, method) {
      const value = target[method];
      if (typeof value !== 'function') return value;
      return async function (...args) {
        const start = process.hrtime.bigint();
        try {
          return await value.apply(target, args);
        } finally {
          const durationMs = Number(process.hrtime.bigint() - start) / 1e6;
          if (durationMs > THRESHOLD_MS) {
            (this?.log ?? logger).warn(
              { repository: name, method: String(method), durationMs }, 'slow query');
          }
        }
      };
    },
  });
}

module.exports = { instrumentRepository };

Exercise 2. In the use case, after confirming the purchase inside the transaction, ticketsSold.inc({ eventId: purchase.eventId, venue: purchase.venue }, purchase.tickets.length). The query uses rate over five minutes and multiplies by 60 to go from seconds to minutes: sum by (venue) (rate(escenaviva_tickets_sold_total{eventId="evt-003"}[5m])) * 60. rate handles counter resets correctly when a worker is recycled, which is precisely why you do not use bare increase or a raw difference. And for the probes, by injecting fake checkers into the application factory:

const request = require('supertest');
const { expect } = require('chai');
const { createApplication } = require('../../src/app.js');

describe('health probes', () => {
  const ok = () => Promise.resolve(true);
  const down = () => Promise.reject(new Error('down'));
  const create = (redis, rest = ok) =>
    createApplication({ checkPostgres: rest, checkMongo: rest, checkRedis: redis });

  it('live answers 200 even when the dependencies fail', async () => {
    await request(create(down, down).app).get('/health/live').expect(200);
  });
  it('ready answers 503 when Redis fails', async () => {
    const response = await request(create(down).app).get('/health/ready').expect(503);
    expect(response.body.detail).to.deep.include({ redis: false, postgres: true });
  });
  it('ready answers 503 during shutdown', async () => {
    const { app, healthState } = create(ok);
    healthState.acceptingTraffic = false;
    await request(app).get('/health/ready').expect(503);
  });
});

Conclusion

Escena Viva no longer talks to itself. It writes structured JSON logs with pino, every line tagged with the requestId we have been carrying since Module 6 — so a failed purchase can be reconstructed in full with one query — run through the Module 8 redaction so that neither a password nor a token ends up in the aggregator. It publishes metrics on a protected /metrics: requests, latency with real percentiles, tickets sold, queue size and the event loop lag Module 10 taught us to measure. It exposes /health/live and /health/ready, deliberately different, ready for a supervisor to interrogate. And it all goes out through stdout, because deciding where the logs end up is not the application's business.

That is exactly what we still lack: somebody outside collecting the output, watching the process and bringing it back up if it dies. Right now Escena Viva is started by hand with npm start in a terminal, and if you close the SSH session it dies with it. In the next lesson, Using PM2 for Process Management, we put a supervisor in front: an ecosystem file with the project's two applications, cluster mode without maintaining our own src/cluster.js, zero-downtime reloads built on graceful shutdown, protection against restart loops and automatic startup when the machine boots.

Node.js Course: From Beginner to Advanced

Module 1: Introduction to Node.js

Module 2: Core Concepts

Module 3: File System and I/O

Module 4: HTTP and Web Servers

Module 5: NPM and Package Management

Module 6: The Express.js Framework

Module 7: Databases and ORMs

Module 8: Authentication and Authorization

Module 9: Testing and Debugging

Module 10: Advanced Topics

Module 11: Deployment and DevOps

Module 12: Real-World Projects

© Copyright 2026. All rights reserved