POST /api/orders still does recordOrder(req.body) with whatever arrives. Helmet, CORS and the rate limiter do not look at the content: a body with quantity: -5, a sessionId that is an array or a 40,000-character e-mail address sails through all of them without breaking a sweat. This is the hole we close now. Validation is one of the things Express does not do for you (we said so in 06-01, table included) and it is also one of the ones with the biggest consequences when it is missing. In this lesson we decide where it lives, pick a tool, write Escena Viva's schemas and build a generic middleware that sits on the route and hands over data already clean and converted.

Contents

  1. Why you never trust the client
  2. What gets validated and where validation must live
  3. Validation by hand and why it becomes unmaintainable
  4. zod: the validator of choice
  5. Escena Viva's schemas
  6. The generic validate(schema, source) middleware
  7. Coercion and normalization
  8. The useful error response
  9. Less obvious uses: responses and configuration
  10. Sanitization versus validation

  1. Why you never trust the client

public/app.js validates the form before sending: it checks that the quantity is between 1 and 6, that the e-mail address has an at sign and that a session is selected. And even so, the server must validate everything again. The reasons, from the most obvious to the least:

  1. The client is not your form. Anyone can call the API with curl, with Postman or with a script: your HTML is a suggestion, not a control.
  2. The client can be modified. The browser developer tools let you change an input's max in two seconds.
  3. There may be more clients. A mobile app, a box-office dashboard, an integration with a promoter. Each one with its own bugs.
  4. Intermediaries alter data. Proxies, redirects, retries with a truncated body.
  5. The client validates for experience; the server validates for integrity. The first avoids frustration, the second avoids corruption.

The classic formulation: client-side validation is courtesy; server-side validation is obligation. And it is not only about attackers: a front-end deployment that sends quantity as a string instead of a number because of a typing bug causes exactly the same wreckage, with no bad intent.

  1. What gets validated and where validation must live

Source What arrives Typical risk
req.body The order's JSON body Wrong types, missing fields, extra fields
req.params :id, :sessionId Arbitrary format, enormous strings, control characters
req.query Filters and pagination A perPage=999999 that takes the response down
req.headers X-Sale-Channel, Accept Log injection, outsized values

And the most important question: where do you put that validation? The journey is Client → validate() → Controller → Service → Domain, and there are two very different rejection points: the validate() middleware returns 400 when the shape is invalid, and the domain returns 409 when there is not enough capacity. They are two levels with different responsibilities, and confusing them is a frequent design mistake:

Input validation (the edge) Domain invariants
Question it answers Does the data have the right shape? Is this operation valid right now?
Example quantity is an integer between 1 and 6 Session ses-001-1 has 3 tickets available
Where it lives src/schemas/ + the validate middleware src/domain/session.js, sales-manager.js
Depends on state / knows HTTP No / yes, it is the edge Yes / no, never
Resulting HTTP status 400 or 422 409

The rule: the edge guarantees the shape; the domain guarantees the rules. Session.sell() must not check whether quantity is a number —that is already guaranteed by the time it gets there— but it must check whether there is capacity, because that depends on state and can change between the validation and the sale. And the other way around: if you delete the validation middleware, the domain must not break catastrophically, but it is not its job either to produce friendly messages for an HTTP client.

  1. Validation by hand and why it becomes unmaintainable

Remember how order-routes.js ended up in Module 4:

// Extract from module 4. It worked. And it would grow out of control.
function validateOrder(data) {
  if (typeof data !== 'object' || data === null || Array.isArray(data)) {
    throw Object.assign(new Error('The body must be an object'), { appCode: 'INVALID_JSON' });
  }
  if (typeof data.sessionId !== 'string' || !/^ses-\d{3}-\d$/.test(data.sessionId)) {
    throw Object.assign(new Error('invalid sessionId'), { appCode: 'INVALID_PARAMETER' });
  }
  if (!Number.isInteger(data.quantity) || data.quantity < 1) {
    throw Object.assign(new Error('invalid quantity'), { appCode: 'INVALID_QUANTITY' });
  }
  if (data.quantity > MAX_TICKETS_PER_ORDER) {
    throw Object.assign(new Error('6 tickets maximum'), { appCode: 'ORDER_LIMIT_EXCEEDED' });
  }
  // ...and now the email, and the channel, and the optional fields, and the discount...
}

Five concrete problems, and all of them get worse over time:

  1. It stops at the first error. The client fixes sessionId, resends, and discovers that quantity was wrong too: three round trips for three errors.
  2. It does not return the converted result. You validate data.quantity but you keep using the original, unnormalized object.
  3. Extra fields get through. If the client sends { quantity: 2, isAdministrator: true }, that field reaches the domain untouched.
  4. It is impossible to reuse. The same order validated from a message queue or an import script needs copy and paste.
  5. It cannot be documented automatically. There is no description of the expected shape; only imperative code you have to read in full.

Fifty lines for four fields: with twenty endpoints, this is half the project.

  1. zod: the validator of choice

zod flips the approach: instead of writing checks, you describe the shape of the data, and the library derives validation, conversion and messages from it. It is installed with npm install zod, and npm view zod license dependencies confirms what you were looking for in Module 5: MIT and zero dependencies.

A short and honest comparison:

zod express-validator Joi
Coupling to Express None High: it is middleware None
Style Declarative schema A chain of checks per field Declarative schema
Output and reuse outside HTTP A converted, pruned object; yes It modifies req and exposes validationResult; no A converted object; yes
Type inference Excellent (with TypeScript, for free) Poor Good
Dependencies and learning curve Zero, gentle curve Several, gentle curve Several, medium curve

We choose zod for three reasons that fit the course: zero dependencies (the Module 5 criterion), independence from Express (the same schema serves the API, an import script and the Module 9 tests) and because the schema is a single source of truth that documents the input better than any comment. The essentials of its API: you declare an object with z.object({ name: z.string().min(1).max(80), age: z.number().int().positive(), email: z.email(), channel: z.enum([...]), notes: z.string().optional(), active: z.boolean().default(true) }) and you run it with schema.safeParse(data) —which returns { success, data } or { success, error }— or with schema.parse(data), which returns the data or throws a ZodError.

A key detail: z.object() discards undeclared keys by default. { quantity: 2, isAdministrator: true } comes out of the validator as { quantity: 2 }, so problem 3 from the previous section is solved out of the box. And if you prefer to reject rather than discard, .strict() makes the extra key an error.

  1. Escena Viva's schemas

// src/schemas/order.js
const { z } = require('zod');
const MAX_TICKETS_PER_ORDER = 6;

// Reusable pieces. Session identifier: ses-NNN-M (ses-001-1...).
const sessionId = z
  .string({ error: 'sessionId must be a string' })
  .trim()
  .regex(/^ses-\d{3}-\d$/, 'sessionId must have the format ses-NNN-M');
const eventId = z.string().trim().regex(/^evt-\d{3}$/, 'The event must be evt-NNN');
const ticketQuantity = z.coerce
  .number({ error: 'quantity must be a number' })
  .int('quantity must be an integer')
  .min(1, 'You must order at least 1 ticket')
  .max(MAX_TICKETS_PER_ORDER, `${MAX_TICKETS_PER_ORDER} tickets per order maximum`);
const buyerEmail = z
  .email('The buyer e-mail address is not valid')
  .max(254, 'The e-mail address is too long')
  .transform((value) => value.trim().toLowerCase());
const saleChannel = z.enum(['web', 'box-office', 'phone'], { error: 'invalid channel' });

/** Body of POST /api/orders */
const createOrderSchema = z.object({
  sessionId,
  quantity: ticketQuantity,
  email: buyerEmail,
  channel: saleChannel.default('web'),
  comment: z.string().trim().max(280).optional(),
});
/** Params of /api/events/:id and query of GET /api/events */
const eventParamsSchema = z.object({ id: eventId });
const eventsQuerySchema = z.object({
  venue: z.enum(['Teatro Almendra', 'Sala Boveda', 'Auditorio Ribera']).optional(),
  soldOut: z.enum(['true', 'false']).transform((v) => v === 'true').optional(),
  page: z.coerce.number().int().min(1).default(1),
  perPage: z.coerce.number().int().min(1).max(50).default(10),
});
module.exports = { createOrderSchema, eventParamsSchema, eventsQuerySchema };

Compare it with the fifty imperative lines from section 3: there are more rules here, in less space, and they read like a specification. And MAX_TICKETS_PER_ORDER is still the same constant from Module 4: the rule has not changed, only the place where it is expressed.

  1. The generic validate(schema, source) middleware

A single configurable middleware that goes on any route. It is a factory, like the ones in 06-04.

// src/middleware/validate.js
const { ValidationError } = require('../errors.js'); // 06-07
const VALID_SOURCES = ['body', 'params', 'query', 'headers'];

// Returns a middleware that validates one part of the request against a schema.
// The converted, pruned result is left on req.validatedData[source].
function validate(schema, source = 'body') {
  // A programming error: it is caught at startup, not in production.
  if (!VALID_SOURCES.includes(source)) {
    throw new Error(`Unsupported validation source: ${source}`);
  }
  return function validateRequest(req, res, next) {
    const result = schema.safeParse(req[source]);
    if (!result.success) {
      // We convert zod's issues into our own details format.
      const details = result.error.issues.map((i) => ({
        field: i.path.join('.') || source, reason: i.message, type: i.code,
      }));
      next(new ValidationError('The data you sent is not valid', details));
      return;
    }
    // IMPORTANT: we do not reassign req.query (read-only in Express 5)
    // or req.params. We store the result in a property of our own.
    req.validatedData = req.validatedData ?? {};
    req.validatedData[source] = result.data;
    next();
  };
}
module.exports = { validate, VALID_SOURCES };

Usage in the routes, and a controller with no checks left:

orderRoutes.post('/', validate(createOrderSchema, 'body'), createOrder);
eventRoutes.get('/', validate(eventsQuerySchema, 'query'), listEvents);
eventRoutes.get('/:id', validate(eventParamsSchema, 'params'), getEvent);

// src/controllers/orders.js — the data already has the right shape:
// integers are integers, the e-mail is lowercased and there are no extra fields.
async function createOrder(req, res) {
  const order = await recordOrder(req.validatedData.body);
  res.status(201).location(`/api/orders/${order.id}`).json(order.toJSON());
}

Why req.query is not reassigned

In Express 4 the usual pattern was req.query = result.data, and it worked. In Express 5 req.query is a getter with no setter, and that line throws TypeError: Cannot set property query of #<IncomingMessage> which has only a getter; if you migrate old code, it is one of the first errors you will see. Storing on req.validatedData is not just a way of dodging the limitation: it is better design, because it makes it explicit in the controller that the data it uses has been through a schema, and it keeps what came from the client intact in case the logger needs to compare it.

  1. Coercion and normalization

Everything that arrives over HTTP is text. req.params and req.query are always strings —in GET /api/events?page=2, req.query.page is '2', not 2—; only req.body with Content-Type: application/json preserves types, and only if the client sent them properly. That is why converting is part of validating, not a later step:

// Coercion: accepts '2' and 2, always returns the number 2.
page: z.coerce.number().int().min(1).default(1),
// Normalization: trims whitespace and lowercases.
email: z.email().transform((value) => value.trim().toLowerCase()),
// Boolean coercion from the query, which never carries booleans.
soldOut: z.enum(['true', 'false']).transform((value) => value === 'true').optional(),
Operation What it does Example in Escena Viva
Coercion Converts the type '2' → 2 in quantity
Normalization Unifies equivalent representations ' [email protected] ' → '[email protected]'
Default value and pruning Fills in what is missing and discards what is undeclared Missing channel → 'web'; isAdministrator disappears

Without lowercasing the e-mail address, [email protected] and [email protected] are two different buyers: duplicates in master data almost always come from a normalization that was missing.

Careful with excessive coercion. z.coerce.number() accepts '' as 0 and true as 1, because it uses Number() underneath. If that does not suit you, be explicit: z.string().regex(/^\d+$/).transform(Number). Friendly coercion in the query (where everything is necessarily text), strict coercion in the JSON body (where the client could have sent the right type).

  1. The useful error response

We keep the format set in Module 4 and add details to it:

{
  "error": {
    "code": "INVALID_DATA",
    "message": "The data you sent is not valid",
    "status": 400,
    "details": [
      { "field": "quantity", "reason": "6 tickets per order maximum", "type": "too_big" },
      { "field": "email", "reason": "The e-mail address is not valid", "type": "invalid_format" }
    ],
    "requestId": "9f2a1c48-3c7e-4a1b-9c62-1d0f8b4a77e1"
  }
}

What makes this response useful: it gives all the errors at once (one round trip instead of three), the exact field with a full path in nested structures (tickets.0.quantity), a readable reason written by you in the schema, a stable type (too_big, invalid_type) that a client can handle programmatically without depending on the language, and the request identifier from 06-04 so it can be cross-referenced with the server logs.

And what must never show up:

Do not reveal Why
The exception's stack trace It gives away file system paths and library versions
The received value as-is It may contain passwords or personal data, and it ends up in the logs
Table, column, file or internal query names A free map of your infrastructure and help for the next attempt
Whether an e-mail address exists in the system It allows user enumeration (critical in Module 8)

400 versus 422

Both are legitimate and the debate is open. Our policy, consistent with the STATUS_BY_CODE table from Module 4:

Situation Domain code Status
Malformed JSON (it cannot even be parsed) INVALID_JSON 400
A field is missing or has the wrong type INVALID_DATA 400
Correct format but an input rule violated (more than 6 tickets) ORDER_LIMIT_EXCEEDED 422
Correct format but the state does not allow it (not enough capacity) INSUFFICIENT_CAPACITY 409

What matters is not which one you choose, but that it is consistent across the whole API and documented.

  1. Less obvious uses: responses and configuration

Schemas are not only for what comes in.

Validating the startup configuration

In 06-02 you wrote functions by hand (required, integer, list); a schema does the same in less space and with better messages:

// src/config/schema.js
const configurationSchema = z.object({
  NODE_ENV: z.enum(['development', 'test', 'production']).default('development'),
  PORT: z.coerce.number().int().min(1).max(65535).default(3000),
  BODY_LIMIT: z.string().regex(/^\d+(kb|mb)$/i).default('100kb'),
  TRUST_PROXY: z.coerce.number().int().min(0).max(10).default(0),
  ALLOWED_ORIGINS: z.string().default('')
    .transform((v) => v.split(',').map((o) => o.trim()).filter(Boolean)),
});

The same "fail fast" principle, with a single place describing everything the application needs in order to start.

Validating an external service's response

src/services/currency-exchange.js does a fetch against an API you do not control: the fact that it returns { rates: { USD: 1.08 } } today does not guarantee that tomorrow it will not return { data: [...] } after a version change.

// src/services/currency-exchange.js (fragment)
const ratesResponseSchema = z.object({
  base: z.literal('EUR'),
  rates: z.record(z.string().length(3), z.number().positive()),
});

async function getExchangeRate(currency) {
  const response = await fetch(RATES_URL, { signal: AbortSignal.timeout(3000) });
  const result = ratesResponseSchema.safeParse(await response.json());
  if (!result.success) {
    // Better to fail with a clear code than to propagate undefined across the system.
    const message = 'The currency API returned an unexpected format';
    throw Object.assign(new Error(message), { appCode: 'EXTERNAL_SERVICE_DOWN' });
  }
  return result.data.rates[currency];
}

The cost is minimal and it avoids the worst kind of failure: the one that does not blow up where it happens, but three layers down with an incomprehensible undefined.

  1. Sanitization versus validation

They get confused constantly, but they are different things:

Validation Sanitization
Question Is this piece of data acceptable? How do I make it harmless in this context?
Timing and result On the way in; rejection with a 400 On the way out towards a destination; a transformation of the value
Example A comment of at most 280 characters Escaping < and > when rendering it in HTML

The rule that prevents most mistakes: validate on the way in, escape on the way out. Why not escape when storing? Because escaping depends on the destination and by storing escaped data you lose the original: if you store &lt;script&gt; and send it as JSON to a mobile app, that app literally shows &lt;script&gt;; if the same comment goes to an HTML page, to a CSV and to an e-mail, each destination needs a different escaping; and if you change templating systems, your data is already contaminated with the previous one's escaping.

Store the data as it is (validated, normalized) and escape at the rendering point. In Escena Viva, the public/app.js that renders comments must use textContent instead of innerHTML: that is the correct escaping for that destination.

The honest note about injection

You will see it repeated that "validation prevents SQL injection". That is false as a primary defense: a denylist of words (DROP, --, ;) can be dodged in a thousand ways and rejects legitimate comments from buyers called O'Brien. The real protection is parameterized queries, which separate code from data so the engine never interprets a value as an instruction; that arrives in Module 7, with Mongoose and Sequelize.

Meanwhile, in Escena Viva persistence is a JSON file and the equivalent risk is different but real:

  • Path traversal if an identifier ends up being part of a file path. That is why resolveWithin is still mandatory (Module 3).
  • Prototype pollution if you do Object.assign({}, req.body) with a __proto__ key. zod protects you because z.object() discards what is undeclared.
  • Denial of service with enormous bodies or catastrophic regexes. That is why limit is there in express.json() and patterns are bounded in the schemas.

Validating reduces surface, but it does not replace the specific defense of each destination.

Common Mistakes and Tips

  • Trusting the form's validation or reassigning req.query/req.params (a TypeError in Express 5): the HTML is courtesy, the server is the real border, and the result goes to req.validatedData.
  • Validating inside the domain. Session.sell() must not check types: by the time it gets there they are already correct. Its job is capacity.
  • Returning only the first error (it forces the client into one round trip per field; safeParse gives them all) or failing to filter the internal error: never send stack traces or received values, log them on the server and return just what is needed.
  • Indiscriminate coercion (z.coerce.number() turns '' into 0 and true into 1) or forgetting .trim() (a ' ses-001-1 ' fails against the regex with a baffling message).
  • Duplicating schemas across routes. Extract the common fragments (sessionId, buyerEmail) and compose them.
  • Tip: the schema is executable documentation. When someone asks what POST /api/orders accepts, the answer is src/schemas/order.js, and it cannot be out of date because it is the code that runs.

Exercises

Exercise 1: an order schema with several sessions

Escena Viva wants to allow orders with tickets for several sessions at once. Write createMultiOrderSchema with email, channel and tickets, an array of 1 to 4 items with sessionId and quantity. Add two rules that a per-field schema cannot express: the total number of tickets cannot exceed 6, and the same sessionId cannot be repeated. Hint: .superRefine() on the whole object.

Exercise 2: validating the events query end to end

Apply validate(eventsQuerySchema, 'query') to GET /api/events and adapt the controller to read from req.validatedData.query. Check with curl that ?perPage=999 and ?page=abc return 400 with the field and the reason, that ?soldOut=true arrives as a boolean and that with no query the defaults are page: 1, perPage: 10.

Exercise 3: the extra field

Send POST /api/orders a body with an undeclared field, for example {"sessionId":"ses-002-1","quantity":2,"email":"[email protected]","priceCents":1}. Check that the order is created and that priceCents does not reach the service. Then modify the schema with .strict() and check that it now returns 400 pointing at the extra key. Reason about which cases you would prefer each behavior for.

Solutions

Solution 1

// src/schemas/order.js (extension)
const orderLineSchema = z.object({ sessionId, quantity: ticketQuantity });
const createMultiOrderSchema = z
  .object({
    email: buyerEmail,
    channel: saleChannel.default('web'),
    tickets: z.array(orderLineSchema).min(1, 'At least one line').max(4, '4 sessions maximum'),
  })
  .superRefine((order, context) => {
    // Rule 1: the order's total number of tickets.
    const total = order.tickets.reduce((sum, line) => sum + line.quantity, 0);
    if (total > MAX_TICKETS_PER_ORDER) {
      context.addIssue({
        code: 'custom', path: ['tickets'],
        message: `The order adds up to ${total} tickets and the maximum is ${MAX_TICKETS_PER_ORDER}`,
      });
    }
    // Rule 2: no repeated sessions.
    const seen = new Set();
    order.tickets.forEach((line, i) => {
      if (seen.has(line.sessionId)) {
        context.addIssue({
          code: 'custom', path: ['tickets', i, 'sessionId'],
          message: `Session ${line.sessionId} is repeated; group the quantities`,
        });
      }
      seen.add(line.sessionId);
    });
  });

superRefine is the place for rules that cross fields. Careful: they are still shape rules, not state rules. "Is there capacity for these 6 tickets?" still belongs to the domain, because it depends on sales at that moment.

Solution 2

// src/controllers/events.js
async function listEvents(req, res) {
  const { venue, soldOut, page, perPage } = req.validatedData.query;
  let events = await getCatalog();
  if (venue) events = events.filter((e) => e.venue === venue);
  if (soldOut !== undefined) events = events.filter((e) => e.soldOut === soldOut);
  const total = events.length;
  const from = (page - 1) * perPage;
  res.json({
    page, perPage, total,
    totalPages: Math.max(1, Math.ceil(total / perPage)),
    events: events.slice(from, from + perPage).map((e) => e.toJSON()),
  });
}
# ?perPage=999 and ?page=abc return 400 INVALID_DATA with the field and the reason:
curl -s 'localhost:3000/api/events?perPage=999' | head -c 200
# {"error":{"code":"INVALID_DATA",...,"details":[{"field":"perPage",...,"type":"too_big"}]}}
curl -s 'localhost:3000/api/events' | head -c 60   # no query, default values
# {"page":1,"perPage":10,"total":3,"totalPages":1,...

Compare the controller with the version from the 06-03 exercise: the parseInt calls, the ?? and the manual caps are gone. That logic has not been lost, it has moved into the schema.

Solution 3

# With a plain z.object(): the extra field is silently discarded.
curl -s -X POST localhost:3000/api/orders -H 'Content-Type: application/json' \
  -d '{"sessionId":"ses-002-1","quantity":2,"email":"[email protected]","priceCents":1}'
# 201: the order is created with the session's real price, never with the one sent.
# With .strict() on the schema, it is rejected with 400 INVALID_DATA and the detail
# {"field":"priceCents","type":"unrecognized_keys"}.

When to prefer each one: pruning (the default behavior) fits a public API with varied clients, because an old client that sends obsolete fields keeps working; .strict() fits an internal API or one with sensitive fields, because an unexpected field is usually a client bug and it is better to flag it early. The crucial part in both cases: priceCents must never come out of the client's body; the price is set by the server by reading the session. If a body field can alter the money, you have a much bigger problem than a misconfigured schema.

Conclusion

Validation is your application's border, and Escena Viva now has it properly placed. You know why the browser form does not count, what gets validated (body, params, query, headers) and —conceptually the most important part— where each thing lives: the edge guarantees the shape, the domain guarantees the rules that depend on state; the first returns 400, the second 409. You have swapped Module 4's fifty imperative lines for declarative zod schemas that are validation, conversion, normalization, pruning of extra fields and executable documentation all at once. The validate(schema, source) middleware goes on any route, leaves the result on req.validatedData respecting that req.query is read-only in Express 5, and frees the controllers from every check. And you have seen the less obvious uses —validating the startup configuration and external services' responses— plus the distinction between validating on the way in and escaping on the way out, with the honest warning that the real defense against injection is the parameterized queries of Module 7.

One detail remains unresolved, and it is the one that closes the module: the validate middleware does not respond, it calls next(new ValidationError(...)), and that ValidationError does not exist yet and nobody is picking it up. The same goes for the EVENT_NOT_FOUND that loadEvent propagates, the INSUFFICIENT_CAPACITY the domain throws, the malformed JSON that express.json() rejects and the 404 for nonexistent routes.

In the next lesson, Error Handling, we build the destination for all of them: the four-argument error middleware, a hierarchy of our own in src/errors.js, the definitive home of the STATUS_BY_CODE table you wrote in Module 4, the distinction between operational errors and programming errors, and the policy for what to do when something fails beyond Express's reach.

Node.js Course: From Beginner to Advanced

Module 1: Introduction to Node.js

Module 2: Core Concepts

Module 3: File System and I/O

Module 4: HTTP and Web Servers

Module 5: NPM and Package Management

Module 6: The Express.js Framework

Module 7: Databases and ORMs

Module 8: Authentication and Authorization

Module 9: Testing and Debugging

Module 10: Advanced Topics

Module 11: Deployment and DevOps

Module 12: Real-World Projects

© Copyright 2026. All rights reserved