In 05-01 we saw that a single neuron draws a straight line and that stacking neurons in layers breaks that limit. Now comes the practical question Marta asks herself in front of the laptop: how many layers, how many neurons, which activation function, which output? Those decisions are the architecture of the network, and they are the most important hyperparameters in deep learning. In this lesson we take a multilayer perceptron (MLP) apart piece by piece: dense layers, width and depth, how to count parameters by hand; the modern activation functions (ReLU and its family, softmax) and why the sigmoid fell out of use in hidden layers (the vanishing gradient, at an intuitive level); how to choose the output layer and the loss function according to the problem; and how to represent the inputs, with embeddings as the key idea for text and products. We will close with a map of the architecture families (convolutional, recurrent, transformers, autoencoders) that 05-04 and 05-05 develop, and with code: we will build in PyTorch the MLP of the NovaMarket returns predictor on the 21 columns prepared in module 4, count its parameters and make a prediction with it without training. It matters because the architecture fixes what the network can learn: a bad choice is not fixed with more epochs.

Contents

  1. Anatomy of a multilayer perceptron: dense layers, width and depth
  2. Counting parameters by hand: 21 → 16 → 8 → 1
  3. Activation functions: sigmoid, tanh, ReLU and variants, softmax
  4. The vanishing gradient (intuition) and why ReLU
  5. Output layer and loss function according to the problem
  6. How to represent the inputs: numeric, categorical and embeddings
  7. Overview of architecture families
  8. Code: the MLP of the returns predictor in PyTorch
  9. Common Mistakes and Tips
  10. Exercises
  11. Conclusion

  1. Anatomy of a multilayer perceptron: dense layers, width and depth

A multilayer perceptron (MLP) is the network of 05-01 generalised: an input layer, one or more dense hidden layers (each neuron connected to all the outputs of the previous layer) and an output layer. Each dense layer computes, for all its neurons at once, h = f(W · x + b), where W is a matrix with one row per neuron and one column per input, b a vector of biases and f the activation applied element by element. It is the same neuron as in 05-01, written in matrix form so that the computer can evaluate it in one go.

Two dimensions describe the shape of an MLP:

  • Width: how many neurons each layer has. More neurons = more "steps" to draw the function (the universal approximator of 05-01), and more parameters.
  • Depth: how many hidden layers there are. Each layer rewrites the output of the previous one, so a deep network composes simple transformations into a complex one: the first layer detects elementary features, the next one combinations of features, and so on. For the same number of parameters, depth is usually more efficient than width for complex functions; and it is what gives deep learning its name (05-04).
Decision Effect of increasing it Effect of decreasing it Practical starting rule
Width (neurons per layer) More capacity, more parameters, more risk of overfitting May underfit Powers of 2 between 8 and 512; decreasing towards the output ("funnel")
Depth (hidden layers) More abstract representations; harder to train (section 4) Less capacity for composition 1-3 layers for tabular data; tens or hundreds in vision and language
Total parameters Of the order of the training examples or fewer, unless regularisation is strong (05-03)

For the NovaMarket returns predictor (2,249 training orders, 21 columns) we will use a modest funnel: 21 → 16 → 8 → 1.

  1. Counting parameters by hand: 21 → 16 → 8 → 1

Each dense layer with n_in inputs and n_out neurons has n_in × n_out weights plus n_out biases. Let us count:

Layer Inputs Neurons Weights Biases Parameters
Hidden 1 21 16 21 × 16 = 336 16 352
Hidden 2 16 8 16 × 8 = 128 8 136
Output 8 1 8 × 1 = 8 1 9
Total 497

Against the 22 parameters of logistic regression (21 weights + 1 bias) that is 22 times more, and yet it is a tiny network: a typical vision model has tens of millions and a large language model billions. Knowing how to count parameters is useful for three things: estimating memory (each parameter is a 32-bit number, 4 bytes: 497 parameters take up 2 KB; 7 billion, 28 GB), estimating the risk of overfitting (497 parameters for 2,249 examples is reasonable; 50,000 would not be without heavy regularisation) and catching mistakes when building the network (if PyTorch reports a different number from the one you computed, something is not wired the way you think).

  1. Activation functions: sigmoid, tanh, ReLU and variants, softmax

The activation of the hidden layers is what provides the non-linearity (without it, 05-01, the network collapses into a linear model). The main options:

Function Formula Range Values at z = −3, −1, 0, 1, 3 Use today
Sigmoid 1 / (1 + e^(−z)) (0, 1) 0.047; 0.269; 0.5; 0.731; 0.953 Only at the output of binary classification
Tanh (e^z − e^(−z)) / (e^z + e^(−z)) (−1, 1) −0.995; −0.762; 0; 0.762; 0.995 Recurrent networks (05-04); centred at 0, better than sigmoid in hidden layers
ReLU max(0, z) [0, ∞) 0; 0; 0; 1; 3 Default in hidden layers
Leaky ReLU z if z > 0; 0.01·z otherwise (−∞, ∞) −0.03; −0.01; 0; 1; 3 When many ReLUs "die" (section 4)
ELU / GELU / SiLU Smooth variants of ReLU ≈ (−1, ∞) GELU is the usual one in transformers (05-05)
Softmax e^(zᵢ) / Σⱼ e^(zⱼ) (over a vector) Each output in (0, 1), summing to 1 See example Output of multiclass classification

Observations:

  • The sigmoid and the tanh saturate: for |z| > 3 the output hardly changes any more (0.953 → 0.999...). That makes them poor for hidden layers, as we will see in section 4.
  • The ReLU (rectified linear unit) is almost offensively simple: it lets positives through and zeroes out negatives. It is cheap to compute, does not saturate on the right and its slope is exactly 1 for z > 0. Since 2012 it has been the default activation for hidden layers.
  • The softmax is not a per-neuron activation but one over a vector: it turns k arbitrary numbers (called logits) into k probabilities that sum to 1, exaggerating the differences. Example with three classes and z = (2, 1, −1): the exponentials are e² = 7.389, e¹ = 2.718, e⁻¹ = 0.368, which sum to 10.475; dividing, softmax = (0.705, 0.259, 0.035). The class with the largest logit takes the largest probability; if you add a constant to all the logits the result does not change. With two classes, softmax is equivalent to the sigmoid.

  1. The vanishing gradient (intuition) and why ReLU

In 05-03 we will see that the network learns by computing the slope of the loss with respect to each weight, and that this slope is obtained by multiplying, layer by layer from the output backwards, the slopes of every function the signal passes through. Here the numerical intuition is enough:

  • The slope of the sigmoid is σ(z)·(1 − σ(z)), which is at most 0.25 (at z = 0) and falls quickly: at z = 2 it is 0.105 and at z = 5, 0.0066. In a 10-layer network with sigmoids, in the best case the slope reaching the first layer has been multiplied by 0.25¹⁰ ≈ 0.000001; in practice, much less, because hardly any neuron sits at z = 0. The first layers receive an almost null correction signal: they do not learn. This is the vanishing gradient problem, and it is one of the reasons why in the 1990s networks stayed at two or three layers.
  • The slope of the ReLU is 1 for z > 0 and 0 for z < 0. Multiplying ones reduces nothing: the signal reaches the deep layers intact along the active paths. That, together with more data and GPUs, is what made it possible to train networks of tens of layers from 2012 onwards (05-04).

The ReLU has its own risk: a neuron whose z is negative for every example always returns 0 and its slope is 0 too, so it never gets updated again (a "dead ReLU"). It usually happens with learning rates that are too high; Leaky ReLU and ELU avoid it by letting a small slope through on the negative side. Practical rule: ReLU in the hidden layers; if you see many neurons stuck at zero, try Leaky ReLU or lower the learning rate.

  1. Output layer and loss function according to the problem

The output layer and the loss function (04-01: the number that training minimises) always come as a pair, dictated by the type of problem:

Problem Output neurons Output activation Loss NovaMarket example
Binary classification 1 Sigmoid Binary cross-entropy (BCE, the log loss of 04-01) Returned / not returned; positive / negative review; photo of damaged / intact package
Multiclass classification (one correct class) k (one per class) Softmax (Categorical) cross-entropy Return reason (4 classes); category of an incident
Multi-label classification (several at once) k Sigmoid on each BCE per output Tags of a review: "price", "shipping", "quality"
Regression 1 (or several) None (linear) Mean squared error (MSE) or MAE Units of the NovaClean next week; days until delivery

The binary cross-entropy for an example with label y and predicted probability p is −[y·log(p) + (1 − y)·log(1 − p)]. If the order was returned (y = 1) and the network said p = 0.9, the loss is −log(0.9) = 0.105; if it said p = 0.5, 0.693; if it said p = 0.1, 2.303. It heavily punishes misplaced confidence, exactly as in logistic regression. An implementation detail you will see in 05-03: PyTorch offers nn.BCEWithLogitsLoss, which merges the sigmoid and the BCE into one numerically more stable operation; with it the output layer is left without a sigmoid and it is applied afterwards only to read probabilities. In this lesson, so that the output is readable, we keep the sigmoid explicit.

  1. How to represent the inputs: numeric, categorical and embeddings

A network only eats numbers in a reasonable range, so the preparation of 04-03 still applies and is even more critical:

  • Numeric: standardised (mean 0, standard deviation 1) or scaled to [0, 1]. Without scaling, the large columns dominate the weighted sums and training becomes unstable (we saw it with distance in 05-01). We will reuse the ColumnTransformer from module 4 as is.
  • Binary: 0/1 as they are.
  • Categorical with few categories: one-hot (one column per category), like category or payment_method in module 4.
  • Categorical with many categories (NovaMarket's 12,000 products, the words of a language, postcodes): one-hot is unworkable (12,000 columns almost all zero) and, besides, it treats all categories as equally different from each other. Deep learning's solution is the embedding: assign each category a dense vector of few dimensions (for instance 8 or 64) whose values are parameters the network learns together with the rest. After training, categories that behave alike end up with similar vectors: the NovaClean and another robot vacuum will sit close together; so will "excellent" and "great".
import torch, torch.nn as nn

table = nn.Embedding(num_embeddings=12000, embedding_dim=8)   # 12,000 products -> vectors of 8
print(table.weight.shape)                                       # torch.Size([12000, 8]): 96,000 parameters
ids = torch.tensor([3, 4587])                                   # two products by their index
print(table(ids).shape)                                         # torch.Size([2, 8]): one vector each

nn.Embedding is literally a lookup table: row i = vector of product i. At first it is random; training moves it around. The idea is central in recommendation (use case 1: customer and product as embeddings whose dot product predicts affinity, 05-04) and in language (each word or token is an embedding; it is the first layer of any transformer, 05-05).

  1. Overview of architecture families

The MLP is the "generic" architecture. For each type of data a family has emerged that builds in assumptions about the structure of the data and therefore learns with far fewer parameters and examples:

flowchart TD
    RN["Neural networks"] --> MLP["MLP / dense<br/>tabular data<br/>(this lesson and 05-03)"]
    RN --> CNN["Convolutional (CNN)<br/>images, signals<br/>incident photos (05-04)"]
    RN --> RNN["Recurrent (RNN, LSTM, GRU)<br/>sequences, time series<br/>weekly demand (05-04)"]
    RN --> AE["Autoencoders<br/>compression, anomalies<br/>fraud (05-04)"]
    RN --> TR["Transformers<br/>text, and today almost everything<br/>reviews, assistant (05-05)"]
Family Assumption it builds in Typical data Lesson
MLP (dense) None: each input is independent Table of columns 05-02, 05-03
Convolutional Patterns are local and repeat at any position Images, audio, series 05-04
Recurrent Order matters; there is a state carried along Time series, text (historically) 05-04
Autoencoder The data live in a lower-dimensional space Any (no labels) 05-04
Transformer Each element must "look at" all the others (attention) Text, images, audio, code 05-05

They all share the ingredients of this lesson: layers of neurons, ReLU or similar activations, an output layer with its loss and embeddings for categorical data. What changes is how the neurons are connected.

  1. Code: the MLP of the returns predictor in PyTorch

We reuse the preparation from module 4 (generate_orders_ml, dirty_orders, prepare_orders and build_preprocessing from novamarket_ml.py, which produce the 21 columns) and build the network 21 → 16 → 8 → 1.

import numpy as np
import torch
import torch.nn as nn
from sklearn.model_selection import train_test_split
from novamarket_ml import generate_orders_ml, dirty_orders, prepare_orders, build_preprocessing

# 1. Data: the same split as in 04-05 (750 test orders)
X, y = prepare_orders(dirty_orders(generate_orders_ml(3000, 42), 42))
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.25, random_state=42, stratify=y)
prep = build_preprocessing().fit(Xtr)               # fits imputation, scaling and one-hot ONLY on train
Xtr_m = prep.transform(Xtr).astype(np.float32)      # numpy matrix 2249 x 21, in float32 (what PyTorch uses)
Xte_m = prep.transform(Xte).astype(np.float32)
print(Xtr_m.shape, Xte_m.shape)

# 2. Architecture
torch.manual_seed(42)                                # initial weights are random: we fix the seed
net = nn.Sequential(
    nn.Linear(21, 16), nn.ReLU(),                    # hidden layer 1: 21 inputs -> 16 neurons, ReLU
    nn.Linear(16, 8),  nn.ReLU(),                    # hidden layer 2: 16 -> 8, ReLU
    nn.Linear(8, 1),   nn.Sigmoid())                 # output: 8 -> 1 probability of return
print(net)

# 3. Parameters
for name, p in net.named_parameters():
    print(f"{name:10s} {str(tuple(p.shape)):10s} {p.numel():4d}")
print("Total:", sum(p.numel() for p in net.parameters()))

Output:

(2249, 21) (750, 21)
Sequential(
  (0): Linear(in_features=21, out_features=16, bias=True)
  (1): ReLU()
  (2): Linear(in_features=16, out_features=8, bias=True)
  (3): ReLU()
  (4): Linear(in_features=8, out_features=1, bias=True)
  (5): Sigmoid()
)
0.weight   (16, 21)    336
0.bias     (16,)        16
2.weight   (8, 16)     128
2.bias     (8,)          8
4.weight   (1, 8)        8
4.bias     (1,)          1
Total: 497

Explanation: nn.Linear(21, 16) creates the 16 × 21 matrix W (PyTorch stores one row per neuron) and the vector b of 16; named_parameters() walks through all the trainable tensors, and the count matches the table of section 2 (497). The indices 0, 2, 4 are the positions inside the Sequential (the activations occupy 1, 3, 5 and have no parameters).

A small function that prints the architecture with its count (handy for checking any Sequential):

def summary(net):
    print(f"{'layer':<10}{'input':>8}{'output':>8}{'parameters':>12}")
    total = 0
    for i, layer in enumerate(net):
        if isinstance(layer, nn.Linear):
            n = layer.in_features * layer.out_features + layer.out_features
            total += n
            print(f"Linear {i:<3}{layer.in_features:>8}{layer.out_features:>8}{n:>12}")
        else:
            print(f"{layer.__class__.__name__:<10}{'':>8}{'':>8}{0:>12}")
    print(f"{'TOTAL':<10}{'':>8}{'':>8}{total:>12}")

summary(net)

Output:

layer        input  output  parameters
Linear 0        21      16         352
ReLU                                 0
Linear 2        16       8         136
ReLU                                 0
Linear 4         8       1           9
Sigmoid                              0
TOTAL                              497

Now the forward pass: we push five test orders through the network without having trained it:

from sklearn.metrics import roc_auc_score

X_t = torch.tensor(Xte_m[:5])                        # numpy -> PyTorch tensor (5 x 21)
with torch.no_grad():                                # we do not need slopes yet
    prob = net(X_t)                                  # forward: applies the 6 pieces in order
print("Probabilities:", prob.numpy().ravel().round(3), " actual:", yte.values[:5])

with torch.no_grad():
    all_probs = net(torch.tensor(Xte_m)).numpy().ravel()
print("Mean:", all_probs.mean().round(3), " min:", all_probs.min().round(3), " max:", all_probs.max().round(3))
print("AUC untrained:", round(roc_auc_score(yte, all_probs), 3))

Output (varies with the seed and the PyTorch version):

Probabilities: [0.546 0.547 0.54  0.558 0.549]  actual: [0 1 0 0 1]
Mean: 0.551  min: 0.531  max: 0.605
AUC untrained: 0.517

Reading: the network knows nothing. With small random weights, all the output z values hover around 0 and the sigmoid returns ≈ 0.55 for every order; the AUC of 0.517 is that of a coin toss (0.5). We can look inside to see the flow of one order:

with torch.no_grad():
    z1 = net[0](X_t[:1])          # weighted sum of layer 1 (16 values)
    h1 = net[1](z1)               # after ReLU: negatives are zeroed
    h2 = net[3](net[2](h1))       # layer 2 + ReLU (8 values)
    z3 = net[4](h2)               # output logit
print("z1:", z1.numpy().round(2))
print("h1:", h1.numpy().round(2))
print("h2:", h2.numpy().round(2))
print("z3:", z3.numpy().round(3), " sigmoid:", torch.sigmoid(z3).numpy().round(3))

Output:

z1: [[-0.2   0.03  0.51 -0.63 -0.01 -0.71 -0.   -0.78 -0.31 -0.13 -0.34 -0.46  0.34 -0.64 -0.21  0.21]]
h1: [[0.   0.03 0.51 0.   0.   0.   0.   0.   0.   0.   0.   0.   0.34 0.    0.    0.21]]
h2: [[0.01 0.12 0.   0.11 0.   0.   0.07 0.  ]]
z3: [[0.187]]  sigmoid: [[0.546]]

You can see the ReLU in action: of the 16 weighted sums of the first layer, 12 were negative and are set to 0; only 4 neurons "speak" for this order. In the second layer, 4 out of 8. That is normal (and desirable: sparse representations), as long as they are not always the same neurons for every order (dead ReLUs). The specific weights (net[0].weight[0, :5] gives something like [0.167, 0.181, -0.051, 0.2, -0.048]) mean nothing yet: they are the random starting point of the descent we will run in 05-03.

Common Mistakes and Tips

  • Sigmoid in the hidden layers. It is the natural reflex after module 4, but it causes the vanishing gradient. Reserve the sigmoid for the binary output; ReLU (or Leaky ReLU/GELU) in the hidden layers.
  • Softmax with a single neuron, or sigmoid with several mutually exclusive classes. The softmax of a single value is always 1; and for "return reason" (one class out of four) independent sigmoids can give 0.9 to two classes at once. Respect the table in section 5.
  • Mismatching output and loss. Sigmoid + MSE trains poorly (the loss flattens in the saturated regions); use BCE. And do not apply the sigmoid twice (once in the network and again inside BCEWithLogitsLoss).
  • Forgetting the scaling. With inputs on disparate scales, ReLU and gradient descent misbehave. Reuse the ColumnTransformer of 04-03 and fit it only on training data.
  • Oversizing by default. 3 layers of 512 neurons for 2,249 orders is almost 400,000 parameters: it will memorise the training set. Start small (like 21 → 16 → 8 → 1) and grow only if validation asks for it (04-06).
  • Trusting the output before training. As we have just seen, the network returns ≈ 0.55 for everything; a correct architecture is not a model. Until 05-03 there is no prediction.

Exercises

Exercise 1. Work out by hand the parameters of a network 21 → 32 → 16 → 8 → 1 (all dense, ReLU in the hidden layers, sigmoid at the output). Build it with nn.Sequential, check the figure with summary and compare it with the number of training orders (2,249). Does it seem prudent to you without regularisation?

Exercise 2. Replace the output nn.Linear(8, 1), nn.Sigmoid() with nn.Linear(8, 2) followed by torch.softmax(..., dim=1) (two neurons: "not returned" and "returned"). Count the parameters and check that the two probabilities sum to 1 for any order. Is it a different network from the binary one with a sigmoid? Which loss would it need according to the table in section 5?

Exercise 3. For each NovaMarket problem, state the number of output neurons, the output activation, the loss and how you would represent the main input: (a) predicting the 1-to-5-star rating of a review from its text; (b) predicting the weekly units of the NovaClean from the previous 8 weeks; (c) predicting the return reason (4 categories) from the 21 order columns plus the product identifier (12,000 values).

Solutions

Solution 1. 21×32 + 32 = 704; 32×16 + 16 = 528; 16×8 + 8 = 136; 8×1 + 1 = 9; total 1,377, and summary confirms it. That is 0.6 parameters per training example: manageable, but already in the range where early stopping and dropout are advisable (05-03); with the 497-parameter network we have headroom, which is why we chose it.

Solution 2. The network has 497 − 9 + (8×2 + 2) = 506 parameters. With torch.softmax(net2(x), dim=1) you get, for instance, [[0.535 0.465]] for an order: they sum to 1. Mathematically it is equivalent to the binary one with a sigmoid (the softmax of two logits z₀, z₁ is the sigmoid of z₁ − z₀), only with one redundant parameter per input of the last layer. It would need the categorical cross-entropy (nn.CrossEntropyLoss, which in PyTorch includes the softmax). In practice, for binary problems the single-neuron version is used.

Solution 3. (a) 5 neurons, softmax, categorical cross-entropy; the input, the text, is represented as a sequence of word embeddings (or, as a baseline, a bag of words: 05-05). A reasonable alternative: treat it as a regression from 1 to 5 with a linear output + MSE, if it matters "by how much" it is wrong (predicting 4 when it was 5 is less serious than predicting 1). (b) 1 neuron, linear output, MSE or MAE; input: the previous 8 units standardised (an MLP with 8 inputs or, better, a recurrent network that respects the order, 05-04). (c) 4 neurons, softmax, cross-entropy; the 21 columns from the ColumnTransformer concatenated with an nn.Embedding(12000, 8) of the product (8 learned dimensions instead of 12,000 one-hot columns).

Conclusion

We have taken apart the architecture of a multilayer perceptron: dense layers that compute f(W·x + b), width and depth as design decisions, and counting parameters by hand (21 → 16 → 8 → 1 = 497, against the 22 of logistic regression). We compared the activation functions and understood why the ReLU replaced the sigmoid in the hidden layers: the sigmoid saturates and its maximum slope of 0.25 makes the correction signal vanish in deep networks, whereas the ReLU passes it on intact. We paired output layer and loss according to the problem (sigmoid + BCE, softmax + cross-entropy, linear + MSE), reviewed how to represent the inputs and introduced embeddings as the way to give learned vectors to products and words. And we built in PyTorch the MLP of the NovaMarket returns predictor on the 21 columns of module 4, counted its parameters and inspected a forward pass: with random weights, the network says ≈ 0.55 to everything and its AUC is 0.517. The network knows nothing yet.

In the next lesson, How a Network Learns: Gradient Descent and Backpropagation, we will teach it: we will see the loss as a landscape whose slope we now do know, gradient descent step by step, the chain rule that distributes the error backwards layer by layer, and we will write the complete training loop in PyTorch (DataLoader, Adam, epochs, early stopping and dropout) to check whether this 497-parameter network improves, or not, on the 0.844 AUC of logistic regression.

Fundamentals of Artificial Intelligence (AI)

Module 1: Introduction to Artificial Intelligence

Module 2: Basic Principles of AI

Module 3: Algorithms in AI

Module 4: Machine Learning

Module 5: Neural Networks and Deep Learning

Module 6: Logic and Expert Systems

Module 7: Tools and Programming Languages in AI

Module 8: Projects and Case Studies

Module 9: Exercises and Practice

Module 10: Additional Resources

© Copyright 2026. All rights reserved