For two lessons now we've been putting off the big question: how does a network with hidden layers learn its weights? The perceptron rule is no use (nobody tells a hidden neuron what its "correct" output was), and in the previous lesson we prepared the missing ingredient: activations with a useful derivative. In this lesson we close the circle. First we'll formalize the forward pass as a chain of matrix multiplications, revisiting our 3-4-2-1 network with 29 parameters. Then we'll build up, in an accessible way, the notion of the gradient and the chain rule, and with them we'll understand backpropagation: the 1986 algorithm that distributes the blame for the error across all the layers. We'll trace it by hand with a small numerical example, see why the vanishing and exploding gradient problems are born from this very mechanism, and finish by training a 2-layer network from scratch in numpy on fictional TecnoMarket data.

Contents

  1. Forward pass with matrices: the 3-4-2-1 network in action
  2. The gradient: the error's compass
  3. The chain rule, explained painlessly
  4. Backpropagation: a numerical example traced by hand
  5. Vanishing and exploding gradients
  6. Complete numpy implementation: a 2-layer network for TecnoMarket

Forward pass with matrices: the 3-4-2-1 network in action

The forward pass is the journey of the data from input to output. We already did it on a small scale with the XOR MLP; now we write it out in general. For a layer with weight matrix $W$ (one column per neuron), biases $b$ and activation $f$:

$$z = xW + b \qquad a = f(z)$$

where $z$ is the pre-activation (each neuron's weighted sum) and $a$ the activation (what goes out to the next layer). An entire network is simply this operation repeated layer by layer. Let's bring back the 3-4-2-1 network we counted by hand in module 1 (29 parameters: 16 + 10 + 3) and give it life:

import numpy as np

rng = np.random.default_rng(0)

# Small random weights (still untrained), with the shapes of the 3-4-2-1 network
W1, b1 = rng.normal(0, 0.5, (3, 4)), np.zeros(4)   # input(3) -> hidden1(4)
W2, b2 = rng.normal(0, 0.5, (4, 2)), np.zeros(2)   # hidden1(4) -> hidden2(2)
W3, b3 = rng.normal(0, 0.5, (2, 1)), np.zeros(1)   # hidden2(2) -> output(1)

def relu(z):    return np.maximum(0, z)
def sigmoid(z): return 1 / (1 + np.exp(-z))

# A TecnoMarket order: [amount_norm, account_age_norm, address_mismatch]
x = np.array([[0.9, 0.1, 1.0]])          # shape (1, 3): 1 example, 3 features

a1 = relu(x @ W1 + b1)                   # shape (1, 4)
a2 = relu(a1 @ W2 + b2)                  # shape (1, 2)
y_hat = sigmoid(a2 @ W3 + b3)            # shape (1, 1): fraud probability
print(y_hat)                             # e.g. [[0.55]] -> it knows nothing yet

Three key observations:

  • The shapes fit together like dominoes: (1,3)·(3,4) → (1,4); (1,4)·(4,2) → (1,2); (1,2)·(2,1) → (1,1). Checking .shape at each step is the best bug detector.
  • Processing a batch is free: if x has shape (32, 3) — 32 orders at once — the same code returns (32, 1) without changing a single line. This is the practical reason for organizing data into batches (vocabulary from 01-04) and for GPUs shining here.
  • The output with random weights hovers around 0.5: the network is still a coin flip. Training consists of moving those 29 parameters until the outputs match the labels.
graph LR
    X[x - 1x3] -->|"W1 (3x4)"| A1[a1 - 1x4]
    A1 -->|"W2 (4x2)"| A2[a2 - 1x2]
    A2 -->|"W3 (2x1)"| Y[y_hat - 1x1]

The gradient: the error's compass

To train we need to measure the error with a number. For now we'll use the squared error $L = (\hat{y} - y)^2$ (loss functions in detail are the subject of the next lesson; this one is enough for us).

The central question of learning is: if I nudge one particular weight $w$ a little, does the error go up or down, and by how much? The answer is the partial derivative $\frac{\partial L}{\partial w}$. The vector of all those derivatives — one per parameter — is the gradient, and it points in the direction in which the error grows fastest. Therefore, to reduce the error we walk in the opposite direction:

$$w \leftarrow w - \eta \frac{\partial L}{\partial w}$$

This is gradient descent, and it is the exact mathematical version of the shower faucet analogy: the derivative says which way to turn (is the water coming out too hot or too cold?) and $\eta$ (the learning rate) how much to turn each time. Compare it with the perceptron rule from 02-01: that one was a rigid special case; this one works for any differentiable network.

The practical problem: our little 3-4-2-1 network has 29 derivatives to compute, and a real network has millions. Computing each one separately would be unfeasible. Backpropagation is, simply, the efficient way to compute them all in one pass, reusing computations. And its engine is the chain rule.

The chain rule, explained painlessly

Think of a chain of causes and effects inside the network:

the weight $w_1$ affects the pre-activation $z_1$, which affects the activation $h$, which affects the output $\hat{y}$, which affects the error $L$.

The chain rule says that the total effect of $w_1$ on $L$ is the product of the effects of each link:

$$\frac{\partial L}{\partial w_1} = \frac{\partial L}{\partial \hat{y}} \cdot \frac{\partial \hat{y}}{\partial z_2} \cdot \frac{\partial z_2}{\partial h} \cdot \frac{\partial h}{\partial z_1} \cdot \frac{\partial z_1}{\partial w_1}$$

A TecnoMarket analogy: if the purchasing department negotiates a price badly (cause), the product's cost goes up, the margin goes down, the quarterly profit falls. How much blame does purchasing bear for the quarterly result? The blame propagates by multiplying link by link along the chain. Backpropagation does exactly that: it starts from the final error and distributes blame backwards, layer by layer — hence the name, backward propagation. And it's efficient because the first factors in the chain (the ones nearest the output) are computed once and reused for all the earlier weights.

You only need three elementary derivatives, all of them already familiar:

Link Derivative Reading
$L = (\hat{y}-y)^2$ with respect to $\hat{y}$ $2(\hat{y}-y)$ The further from the target, the more blame comes in
Sigmoid $\sigma(z)$ with respect to $z$ $\sigma(z)(1-\sigma(z))$ You computed it in exercise 1 of the previous lesson
$z = wh + b$ with respect to $w$, $b$, $h$ $h$, $1$, $w$ A weight's blame is proportional to the signal it received

Backpropagation: a numerical example traced by hand

Let's do one complete iteration, with numbers, on the smallest possible network with a hidden layer: 1 input → 1 hidden neuron (sigmoid) → 1 output (sigmoid). Data: input $x = 1.0$, label $y = 1$ (it's fraud). Initial weights: $w_1 = 0.6$, $b_1 = 0$, $w_2 = 0.4$, $b_2 = 0$. Learning rate $\eta = 0.5$.

Forward pass (we store all the intermediate values: we'll need them on the way back):

Step Calculation Value
$z_1 = w_1 x + b_1$ $0.6 \times 1.0$ $0.600$
$h = \sigma(z_1)$ $\sigma(0.6)$ $0.646$
$z_2 = w_2 h + b_2$ $0.4 \times 0.646$ $0.258$
$\hat{y} = \sigma(z_2)$ $\sigma(0.258)$ $0.564$
$L = (\hat{y}-y)^2$ $(0.564-1)^2$ $0.190$

The network says "56% probability of fraud" when it should say almost 100%. Let's distribute the blame.

Backward pass, from the output toward the input:

  1. Blame at the output: $\frac{\partial L}{\partial \hat{y}} = 2(\hat{y}-y) = 2(0.564-1) = -0.872$ (negative: $\hat{y}$ must go up).
  2. We cross the output sigmoid: $\frac{\partial \hat{y}}{\partial z_2} = \hat{y}(1-\hat{y}) = 0.564 \times 0.436 = 0.246$. Accumulated: $\delta_2 = -0.872 \times 0.246 = -0.214$.
  3. Blame for the layer-2 parameters: $\frac{\partial L}{\partial w_2} = \delta_2 \cdot h = -0.214 \times 0.646 = -0.139$; $\frac{\partial L}{\partial b_2} = \delta_2 = -0.214$.
  4. The blame descending to the hidden layer travels through the weight: $\frac{\partial L}{\partial h} = \delta_2 \cdot w_2 = -0.214 \times 0.4 = -0.086$.
  5. We cross the hidden sigmoid: $\frac{\partial h}{\partial z_1} = h(1-h) = 0.646 \times 0.354 = 0.229$. Accumulated: $\delta_1 = -0.086 \times 0.229 = -0.020$.
  6. Blame for layer 1: $\frac{\partial L}{\partial w_1} = \delta_1 \cdot x = -0.020$; $\frac{\partial L}{\partial b_1} = \delta_1 = -0.020$.

Update ($w \leftarrow w - \eta \cdot \text{gradient}$):

$$w_2 = 0.4 - 0.5(-0.139) = 0.470 \qquad b_2 = 0 - 0.5(-0.214) = 0.107$$ $$w_1 = 0.6 - 0.5(-0.020) = 0.610 \qquad b_1 = 0 - 0.5(-0.020) = 0.010$$

Check: repeating the forward pass with the new weights, $\hat{y}$ rises from $0.564$ to $\approx 0.594$ and the loss drops from $0.190$ to $\approx 0.165$. The network has learned a little. Repeating this thousands of times (over many examples) is training.

Note a revealing detail in step 3 versus step 6: the gradient of $w_2$ ($-0.139$) is about 7 times larger than that of $w_1$ ($-0.020$). The layer furthest from the output receives less correction signal. That detail is the doorway to the next section.

Vanishing and exploding gradients

Every time the blame steps back one layer, it gets multiplied by two things: the activation's derivative and the weights. In our hand trace, crossing each sigmoid multiplied the signal by ~0.23–0.25. Remember from the previous lesson that the sigmoid's derivative never exceeds 0.25. So, in a network of 10 sigmoid layers:

$$0.25^{10} \approx 0.00000095$$

The blame reaching the earliest layers is practically zero: they don't learn. This is the vanishing gradient, and it's the historical reason why training deep networks was nearly impossible during the 1990s, and one of the reasons for ReLU's triumph (its derivative is 1 in the positive zone: the signal doesn't shrink as it passes through).

The opposite phenomenon also exists: if the weights are large (greater than 1 in magnitude), the signal gets amplified at each layer and the gradients blow up to enormous values — the exploding gradient — producing giant updates that wreck the training (you'll see losses jumping to nan). It's especially typical of recurrent networks, where the "depth" is the length of the sequence (we'll pick this up again in module 4), and the modern mitigation techniques are covered in module 5.

Problem Cause Symptom Mitigation (preview)
Vanishing Derivatives < 1 multiplied many times The earliest layers don't change; the loss stalls ReLU, good initialization, module 5 architectures
Exploding Large weights multiplied many times Loss oscillates wildly or becomes nan Lower $\eta$, gradient clipping (module 5)

Complete numpy implementation: a 2-layer network for TecnoMarket

Let's put it all together: a 3-8-1 network (3 inputs, 8 hidden with sigmoid, 1 sigmoid output) trained with gradient descent to estimate an order's fraud probability. It's the quality leap over the perceptron from 02-01: continuous inputs, a nuanced output, and hidden layers genuinely trained.

First, synthetically generated fictional TecnoMarket data (400 orders):

import numpy as np

rng = np.random.default_rng(7)
n = 400

amount      = rng.uniform(0, 1, n)      # normalized amount (0 = 0 EUR, 1 = 2000 EUR)
account_age = rng.uniform(0, 1, n)      # account age (0 = new, 1 = 10 years)
mismatch    = rng.integers(0, 2, n)     # 1 if shipping/billing addresses differ

# Hidden rule (which the network must discover): high risk if the order is expensive
# AND the account is new, aggravated by the address mismatch
risk = 2.2*amount - 2.0*account_age + 1.2*mismatch - 0.6
y = (risk + rng.normal(0, 0.3, n) > 0).astype(float).reshape(-1, 1)

X = np.column_stack([amount, account_age, mismatch])   # shape (400, 3)

Now the network. The entire backward pass is the exact matrix version of the 6 steps we traced by hand:

def sigmoid(z): return 1 / (1 + np.exp(-z))

# Initialization: small random values (never zeros in the hidden layers)
W1 = rng.normal(0, 0.5, (3, 8)); b1 = np.zeros((1, 8))
W2 = rng.normal(0, 0.5, (8, 1)); b2 = np.zeros((1, 1))

eta, epochs = 0.5, 3000
m = X.shape[0]                              # number of examples

for epoch in range(epochs):
    # ---- FORWARD: store the intermediates for the backward pass ----
    z1 = X @ W1 + b1                        # (400, 8)
    h  = sigmoid(z1)                        # (400, 8)
    z2 = h @ W2 + b2                        # (400, 1)
    y_hat = sigmoid(z2)                     # (400, 1)
    loss = np.mean((y_hat - y) ** 2)        # mean squared error

    # ---- BACKWARD: the same 6 steps as the hand trace ----
    d_yhat = 2 * (y_hat - y) / m            # step 1 (the /m averages the batch)
    delta2 = d_yhat * y_hat * (1 - y_hat)   # step 2: cross the output sigmoid
    dW2 = h.T @ delta2                      # step 3: blame for W2 (8, 1)
    db2 = delta2.sum(axis=0, keepdims=True)
    d_h = delta2 @ W2.T                     # step 4: blame descends through the weights
    delta1 = d_h * h * (1 - h)              # step 5: cross the hidden sigmoid
    dW1 = X.T @ delta1                      # step 6: blame for W1 (3, 8)
    db1 = delta1.sum(axis=0, keepdims=True)

    # ---- UPDATE: gradient descent ----
    W1 -= eta * dW1; b1 -= eta * db1
    W2 -= eta * dW2; b2 -= eta * db2

    if epoch % 500 == 0:
        accuracy = np.mean((y_hat > 0.5) == y)
        print(f"Epoch {epoch:4d} | loss {loss:.4f} | accuracy {accuracy:.2%}")

Typical output:

Epoch    0 | loss 0.2531 | accuracy 52.25%
Epoch  500 | loss 0.1025 | accuracy 86.50%
Epoch 1000 | loss 0.0668 | accuracy 91.75%
Epoch 2000 | loss 0.0525 | accuracy 93.50%
Epoch 2500 | loss 0.0499 | accuracy 94.00%

From a coin flip (52%) to a decent classifier (94%), without anyone programming the risk rule: the network extracted it from the data via forward + backward + update. Let's test it the way the fraud team would:

def fraud_prob(order):
    h = sigmoid(order @ W1 + b1)
    return sigmoid(h @ W2 + b2)[0, 0]

print(fraud_prob(np.array([[0.95, 0.05, 1]])))  # expensive, new account, different addresses -> ~0.98
print(fraud_prob(np.array([[0.30, 0.90, 0]])))  # cheap, long-time customer -> ~0.02

The code ↔ hand-trace correspondences deserve a second read: h.T @ delta2 is "a weight's blame is the signal it received times its neuron's blame", summed over the batch's 400 examples; delta2 @ W2.T is "the blame descends to the previous layer traveling through the weights". If you understand those two lines, you understand backpropagation.

Common Mistakes and Tips

  • Not storing the forward pass's intermediate values. The backward pass needs h, y_hat, etc. If you recompute or lose them, the code either gets complicated or gets slow. Always structure it as: forward stores → backward consumes.
  • Initializing all weights to zero. It worked for the perceptron; with hidden layers it's fatal: all the neurons in a layer receive the same gradient, learn the same thing, and the layer behaves like a single neuron (the symmetry problem). Initialize with small random values.
  • Getting the update sign wrong. It's $w \mathrel{-}= \eta \cdot dW$ (subtract: descend). An accidental += makes the loss go up and is a surprisingly frequent slip.
  • Dimension errors in the backward pass. Golden rule: the gradient dW must have exactly the same shape as W. Check dW1.shape == W1.shape as a quick test.
  • Loss turning into nan. Almost always an exploding gradient or a learning rate that's too high. Lower $\eta$ by an order of magnitude and review the initialization.
  • Verify the gradient numerically when in doubt. Approximate $\frac{\partial L}{\partial w} \approx \frac{L(w+\epsilon) - L(w-\epsilon)}{2\epsilon}$ with $\epsilon = 10^{-5}$ for one specific weight and compare it against your backward pass. If they don't agree to several digits, there's a bug in the backward pass.

Exercises

  1. Second iteration by hand. Continue the hand-traced example: with the updated weights ($w_1=0.610$, $b_1=0.010$, $w_2=0.470$, $b_2=0.107$), do the full forward pass and compute the new loss. Verify that it has dropped below 0.190.
  2. Feeling the vanishing gradient. In the TecnoMarket numpy network, change the architecture to 4 sigmoid hidden layers of 8 neurons (instead of 1) and train with the same hyperparameters. At epoch 0, print the mean magnitude of the gradients of the first and the last hidden layer (np.abs(dW).mean()). How many times smaller is the first layer's? What happens to the learning speed?
  3. Extreme learning rate. Go back to the original 3-8-1 network and train with eta = 20. Describe what you observe in the loss and explain it using this lesson's vocabulary. And with eta = 0.001?

Solutions

Exercise 1. Forward with the new weights: $z_1 = 0.610 \times 1.0 + 0.010 = 0.620$; $h = \sigma(0.620) = 0.650$; $z_2 = 0.470 \times 0.650 + 0.107 = 0.413$; $\hat{y} = \sigma(0.413) = 0.602$; $L = (0.602 - 1)^2 = 0.158$. The loss drops from 0.190 to ≈0.158 (the exact decimals may vary slightly depending on the rounding you carry along): gradient descent has worked.

Exercise 2. When stacking 4 sigmoid layers, at epoch 0 the mean gradient of the first layer is usually between 20 and 100 times smaller than that of the last one (each sigmoid crossed multiplies the signal by ≤ 0.25, plus the effect of the small weights). Training becomes much slower and the loss can stay stuck near its initial value for hundreds of epochs: you are watching the vanishing gradient live. Replacing the hidden sigmoids with ReLU (and its derivative, (z > 0).astype(float)) improves the situation noticeably.

Exercise 3. With eta = 20 the loss oscillates wildly or grows until it settles at a bad value (or nan appears outright if the gradients explode): each update is such a big turn of the faucet that we jump from "freezing" to "scalding" without ever passing through the sweet spot. With eta = 0.001 the loss decreases monotonically but extremely slowly: after 3000 epochs the accuracy has barely improved. The learning rate is a trade-off; in the next lesson we'll see techniques (and optimizers) that manage it better than a fixed value.

Conclusion

You now have the complete heart of deep learning: the forward pass turns inputs into predictions through chained matrix multiplications; the loss summarizes the error in one number; the gradient says which way to move each parameter; the chain rule lets you compute it layer by layer, distributing blame backwards — that is backpropagation, the algorithm that pulled networks out of their first winter in 1986; and gradient descent applies the correction, one gentle faucet turn after another. You've seen it in formulas, traced by hand number by number, and working in numpy on TecnoMarket orders. You've also seen that the backward pass's own multiplicative mechanics breeds its pathologies: gradients that vanish or explode.

In our training we used the simplest possible loss and optimizer: squared error and pure gradient descent over all the data at once. The next lesson examines those two choices under a magnifying glass: which loss functions exist and which one fits each problem, and which optimizers (mini-batch, momentum, RMSprop, Adam) make the descent faster and more stable. With that you'll have all the pieces to assemble your first complete network in Keras at the end of the module.

© Copyright 2026. All rights reserved