In the previous lesson we trained a complete network with the most basic tools possible: we measured the error with the squared difference and corrected with pure gradient descent over all the data at once. It worked, but both choices were provisional. This lesson examines them in depth, because they are two of the most important design decisions in any project: the loss function defines what "being wrong" means (and it must match the type of problem and the output activation you chose in 02-02), and the optimizer defines how the gradient is used to correct (and determines whether your training takes minutes or hours, or simply never converges). We'll cover the standard losses with formulas and intuition, the three variants of gradient descent, the learning rate's pathologies, and the modern optimizers — momentum, RMSprop and Adam — you'll use in practice. We close by seeing how all of this is declared in Keras with a single line: compile().

Contents

  1. What a loss function actually is
  2. Losses for regression: MSE and MAE
  3. Losses for classification: binary and categorical cross-entropy
  4. Decision table: which loss for which problem
  5. Gradient descent: batch, stochastic and mini-batch
  6. The learning rate and its pathologies
  7. Modern optimizers: momentum, RMSprop and Adam
  8. All together in Keras: compile()

What a loss function actually is

The loss function takes the prediction $\hat{y}$ and the true label $y$ and returns a single number: how badly the network did on that example (or, averaging, on a batch). All the learning described in the backpropagation lesson starts by differentiating that number; that's why the loss must be differentiable and must penalize sensibly the errors of the specific problem.

This is not a technical detail: the loss is the contract you hand the network. If the contract is badly chosen, the network will diligently optimize the wrong thing. The good news is that for the three standard problem types (regression, binary classification, multiclass classification) there are canonical answers.

Losses for regression: MSE and MAE

For predicting continuous numbers — TecnoMarket's weekly demand for coffee makers — the two basic losses compare prediction and reality over the $m$ examples:

Mean Squared Error (MSE):

$$\text{MSE} = \frac{1}{m}\sum_{i=1}^{m} (\hat{y}_i - y_i)^2$$

Mean Absolute Error (MAE):

$$\text{MAE} = \frac{1}{m}\sum_{i=1}^{m} |\hat{y}_i - y_i|$$

The difference lies in how they treat large errors. By squaring, MSE punishes big misses disproportionately: being off by 10 units weighs 100 times more than being off by 1 (not 10 times more). MAE treats every unit of error equally.

The practical consequence with TecnoMarket: suppose the real demand in an atypical week (Black Friday) was 500 units and your normal history hovers around 50. With MSE, that single extreme point dominates the total loss and the network will warp its normal predictions to fail less badly there; with MAE, the point carries just its fair weight. Rule of thumb:

  • MSE by default: smooth gradients, stable training, and punishing large errors hard is usually what you want.
  • MAE when your data has outliers you don't want hijacking the training.
import numpy as np

y_true = np.array([48, 52, 45, 500])   # real demand; the last week was Black Friday
y_pred = np.array([50, 50, 50, 60])    # prediction from a "normal" model

mse = np.mean((y_pred - y_true) ** 2)  # -> 48405.25  (the outlier dominates everything)
mae = np.mean(np.abs(y_pred - y_true)) # -> 112.25    (more representative of everyday performance)

Losses for classification: binary and categorical cross-entropy

For classification we could use MSE ("the predicted probability minus the label, squared", as we did provisionally in the previous lesson), but there's a loss far better suited to outputs that are probabilities: cross-entropy.

Binary cross-entropy (BCE) — for a sigmoid output with 0/1 labels:

$$\text{BCE} = -\frac{1}{m}\sum_{i=1}^{m} \Big[ y_i \log \hat{y}_i + (1-y_i)\log(1-\hat{y}_i) \Big]$$

The formula is less scary than it looks, because for each example only one of the two terms survives:

  • If $y = 1$ (it was fraud), the loss is $-\log \hat{y}$: it's nearly 0 if the network said $\hat{y} \approx 1$ (a confident hit) and it tends to infinity if it said $\hat{y} \approx 0$ (a fully confident mistake).
  • If $y = 0$, symmetrically, the loss is $-\log(1-\hat{y})$.

That is the key intuition: cross-entropy brutally punishes misplaced confidence. Saying "2% probability of fraud" on an order that was fraud costs $-\log(0.02) \approx 3.9$; saying "40%" costs only $0.92$. With MSE the difference would be much milder (0.96 versus 0.36), and on top of that MSE + sigmoid produces tiny gradients precisely when the network is most wrong (because of the saturation we saw in 02-02); the sigmoid + BCE pairing cancels that problem and learns fast exactly on the serious mistakes.

Categorical cross-entropy (CCE) — for a softmax output with $K$ classes:

$$\text{CCE} = -\frac{1}{m}\sum_{i=1}^{m} \log \hat{y}_{i,c_i}$$

where $\hat{y}_{i,c_i}$ is the probability the network assigned to the correct class of example $i$. Direct reading: the loss only looks at how much probability you gave the right answer. If TecnoMarket's photo classifier sees a coffee maker and its softmax assigns (laptop 0.05, coffee maker 0.90, vacuum cleaner 0.03, monitor 0.02), the loss is $-\log(0.90) = 0.105$: good. If it assigned coffee maker 0.10, the loss is $-\log(0.10) = 2.30$: a serious penalty.

In Keras you'll see two variants of CCE depending on the label format: categorical_crossentropy if they're one-hot ([0,1,0,0]) and sparse_categorical_crossentropy if they're integers (1). It's the same mathematics with different input encoding; the sparse variant will save us a conversion step in the next lesson.

Decision table: which loss for which problem

This table completes the output-activation table from 02-02 — each row is the activation + loss pairing that goes together:

Problem TecnoMarket example Output activation Loss Keras name
Regression Predicting a product's weekly demand None (linear) MSE (or MAE with outliers) "mse" / "mae"
Binary classification Fraudulent order? Positive review? Sigmoid (1 neuron) Binary cross-entropy "binary_crossentropy"
Multiclass (one-hot labels) Photo → 1 of 12 product categories Softmax (K neurons) Categorical cross-entropy "categorical_crossentropy"
Multiclass (integer labels) Same, with labels 0..11 Softmax (K neurons) Sparse CCE "sparse_categorical_crossentropy"
Multilabel Review → topics {shipping, price, quality} Sigmoid (K neurons) BCE on each output "binary_crossentropy"

Gradient descent: batch, stochastic and mini-batch

In the previous lesson we computed the gradient using all 400 orders at once for each update. That is one of three possible strategies, which differ in how many examples they look at before each correction:

Variant Examples per update Advantages Drawbacks
Batch (full batch) All ($m$) Exact gradient, smooth trajectory Very slow and prohibitive in memory on large datasets; only 1 update per epoch
Stochastic (pure SGD) 1 Cheap, frequent updates; the noise helps escape bad corners Very erratic trajectory; wastes the GPU's vectorization
Mini-batch A batch of 32–256 Balance: reasonable gradient, exploits the GPU, many updates per epoch You have to choose the batch size

The TecnoMarket intuition: to adjust the anti-fraud policy you could wait to analyze all the year's orders (batch: a very well-informed decision, once a year), react after each individual order (stochastic: extremely fast but hysterical — one odd order sends you swerving), or review each morning the previous day's batch of orders (mini-batch: frequent and stable). The absolute standard in deep learning is mini-batch, and it's what Keras's batch_size parameter controls — the same "batch" from the 01-04 vocabulary. Terminology note: in practice, when someone says "SGD" they almost always mean mini-batch SGD.

With mini-batches, an epoch is no longer one update but many: with 60,000 examples and batch_size=32, each epoch is 1,875 weight updates.

The learning rate and its pathologies

The learning rate $\eta$ multiplies the gradient in each update: $w \leftarrow w - \eta \nabla L$. You've known it since the perceptron rule and the shower faucet; now, its failure modes. Picture the loss as a valley and yourself descending blindly, taking steps of length $\eta$:

  • $\eta$ too large: you leap from one slope to the opposite one. The loss oscillates or even diverges (grows until nan, as you saw in exercise 3 of the previous lesson).
  • $\eta$ somewhat large: you descend fast at first but then bounce around the bottom without fine-tuning: the loss stalls at a mediocre, noisy value.
  • $\eta$ too small: every step is millimetric. The loss decreases in a pretty, monotonic way… and after hours it's still far from the bottom. You can also get stuck on the first ledge you find.
  • $\eta$ just right: a fast descent at first that gradually settles. In practice you find it by trying powers of 10 (0.1, 0.01, 0.001…) and watching the loss curve.

A classic trick is to reduce $\eta$ during training (large steps at the start, fine ones at the end). The modern optimizers in the next section largely automate this management.

Modern optimizers: momentum, RMSprop and Adam

Pure gradient descent (SGD) uses only the current gradient at each step. Modern optimizers add memory to navigate the loss valley better.

Momentum: the rolling ball

$$v \leftarrow \beta v - \eta \nabla L \qquad w \leftarrow w + v$$

Instead of taking independent steps, we maintain a velocity $v$ that accumulates the recent gradients ($\beta \approx 0.9$: each step keeps 90% of the previous momentum). Intuition: a heavy ball rolling downhill. Two benefits: in directions where the gradient points consistently the same way, the ball accelerates; in directions where the gradient flips sign at every step (the slope-to-slope oscillations), the opposing impulses cancel out and the trajectory smooths. On top of that, the momentum helps push through small ledges and bumps.

RMSprop: a learning rate tailored to each parameter

$$s \leftarrow \beta s + (1-\beta)(\nabla L)^2 \qquad w \leftarrow w - \frac{\eta}{\sqrt{s} + \epsilon} \nabla L$$

RMSprop keeps, for each parameter separately, a moving average $s$ of the square of its recent gradients, and divides the step by its square root. Effect: parameters with historically enormous gradients take moderate steps (they're braked) and those with minuscule gradients take relatively larger ones (they're encouraged). In other words: there's no longer a single $\eta$ for 29 (or 29 million) parameters with different scales — each one receives a tailored step.

Adam: momentum + RMSprop

Adam (Adaptive Moment Estimation) combines both ideas: it maintains the moving average of the gradient (like momentum) and the moving average of its square (like RMSprop), with a small startup correction, and uses them together. Its default hyperparameters (learning_rate=0.001, $\beta_1=0.9$, $\beta_2=0.999$) work well on an enormous variety of problems, which is why Adam is the default optimizer to start any project with.

Optimizer Added idea Memory it keeps When to use it
SGD None (baseline) Nothing Teaching; production with expert tuning
SGD + momentum Inertia / velocity 1 value per parameter A classic in vision, excellent final results when well tuned
RMSprop Per-parameter adapted step 1 value per parameter Historically used in recurrent networks (module 4)
Adam Momentum + adaptation 2 values per parameter Default: robust with almost no tuning

An honest nuance: Adam converges fast "without touching anything", but a finely tuned SGD+momentum sometimes reaches slightly better solutions in vision. For this course: always start with Adam; consider alternatives only when you have a reason.

All together in Keras: compile()

Everything above is declared in Keras in a single call, the compile() step that turns an architecture into a trainable model. We won't build the full model yet (that's the next lesson) — just read this line with fresh eyes:

from tensorflow import keras

# Fraud detection (binary): sigmoid output + BCE + Adam
model.compile(
    optimizer="adam",                  # the optimizer from section 7
    loss="binary_crossentropy",        # the loss from the decision table
    metrics=["accuracy"],              # metrics ONLY to observe, they are not optimized
)

# Variant with explicit hyperparameters:
model.compile(
    optimizer=keras.optimizers.Adam(learning_rate=0.001),
    loss="binary_crossentropy",
    metrics=["accuracy"],
)

# Two other TecnoMarket projects, same structure:
# - Photo classifier (12 classes, integer labels):
#     optimizer="adam", loss="sparse_categorical_crossentropy", metrics=["accuracy"]
# - Demand forecasting (regression):
#     optimizer="adam", loss="mse", metrics=["mae"]

Points you must be clear on when reading it:

  • loss is what gets optimized; metrics is what gets observed. Accuracy isn't differentiable (it's a count of correct answers), so it can't be the loss; it's displayed so that we humans can follow the progress.
  • The string "adam" uses the default values; the keras.optimizers.Adam(...) object lets you adjust the learning rate when needed.
  • The batch_size (mini-batch) doesn't go in compile() but in fit(), together with the epochs — we'll see it in action in the next lesson.
  • The output-activation ↔ loss coherence is your responsibility: Keras won't warn you if you pair softmax with binary_crossentropy; it will simply train badly.

Common Mistakes and Tips

  • Mismatching output and loss. The three valid pairings are: linear+MSE/MAE, sigmoid+BCE, softmax+CCE. Any other combination (softmax with BCE, sigmoid with CCE…) trains strangely or outright badly, often without raising an explicit error.
  • categorical_crossentropy with integer labels. If your labels are 0..K-1 (not one-hot) and you use regular CCE, Keras throws a shape error or learns garbage. Use sparse_categorical_crossentropy or convert with keras.utils.to_categorical.
  • Optimizing accuracy directly. You can't: it isn't differentiable. The loss is its differentiable stand-in; accept that sometimes the loss drops without the accuracy moving (the network gains confidence on answers it already had right).
  • A batch that's too large "because it's faster". Each epoch has fewer updates and the gradient loses its beneficial noise; it can also exhaust the GPU's memory. Start at 32 or 64.
  • Switching optimizers before touching the learning rate. If training goes badly, the first thing to adjust is $\eta$ (try ÷10 and ×10); swapping SGD for Adam without reviewing $\eta$ is treating the symptom.
  • Comparing losses across different problems. An MSE of 0.05 and a BCE of 0.05 are neither comparable nor "good" in the abstract; a loss only makes sense compared against itself over the course of training (and against a dumb baseline, like always predicting the mean).

Exercises

  1. Cross-entropy by hand. The fraud detector evaluates 4 orders with true labels y = [1, 0, 1, 0] and predicts probabilities y_hat = [0.9, 0.2, 0.3, 0.95]. Compute the total BCE (mean of the 4) with numpy and by hand, term by term. Which order contributes the most loss and why?
  2. Reasoned choice. For each TecnoMarket case, state the output activation, the loss (Keras name) and a reasonable metric: (a) estimating an order's delivery time in days; (b) classifying support tickets into {billing, shipping, warranty, other}; (c) deciding whether to show an offer to a customer (yes/no); (d) predicting the buy-back price of used phones knowing there are occasional atypical appraisals.
  3. Optimizer race. Reuse the 3-8-1 numpy network from the previous lesson and add momentum to the update (keep a velocity vW1, vb1, vW2, vb2 initialized to zeros and apply $v = 0.9v - \eta \nabla$; $w = w + v$). Train 3000 epochs with pure SGD and with momentum using the same $\eta = 0.5$ and compare the loss at epochs 100, 500 and 1000.

Solutions

Exercise 1.

y     = np.array([1, 0, 1, 0])
y_hat = np.array([0.9, 0.2, 0.3, 0.95])
bce = -np.mean(y*np.log(y_hat) + (1-y)*np.log(1-y_hat))   # -> 1.135

Term by term: $-\log(0.9)=0.105$; $-\log(1-0.2)=0.223$; $-\log(0.3)=1.204$; $-\log(1-0.95)=2.996$. Mean ≈ 1.13. The biggest contributor is the fourth order (loss ≈ 3.0): the network said "95% fraud" and it was legitimate — high misplaced confidence, exactly what cross-entropy punishes hardest (even more than the third order, which is also wrong but with less conviction).

Exercise 2. (a) Regression: linear output, loss="mse", metric mae (interpretable in days). (b) 4-category multiclass: 4-neuron softmax, sparse_categorical_crossentropy (integer labels), metric accuracy. (c) Binary: 1 sigmoid, binary_crossentropy, accuracy. (d) Regression with declared outliers: linear output, loss="mae" (robust against atypical appraisals), metric mae.

Exercise 3. Minimal changes in the training loop:

vW1 = np.zeros_like(W1); vb1 = np.zeros_like(b1)
vW2 = np.zeros_like(W2); vb2 = np.zeros_like(b2)
# ... after computing dW1, db1, dW2, db2 in the backward pass:
vW1 = 0.9*vW1 - eta*dW1;  W1 += vW1
vb1 = 0.9*vb1 - eta*db1;  b1 += vb1
vW2 = 0.9*vW2 - eta*dW2;  W2 += vW2
vb2 = 0.9*vb2 - eta*db2;  b2 += vb2

Typical result: with momentum, the loss at epoch 100 is already where pure SGD gets to around epoch 400–600; by epoch 1000 both approach the same final value, but momentum got there much sooner. (If with momentum the loss oscillates at the start, that's the inertia accumulating too much impulse: lower $\eta$ to 0.1 and compare again — a good example of how optimizer and learning rate are tuned together.)

Conclusion

You now have both training decisions under control. The loss function translates "being wrong" into a differentiable number, and it's chosen by problem type, forming a pair with the output activation: linear+MSE/MAE for regression, sigmoid+BCE for binary, softmax+CCE for multiclass. The optimizer decides how to use the gradient: mini-batch as the universal data strategy, the learning rate as the most delicate hyperparameter, and Adam (momentum + per-parameter adaptation) as a robust starting point. And you've seen that in Keras all of this is declared in one line of compile().

With this, the module has completed all its pieces: neuron and layers (02-01), activations (02-02), forward and backward (02-03), losses and optimizers (02-04). In the next lesson we assemble them all: you'll build, train and evaluate in Keras your first complete neural network on a real image dataset, following the course methodology — public prototype first, TecnoMarket after.

© Copyright 2026. All rights reserved