In the previous lesson the perceptron made its decisions with a step function: all or nothing, 0 or 1. That abrupt switch has two serious problems: it expresses no nuance ("how suspicious is this order?") and, above all, it cannot be differentiated, which — as we'll see in the next lesson — makes training hidden layers impossible. In this lesson you'll first understand why a network needs non-linear activation functions (without them, stacking a hundred layers is equivalent to having just one), and then we'll walk through the catalog of activations you'll use throughout the course: sigmoid, tanh, ReLU and its variants, and softmax. We'll close with a practical guide to which one to use in each situation, with TecnoMarket's projects as examples, and implement and visualize them all with numpy and matplotlib.
Contents
- Why non-linearity is indispensable
- Catalog of activation functions
- Comparison table
- Which activation to use in each layer depending on the problem
- Implementation in numpy and visualization with matplotlib
Why non-linearity is indispensable
Imagine we remove the activation (or use the "identity activation" $f(z) = z$, which is linear) in the 2-2-1 network that solved XOR. Each layer would perform only a linear operation: multiply by a matrix and add a bias. Let's see what happens when we chain two of them:
$$h = W_1^T x + b_1 \qquad \hat{y} = W_2^T h + b_2$$
Substituting the first into the second:
$$\hat{y} = W_2^T (W_1^T x + b_1) + b_2 = \underbrace{(W_2^T W_1^T)}{W'} x + \underbrace{(W_2^T b_1 + b_2)}{b'} = W' x + b'$$
The result is another linear operation, with a combined matrix $W'$ and bias $b'$. It doesn't matter how many layers you stack: the composition of linear functions is linear. A 100-layer network without non-linear activations is mathematically identical to a perceptron without the step — that is, to plain linear regression — and therefore crashes into XOR, and into any real-world problem, all over again.
You can verify it empirically in numpy:
import numpy as np
rng = np.random.default_rng(42)
W1, b1 = rng.normal(size=(3, 4)), rng.normal(size=4)
W2, b2 = rng.normal(size=(4, 2)), rng.normal(size=2)
x = rng.normal(size=3)
# 2-layer network WITHOUT activation
y_net = (x @ W1 + b1) @ W2 + b2
# A single equivalent, "collapsed" layer
W_eq = W1 @ W2
b_eq = b1 @ W2 + b2
y_eq = x @ W_eq + b_eq
print(np.allclose(y_net, y_eq)) # -> True: they were the same functionThe activation function is the piece that breaks this trap: by applying a non-linear function $f$ after each layer ($h = f(W^T x + b)$), the composition stops collapsing and every new layer adds real capacity. It's what makes possible the hierarchy of representations we saw in the first lesson of the course: edges → shapes → objects.
Catalog of activation functions
For each function we give the formula, the shape of its graph, the output range, and its typical use. They all receive the weighted sum $z = w \cdot x + b$ from the previous lesson.
Step
$$f(z) = \begin{cases} 1 & z \geq 0 \ 0 & z < 0 \end{cases}$$
- Graph: a horizontal line at 0 that jumps abruptly to 1 when crossing the origin.
- Range: {0, 1} (only two values).
- Use today: none in practice; historical interest (the perceptron). Its derivative is 0 everywhere (and infinite at the jump), so it conveys no information about "how much to correct".
Sigmoid (logistic)
$$\sigma(z) = \frac{1}{1 + e^{-z}}$$
- Graph: a smooth "S": close to 0 for very negative $z$, close to 1 for very positive $z$, and exactly 0.5 at $z=0$. It's the "smoothed" version of the step.
- Range: (0, 1).
- Interpretation: its output can be read as a probability. For TecnoMarket's fraud filter, instead of a blunt "fraud yes/no", the sigmoid returns "fraud probability 0.87", which enables graduated policies (block if > 0.9, review manually if 0.5–0.9).
- Problem: at the extremes the curve is almost flat (it saturates): the slope is practically zero, and that hampers learning in deep networks. It's the origin of the vanishing gradient we'll mention at the end and develop in the next lesson.
- Current use: the output layer in binary classification. Almost never in hidden layers anymore.
Hyperbolic tangent (tanh)
$$\tanh(z) = \frac{e^z - e^{-z}}{e^z + e^{-z}}$$
- Graph: the same "S" as the sigmoid, but stretched vertically: it passes through the origin (tanh(0) = 0).
- Range: (−1, 1).
- Advantage over the sigmoid: its output is zero-centered, which usually makes training the next layer easier (the inputs aren't all positive).
- Problem: it saturates at the extremes just like the sigmoid.
- Current use: hidden layers in some recurrent architectures (you'll meet it again inside LSTMs in module 4).
ReLU (Rectified Linear Unit)
$$\text{ReLU}(z) = \max(0, z)$$
- Graph: flat at 0 for negative $z$, and a straight line of slope 1 for positive $z$. An "elbow" at the origin.
- Range: [0, +∞).
- Why it dominates: it's extremely cheap to compute (one comparison), it doesn't saturate for positive values (the slope stays at 1 no matter how large $z$ gets), and in practice it makes deep networks train much faster. It has been the default activation for hidden layers since the AlexNet era (2012, history lesson).
- Although it looks "almost linear", it isn't: the elbow at 0 is enough to break the linear collapse from section 1.
- Problem: dead neurons (dying ReLU): if a neuron ends up with $z$ always negative for all the data, its output is always 0, its slope is always 0, and it never learns again.
ReLU variants: Leaky ReLU and ELU
Leaky ReLU — fixes dead neurons by letting a small slope through in the negative zone:
$$f(z) = \begin{cases} z & z \geq 0 \ \alpha z & z < 0 \end{cases} \qquad (\alpha \approx 0.01)$$
Graph: like ReLU, but the left-hand part is not flat — it's a nearly horizontal line with slope $\alpha$. Range: (−∞, +∞).
ELU (Exponential Linear Unit) — smooths the negative zone with an exponential:
$$f(z) = \begin{cases} z & z \geq 0 \ \alpha(e^z - 1) & z < 0 \end{cases} \qquad (\alpha \approx 1)$$
Graph: identical to ReLU on the right; on the left it descends smoothly and flattens out at $-\alpha$. Range: (−α, +∞). Mean output closer to zero than ReLU and no sharp elbow, at the cost of computing an exponential.
Softmax (for multiclass output)
Unlike the previous ones, softmax doesn't act neuron by neuron but on the entire vector of outputs $z = (z_1, \dots, z_K)$ of the last layer:
$$\text{softmax}(z)i = \frac{e^{z_i}}{\sum{j=1}^{K} e^{z_j}}$$
- Range: each component in (0, 1) and all summing to exactly 1: a probability distribution over the $K$ classes.
- Intuition: exponentiating amplifies the differences (the largest $z_i$ takes almost all the probability) and dividing by the sum normalizes.
- Use: the output layer in multiclass classification. When TecnoMarket's photo classifier looks at an image and has to choose among {laptop, coffee maker, vacuum cleaner, monitor}, softmax will return something like (0.85, 0.02, 0.03, 0.10): 85% confidence in "laptop".
Comparison table
| Function | Formula | Range | Saturates? | Cost | Main use today |
|---|---|---|---|---|---|
| Step | $z \geq 0 \Rightarrow 1$ | {0, 1} | — (not differentiable) | Minimal | Historical (perceptron) |
| Sigmoid | $1/(1+e^{-z})$ | (0, 1) | Yes, both ends | Medium (exp) | Binary output |
| Tanh | $\frac{e^z - e^{-z}}{e^z + e^{-z}}$ | (−1, 1) | Yes, both ends | Medium (exp) | Hidden in RNN/LSTM |
| ReLU | $\max(0, z)$ | [0, +∞) | No (positive side) | Minimal | Hidden, by default |
| Leaky ReLU | $\max(\alpha z, z)$ | (−∞, +∞) | No | Minimal | Hidden if dead neurons appear |
| ELU | $z$ or $\alpha(e^z-1)$ | (−α, +∞) | Negative side only | Medium (exp) | Hidden, alternative to ReLU |
| Softmax | $e^{z_i}/\sum e^{z_j}$ | (0,1), sums to 1 | — | Medium | Multiclass output |
Which activation to use in each layer depending on the problem
The choice follows two different rules depending on the layer:
Hidden layers: always start with ReLU. If you detect many dead neurons, try Leaky ReLU or ELU. Save tanh for recurrent architectures (module 4). Avoid the sigmoid in hidden layers.
Output layer: it's decided by the type of problem, because the output must have the shape of the answer you're looking for:
| Problem type | TecnoMarket example | Output neurons | Output activation |
|---|---|---|---|
| Regression (real number) | Predicting units of a coffee maker sold next week | 1 | None (linear/identity) |
| Binary classification | Is this order fraudulent? / Is this review positive? | 1 | Sigmoid |
| Multiclass classification (a single correct class) | Is this photo a laptop, a coffee maker, a vacuum cleaner or a monitor? | K (one per class) | Softmax |
| Multilabel (several classes at once) | Tagging a review with several topics: "shipping", "price", "quality" | K | Sigmoid on each neuron |
Notice the difference between the last two rows: softmax forces the probabilities to compete (they sum to 1: a photo is one single thing); independent sigmoids allow several labels to be true at once. In the optimization and loss function lesson you'll see that each of these outputs is paired with a specific loss function.
Implementation in numpy and visualization with matplotlib
They all fit in a few lines. Save this code in your tecnomarket-dl/ folder (for example as activations.py), because we'll reuse it in the next lesson:
import numpy as np
import matplotlib.pyplot as plt
def step(z): return (z >= 0).astype(float)
def sigmoid(z): return 1 / (1 + np.exp(-z))
def tanh(z): return np.tanh(z) # numpy already provides it
def relu(z): return np.maximum(0, z)
def leaky_relu(z, alpha=0.01):
return np.where(z >= 0, z, alpha * z)
def elu(z, alpha=1.0):
return np.where(z >= 0, z, alpha * (np.exp(z) - 1))
def softmax(z):
z = z - np.max(z) # numerical stability trick
e = np.exp(z)
return e / e.sum()Two important details:
np.where(condition, a, b)chooses element by element: where the condition holds it takesa, where it doesn't,b. It's the vectorized form of "if z ≥ 0 then... else...".- In
softmaxwe subtract the maximum before exponentiating. It doesn't change the mathematical result (numerator and denominator both end up divided by the same constant $e^{z_{max}}$), but it preventsnp.exp(1000)from overflowing to infinity. Without this trick, softmax fails on large inputs.
Now we plot all the curves to anchor the visual intuition:
z = np.linspace(-5, 5, 200) # 200 points between -5 and 5
functions = [
("Step", step), ("Sigmoid", sigmoid), ("Tanh", tanh),
("ReLU", relu), ("Leaky ReLU", leaky_relu), ("ELU", elu),
]
fig, axes = plt.subplots(2, 3, figsize=(12, 6))
for ax, (name, f) in zip(axes.flat, functions):
ax.plot(z, f(z), linewidth=2)
ax.set_title(name)
ax.axhline(0, color="gray", linewidth=0.5) # horizontal reference axis
ax.axvline(0, color="gray", linewidth=0.5) # vertical reference axis
ax.grid(True, alpha=0.3)
plt.tight_layout()
plt.show()When you run it you'll see, in a single figure, the six shapes described in the catalog: the jump of the step, the two "S" curves (one between 0 and 1, the other between −1 and 1), the ReLU elbow and its two variants with the negative zone "rescued". And a quick softmax test with the photo classifier's scores:
scores = np.array([3.1, -0.5, 0.2, 1.4]) # laptop, coffee maker, vacuum cleaner, monitor
probs = softmax(scores)
print(np.round(probs, 3)) # -> [0.792 0.022 0.044 0.143]
print(probs.sum()) # -> 1.0The class with the highest score (laptop) concentrates the probability, but the others don't drop to zero: the model expresses its uncertainty.
One final note that connects to the next lesson: we've said several times that sigmoid and tanh "saturate" and that this hampers learning. The exact reason is that in the flat zones their slope (derivative) is almost zero, and the training algorithm — backpropagation — multiplies slopes in a chain: many near-zero numbers multiplied together give a practical zero, and the earliest layers stop learning. This is the famous vanishing gradient; we'll see it emerge naturally in the next lesson, and the techniques to fight it in depth in module 5.
Common Mistakes and Tips
- Putting an activation on the output of a regression. If you're predicting units sold and you put a sigmoid on the output, your network can only predict between 0 and 1. For regression, the last layer goes without an activation (or
activation="linear"in Keras, which is the same thing). - Using softmax with a single output neuron. Softmax over a one-element vector always returns 1.0, whatever the input. For binary: 1 neuron + sigmoid, or 2 neurons + softmax — but don't mix them.
- Computing softmax without subtracting the maximum. It works with small numbers and blows up with large ones (
np.exp(1000) = inf, andinf/inf = nan). Always use the stable version. - Sigmoid in many hidden layers. That's the classic recipe for vanishing gradients. In hidden layers, ReLU by default.
- Ignoring dead neurons. If after training you notice that many ReLU units output 0 for every example, lower the learning rate or switch to Leaky ReLU/ELU.
- Choosing the output activation without thinking about the loss. They come in pairs (sigmoid ↔ binary cross-entropy, softmax ↔ categorical cross-entropy). We cover the exact pairing in the optimization lesson.
Exercises
- Visual derivatives. The derivative of the sigmoid is $\sigma'(z) = \sigma(z)(1 - \sigma(z))$. Implement it and plot it alongside the sigmoid on the same graph for $z \in [-6, 6]$. What is its maximum value and where is it reached? What is it worth, approximately, at $z = 5$? Relate what you see to the vanishing gradient.
- The right output classifier. For each TecnoMarket project, state how many output neurons you would use and with which activation: (a) deciding whether a review is positive or negative; (b) classifying a photo among 12 product categories; (c) predicting a customer's average purchase amount next month; (d) tagging a review with whichever labels apply from {shipping, price, quality, service}.
- XOR with sigmoid. Rewrite the 2-2-1 MLP that solved XOR in the previous lesson, replacing the step with the sigmoid (keep the same weights:
W1=[[1,1],[1,1]],b1=[-0.5,-1.5],W2=[[1],[-2]],b2=[-0.5]). First multiply all the weights and biases by 10. Compute the output for the four inputs and check that, after rounding, it's still XOR. Why is multiplying by 10 necessary?
Solutions
Exercise 1.
def sigmoid_derivative(z):
s = sigmoid(z)
return s * (1 - s)
z = np.linspace(-6, 6, 200)
plt.plot(z, sigmoid(z), label="sigmoid")
plt.plot(z, sigmoid_derivative(z), label="derivative")
plt.legend(); plt.grid(True); plt.show()The maximum of the derivative is 0.25 and it's reached at $z = 0$. At $z = 5$ it's worth approximately 0.0066: almost zero. Since the derivative never exceeds 0.25, when several sigmoid layers are chained the slopes multiply (0.25 × 0.25 × ... → 0 quickly): that is the seed of the vanishing gradient.
Exercise 2. (a) 1 neuron with a sigmoid (binary). (b) 12 neurons with softmax (mutually exclusive multiclass). (c) 1 neuron with no activation (regression; the amount isn't bounded to (0,1)). (d) 4 neurons, each with its own sigmoid (multilabel: a review can talk about shipping and price at the same time, so the probabilities must not compete).
Exercise 3. With the weights ×10 (b1=[-5,-15], etc.), the weighted sums land far from zero and the sigmoid behaves almost like the step: for (0,0) the output is ≈0.007 → 0; for (0,1) and (1,0), ≈0.98 → 1; for (1,1), ≈0.01 → 0. Rounding: XOR. Multiplying by 10 is necessary because with the original weights the sums fall in the central, gentle zone of the sigmoid (where the output hovers around 0.3–0.6 instead of approaching 0 or 1) and rounding no longer reproduces the step's behavior. The moral: the sigmoid is a "softened" step, and the magnitude of the weights controls how closely it resembles one.
Conclusion
You now know why activation functions are not decoration but the very condition for depth to be worth anything: without non-linearity, any stack of layers collapses into a single linear transformation. You have the complete catalog — the step (history), sigmoid and tanh (the "S" curves that saturate), ReLU and its variants (the standard for hidden layers), softmax (the multiclass output) — and the two practical rules: ReLU by default in hidden layers, and at the output whatever the problem dictates (nothing for regression, sigmoid for binary, softmax for multiclass).
Moreover, all modern activations share a property the step lacked: they have a useful derivative, a slope that says in which direction and by how much to correct. That slope is exactly what's needed by the algorithm we've been postponing for two lessons: in the next lesson we'll finally open the backpropagation box and see how a complete network — our old friend 3-4-2-1 included — learns its weights layer by layer.
Deep Learning Course
Module 1: Introduction to Deep Learning
- What is Deep Learning?
- History and evolution of Deep Learning
- Applications of Deep Learning
- Basic concepts of neural networks
- Setting up the work environment
Module 2: Neural Network Fundamentals
- Perceptron and Multilayer Perceptron
- Activation functions
- Forward and backward propagation
- Optimization and loss functions
- Your first complete neural network
Module 3: Convolutional Neural Networks (CNN)
- Introduction to CNNs
- Convolutional and pooling layers
- Popular CNN architectures
- CNN applications in image recognition
Module 4: Recurrent Neural Networks (RNN)
- Introduction to RNNs
- LSTM and GRU
- RNN applications in natural language processing
- Sequences and time series
Module 5: Advanced Deep Learning Techniques
- Generative Adversarial Networks (GAN)
- Autoencoders
- Transfer Learning
- Regularization and improvement techniques
- Attention mechanisms and Transformers
Module 6: Tools and Frameworks
- Introduction to TensorFlow
- Introduction to PyTorch
- Framework comparison
- Development environments and additional resources
- Saving, loading and deploying models
Module 7: Hands-On Projects
- Image classification with CNNs
- Text generation with RNNs
- Anomaly detection with Autoencoders
- Building a GAN for image generation
- Fine-tuning a pretrained model
