We closed module 4 with a question from Marta and Diego: the returns predictor is validated (AUC 0.84 with logistic regression), but can a model read the reviews in reviews.csv or look at the incident photos? Classical algorithms need someone to first turn the text or the image into meaningful numeric columns; neural networks, by contrast, learn that conversion themselves. This module explains how. In this first lesson we start with the basic building block: the artificial neuron, which will turn out to be something you already know (the logistic regression of 04-04, seen with fresh eyes); Rosenblatt's perceptron of 1958, which we will implement from scratch in numpy and train on logic gates and on a NovaMarket problem; the famous XOR problem, which a single neuron cannot solve and which helped trigger the first AI winter (01-01); and the solution, stacking neurons in layers, which is what we call a neural network. We will finish by seeing, without training it yet, how that same network is written in PyTorch. This matters because all of deep learning is this idea repeated at scale: understanding one neuron well, and why layers are needed, is understanding 80 % of what comes next.

Contents

  1. From the biological neuron to the artificial neuron
  2. Anatomy of an artificial neuron: inputs, weights, bias, sum and activation
  3. The logistic regression of 04-04 was a neuron
  4. Rosenblatt's perceptron and its learning rule
  5. The perceptron in numpy: logic gates and NovaMarket urgent orders
  6. The XOR problem: the limit of one neuron and the two-layer solution
  7. What a neural network is: input, hidden and output layers
  8. The universal approximator and the linear model vs network comparison
  9. A glimpse of the destination: the same network in PyTorch
  10. Common Mistakes and Tips
  11. Exercises
  12. Conclusion

  1. From the biological neuron to the artificial neuron

A biological neuron receives signals from other neurons through its dendrites, integrates them in the cell body and, if the accumulated signal exceeds a threshold, fires an impulse down the axon towards the next ones. The connections (synapses) are not all alike: some excite and others inhibit, and their strength changes with experience. The human brain has around 86 billion neurons connected in a massively parallel way.

In 1943, McCulloch and Pitts proposed a minimal mathematical model of that neuron: binary inputs, a sum and a threshold. It is the inspiration for everything that follows, but it is worth being honest about the analogy:

Biological neuron Artificial neuron Comment
Dendrites receive signals Input vector x Reasonable analogue
Strength of the synapse Weight w Reasonable analogue: it is what gets "learned"
Firing threshold Bias b and activation function Reasonable analogue
Electrical impulses over time, chemistry, plasticity One real number computed in one go No analogy: the artificial neuron is infinitely simpler
Local learning, no supervisor Gradient descent with labels (05-03) The brain does not do backpropagation

So "neural network" is a historical name, not a claim that we are imitating the brain. What really matters is the mathematics: a very flexible function built by stacking very simple pieces.

  1. Anatomy of an artificial neuron: inputs, weights, bias, sum and activation

An artificial neuron receives n numeric inputs x₁, ..., xₙ and produces an output in three steps:

  1. Weighted sum: multiply each input by its weight and add a bias: z = w₁·x₁ + w₂·x₂ + ... + wₙ·xₙ + b. In vector notation, z = w · x + b. The weights say how much, and in which direction, each input matters; the bias shifts the point at which the neuron "switches on" (without it, the decision boundary would be forced to pass through the origin).
  2. Activation function: transforms z into the output a = f(z). The two historical ones are:
    • Step (Heaviside): f(z) = 1 if z > 0, 0 otherwise. The neuron "fires" or it does not. It is the one in McCulloch-Pitts and in the perceptron.
    • Sigmoid: σ(z) = 1 / (1 + e^(−z)). A smooth version of the step: it gives a number between 0 and 1 interpretable as a probability, and since it has no jumps it will let us compute slopes (05-03).
  3. Output: a is used as the prediction or passed on to other neurons.
flowchart LR
    x1["x₁"] -->|w₁| S(("Σ + b"))
    x2["x₂"] -->|w₂| S
    x3["x₃"] -->|w₃| S
    S -->|z| F["f(z)"]
    F --> a["a = output"]

Numerical example with three inputs: x = (2, 0, 1), weights w = (0.5, −1, 0.8), bias b = −1. Then z = 0.5·2 + (−1)·0 + 0.8·1 − 1 = 0.8. With the step the output is 1 (because 0.8 > 0); with the sigmoid, σ(0.8) = 1 / (1 + e^(−0.8)) ≈ 0.69.

All a single neuron can do is draw a linear boundary: in two dimensions, a straight line (w₁x₁ + w₂x₂ + b = 0); in three, a plane; in 21, a hyperplane. On one side of the boundary it says "1", on the other "0". Hold on to this idea: it is the key to the XOR problem.

  1. The logistic regression of 04-04 was a neuron

Reread the definition of logistic regression in 04-04: "it computes the weighted sum z = a + b₁·x₁ + ... + bₙ·xₙ and passes it through the sigmoid". That is exactly a neuron with a sigmoid activation; only the notation changes (the coefficients are now called weights, the intercept is called the bias). Let us check it with the coefficients we obtained back then for the four-column returns predictor:

import numpy as np

def step(z):
    return np.where(z > 0, 1, 0)

def sigmoid(z):
    return 1 / (1 + np.exp(-z))

# Order from 04-04: €250, 1 item, 5 delivery days, new customer
x = np.array([250.0, 1, 5, 1])
# Coefficients that LogisticRegression learned in 04-04 (we now call them weights)
w = np.array([0.012, -0.2012, 0.3069, 1.8821])
b = -4.906                                    # the intercept is the bias

z = w @ x + b                                 # weighted sum (dot product + bias)
print("z =", round(z, 3), " sigmoid =", round(sigmoid(z), 3), " step =", step(z))

Output:

z = 1.309  sigmoid = 0.787  step = 1

The same 0.79 (up to the rounding of the coefficients) that we computed then with predict_proba. Important conclusion: you have already trained a neuron in module 4, and fit of LogisticRegression did what training a network will do: search for the weights that minimise a loss. The difference between "a neuron" and "a network" is not in the building block but in how many blocks there are and how they are connected.

  1. Rosenblatt's perceptron and its learning rule

In 1958 Frank Rosenblatt built the perceptron: a neuron with a step activation and, above all, the first learning rule that adjusted the weights from examples (01-01). The rule is astonishingly simple. For each example (x, y) in the training set:

  1. Compute the prediction ŷ = step(w · x + b).
  2. Compute the error e = y − ŷ (0 if correct, +1 if it should have said 1 and said 0, −1 the other way round).
  3. Update: w ← w + η · e · x and b ← b + η · e, where η (eta) is the learning rate, a small number such as 0.1.

Read it slowly: if the prediction is right, nothing is touched; if it should have given 1 and gave 0, the weights are pushed in the direction of x (so that next time z is larger); if it should have given 0 and gave 1, they are pushed the opposite way. The set is traversed several times (each complete pass is called an epoch) until no error remains. Rosenblatt proved the perceptron convergence theorem: if the data are linearly separable (there exists a line or hyperplane that separates the classes without error), the rule finds one in a finite number of steps. That "if" is the catch we will see in section 6.

  1. The perceptron in numpy: logic gates and NovaMarket urgent orders

5.1 Implementation

import numpy as np

def train_perceptron(X, y, rate=0.1, epochs=10):
    """Rosenblatt's rule. Returns weights, bias and the errors made in each epoch."""
    w = np.zeros(X.shape[1])          # one weight per column, we start at zero
    b = 0.0
    history = []
    for ep in range(epochs):
        errors = 0
        for xi, yi in zip(X, y):      # example by example
            pred = 1 if xi @ w + b > 0 else 0
            error = yi - pred          # 0, +1 or -1
            if error != 0:
                w = w + rate * error * xi
                b = b + rate * error
                errors += 1
        history.append(errors)
        if errors == 0:               # an epoch without errors: we have converged
            break
    return w, b, history

Line-by-line explanation: X is a matrix with one example per row; w starts at zeros; in each epoch we go through the examples, compute z = xi @ w + b (@ is numpy's dot product) and apply the step (> 0); if there is an error, we correct weights and bias with the rule; we record how many errors the epoch had and stop as soon as an epoch finishes clean.

5.2 AND and OR logic gates

Logic gates are the historical "hello world" of neurons: two binary inputs and one output.

X_gates = np.array([[0, 0], [0, 1], [1, 0], [1, 1]])
y_and = np.array([0, 0, 0, 1])
y_or  = np.array([0, 1, 1, 1])

for name, y in [("AND", y_and), ("OR", y_or)]:
    w, b, hist = train_perceptron(X_gates, y, rate=0.1, epochs=20)
    pred = step(X_gates @ w + b)
    print(f"{name}: w={w.round(2)} b={b:.2f} errors per epoch={hist} prediction={pred}")

Output:

AND: w=[0.2 0.1] b=-0.20 errors per epoch=[1, 3, 3, 2, 1, 0] prediction=[0 0 0 1]
OR: w=[0.1 0.1] b=0.00 errors per epoch=[1, 2, 1, 0] prediction=[0 1 1 1]

Let us follow the AND trace by hand to see the rule in action (only the steps where there was an error are shown; η = 0.1):

Epoch Input x y z = w·x + b Prediction Error w after update b after update
1 (1, 1) 1 0.0 0 +1 (0.1, 0.1) 0.1
2 (0, 0) 0 0.1 1 −1 (0.1, 0.1) 0.0
2 (0, 1) 0 0.1 1 −1 (0.1, 0.0) −0.1
2 (1, 1) 1 0.0 0 +1 (0.2, 0.1) 0.0
3 (0, 1) 0 0.1 1 −1 (0.2, 0.0) −0.1
3 (1, 0) 0 0.1 1 −1 (0.1, 0.0) −0.2
3 (1, 1) 1 −0.1 0 +1 (0.2, 0.1) −0.1
4 (1, 0) 0 0.1 1 −1 (0.1, 0.1) −0.2
4 (1, 1) 1 0.0 0 +1 (0.2, 0.2) −0.1
5 (0, 1) 0 0.1 1 −1 (0.2, 0.1) −0.2
6 (all) 0 (0.2, 0.1) −0.2

With w = (0.2, 0.1) and b = −0.2, the boundary is the line 0.2·x₁ + 0.1·x₂ − 0.2 = 0: only the point (1, 1) lands on the positive side (z = 0.1); the other three give z ≤ 0. Note that the solution is not unique (w = (1, 1), b = −1.5 also works): the perceptron finds a separating line, not the best one.

5.3 NovaMarket urgent orders

A less toy-like case. In the Zaragoza warehouse, an order is considered urgent if it has to go out in the next shift; that depends on the hours left until the delivery commitment and on the distance to the destination (the further away, the earlier it has to leave). We simulate 200 orders with the hidden rule "urgent if hours < 12 + distance/10", which is linear, and check that the perceptron discovers it:

rng = np.random.default_rng(42)
n = 200
hours = rng.uniform(2, 72, n)          # hours until the delivery commitment
distance = rng.uniform(5, 400, n)      # km from the Zaragoza warehouse
urgent = (hours < 12 + distance / 10).astype(int)   # hidden rule (linear)
X = np.column_stack([hours, distance])
print("Share of urgent orders:", urgent.mean())

X_std = (X - X.mean(axis=0)) / X.std(axis=0)   # standardise, as in 04-03
w, b, hist = train_perceptron(X_std, urgent, rate=0.1, epochs=100)
print("w =", w.round(3), " b =", round(b, 3), " epochs:", len(hist), " errors/epoch:", hist)
print("Accuracy:", (step(X_std @ w + b) == urgent).mean())

# Where is the boundary in original units? We solve for hours at three distances
mh, md = X.mean(axis=0); sh, sd = X.std(axis=0)
for d in (50, 200, 350):
    h_limit = mh - sh / w[0] * (w[1] * (d - md) / sd + b)
    print(f"distance {d} km -> urgent if fewer than {h_limit:.1f} h remain (true rule: {12 + d/10:.1f} h)")

Output:

Share of urgent orders: 0.475
w = [-0.862  0.479]  b = -0.2  epochs: 11  errors/epoch: [16, 9, 9, 3, 8, 6, 6, 5, 6, 4, 0]
Accuracy: 1.0
distance 50 km -> urgent if fewer than 17.3 h remain (true rule: 17.0 h)
distance 200 km -> urgent if fewer than 31.8 h remain (true rule: 32.0 h)
distance 350 km -> urgent if fewer than 46.3 h remain (true rule: 47.0 h)

Reading: in 11 epochs the perceptron classifies the 200 orders without error, and the line it has found matches the hidden rule almost exactly. The weight of hours is negative (more hours, less urgent) and that of distance positive, as expected. Two practical observations: (1) the number of errors per epoch does not decrease monotonically (16, 9, 9, 3, 8...): the rule fixes one example and unsettles others; it only guarantees convergence in the end; (2) we standardised the inputs: without scaling, with rate=0.01, after 200 epochs it was still making about 16 errors per epoch, because a step of 0.1 in the weight of distance (which ranges from 5 to 400) moves the boundary far more than the same step in hours. The lesson of 04-03 (scale your inputs) matters even more with networks.

  1. The XOR problem: the limit of one neuron and the two-layer solution

Let us now try the XOR gate (exclusive or: 1 if the inputs differ):

y_xor = np.array([0, 1, 1, 0])
w, b, hist = train_perceptron(X_gates, y_xor, rate=0.1, epochs=20)
print("XOR: errors per epoch =", hist)
print("Prediction:", step(X_gates @ w + b))

Output:

XOR: errors per epoch = [2, 3, 4, 4, 4, 4, 4, 4, 4, 4, 4, 4, 4, 4, 4, 4, 4, 4, 4, 4]
Prediction: [1 1 0 0]

It never converges. It is not a bug in the code: draw the four points on the plane; the ones worth 1, (0,1) and (1,0), sit in opposite corners, and the ones worth 0, (0,0) and (1,1), in the other two. There is no straight line that leaves the two ones on one side and the two zeros on the other. Since a neuron only knows how to draw lines, XOR is out of its reach. This is what Minsky and Papert proved rigorously in Perceptrons (1969), and, combined with the lack of a method to train several layers, it froze research on networks for fifteen years (01-01).

The solution was known in theory: chain neurons together. XOR can be written as "(x₁ OR x₂) AND NOT (x₁ AND x₂)". Each piece is linear, so two neurons in a first layer (one computes OR, the other AND) and a third one that combines them solve it:

x₁ x₂ h₁ = step(x₁ + x₂ − 0.5) (OR) h₂ = step(x₁ + x₂ − 1.5) (AND) y = step(h₁ − h₂ − 0.5) True XOR
0 0 step(−0.5) = 0 step(−1.5) = 0 step(−0.5) = 0 0
0 1 step(0.5) = 1 step(−0.5) = 0 step(0.5) = 1 1
1 0 step(0.5) = 1 step(−0.5) = 0 step(0.5) = 1 1
1 1 step(1.5) = 1 step(0.5) = 1 step(−0.5) = 0 0
def xor_net(x1, x2):
    h1 = step(x1 + x2 - 0.5)      # OR neuron   (weights 1,1; bias -0.5)
    h2 = step(x1 + x2 - 1.5)      # AND neuron  (weights 1,1; bias -1.5)
    return step(h1 - h2 - 0.5)    # output neuron: OR and not AND (weights 1,-1; bias -0.5)

for x1 in (0, 1):
    for x2 in (0, 1):
        print(x1, x2, "->", xor_net(x1, x2))

Output: 0 0 -> 0, 0 1 -> 1, 1 0 -> 1, 1 1 -> 0. The conceptual key: the first layer transforms the inputs into a new space (h₁, h₂) in which the problem is linearly separable (the points become (0,0)→0, (1,0)→1, (1,0)→1, (1,1)→0, and a line separates them). That is what every hidden layer of a network will do: rewrite the data until the last neuron can decide with a straight line. What was not known in 1969 was how to find those weights automatically rather than by hand; the answer, backpropagation, arrived in the 1980s and we will see it in 05-03.

  1. What a neural network is: input, hidden and output layers

A neural network (in its basic form, a multilayer perceptron or MLP) is a set of neurons organised in layers, where the output of each layer is the input of the next:

  • Input layer: computes nothing; it is the features (the 21 columns of the returns predictor, the 256 pixels of a 16×16 photo, etc.).
  • Hidden layers: one or more layers of neurons, each connected to all the outputs of the previous layer (hence they are called "dense" or "fully connected"). They learn intermediate representations, like h₁ and h₂ in the XOR.
  • Output layer: produces the prediction; one neuron with a sigmoid for binary classification, as many as there are classes for multiclass, one linear neuron for regression (details in 05-02).
flowchart LR
    subgraph E["Input layer"]
        x1["x₁"]; x2["x₂"]; x3["x₃"]
    end
    subgraph H["Hidden layer"]
        h1["h₁"]; h2["h₂"]; h3["h₃"]; h4["h₄"]
    end
    subgraph S["Output layer"]
        y["ŷ"]
    end
    x1 --> h1 & h2 & h3 & h4
    x2 --> h1 & h2 & h3 & h4
    x3 --> h1 & h2 & h3 & h4
    h1 & h2 & h3 & h4 --> y

The parameters of the network are all the weights and biases: in the diagram, 3×4 + 4 = 16 in the hidden layer and 4×1 + 1 = 5 in the output layer, 21 in total. The hyperparameters (04-01) are the design decisions: how many layers, how many neurons, which activation. One essential detail: if the hidden layers had no activation function (or a linear one), stacking layers would be pointless, because a composition of linear functions is another linear function and we would be back to a single line. The non-linearity of the activation is what gives layers their power.

  1. The universal approximator and the linear model vs network comparison

A theoretical result from the late 1980s (Cybenko, Hornik) says that a network with a single, sufficiently wide hidden layer and a non-linear activation can approximate any reasonable continuous function to any desired precision. It is the universal approximation theorem. Intuition: each hidden neuron with a sigmoid is a "soft step" placed somewhere in the input space; adding up many steps of different heights and positions you can draw any curve, just as with enough rectangular blocks you can approximate any silhouette. Two important caveats the theorem does not cover: it does not say how many neurons are needed (it may be an enormous number) nor how to find the weights (that is training). In practice, deeper networks (several layers) approximate the same thing with far fewer neurons than a single very wide layer, and that is where the "deep" in deep learning comes from (05-04).

Aspect Linear model (logistic regression) Neural network (MLP)
Decision boundary One hyperplane Any shape (curves, disconnected regions)
Features Whatever you give it; interactions must be created by hand (04-03) Hidden layers learn combinations and interactions
Parameters One per column + bias (22 in the returns predictor) Hundreds, thousands or millions
Interpretability High: one coefficient per variable Low: hidden weights have no direct meaning
Data needed Little A lot (more parameters → more risk of overfitting, 04-06)
Training Fast, single optimum (convex loss) Slower, many local optima, requires care (05-03)
Where it shines Tabular data, small datasets, need to explain Text, images, audio, signals; abundant data

Diego, reading the table, asks the right question: "so for returns we stick with logistic regression?". Probably yes, and in 05-03 we will check it with numbers; the network earns its place in the problems where logistic regression cannot even get started: the reviews and the photos.

  1. A glimpse of the destination: the same network in PyTorch

To show you where we are heading, this is how the two-layer network that solves XOR (two inputs → two hidden neurons → one output) is defined in PyTorch, with a sigmoid instead of a step so that it can be trained in 05-03:

import torch
import torch.nn as nn

torch.manual_seed(0)
net = nn.Sequential(
    nn.Linear(2, 2),      # hidden layer: 2 inputs -> 2 neurons (2x2 weights + 2 biases)
    nn.Sigmoid(),         # activation of the hidden layer
    nn.Linear(2, 1),      # output layer: 2 -> 1 (1x2 weights + 1 bias)
    nn.Sigmoid())         # output between 0 and 1
print(net)
print("Parameters:", sum(p.numel() for p in net.parameters()))

# Untrained, PyTorch has set random weights. We copy the solution of section 6 by hand,
# multiplied by 20 so that the sigmoid behaves almost like a step:
with torch.no_grad():
    net[0].weight[:] = 20 * torch.tensor([[1.0, 1.0], [1.0, 1.0]])   # OR and AND
    net[0].bias[:]   = 20 * torch.tensor([-0.5, -1.5])
    net[2].weight[:] = 20 * torch.tensor([[1.0, -1.0]])              # OR and not AND
    net[2].bias[:]   = 20 * torch.tensor([-0.5])

X_t = torch.tensor(X_gates, dtype=torch.float32)
print(net(X_t).detach().numpy().round(3).ravel())

Output:

Sequential(
  (0): Linear(in_features=2, out_features=2, bias=True)
  (1): Sigmoid()
  (2): Linear(in_features=2, out_features=1, bias=True)
  (3): Sigmoid()
)
Parameters: 9
[0. 1. 1. 0.]

nn.Sequential chains layers; nn.Linear(2, 2) is a dense layer that computes z = W·x + b for 2 neurons (4 weights + 2 biases); nn.Sigmoid() applies the activation; 9 parameters in total (4 + 2 + 2 + 1). With the hand-set weights the network reproduces XOR. What we have not done, and it is the goal of the rest of the module, is have the network find those weights on its own: that requires choosing the architecture and activations well (05-02) and training with gradient descent and backpropagation (05-03).

Common Mistakes and Tips

  • Believing that "neural network" means imitating the brain. It is a flexible mathematical function with a historical name. Avoid explaining it that way to Diego: it creates the wrong expectations (remember the over-hype of 01-01).
  • Forgetting the bias. Without b, the boundary passes through the origin and many trivial problems (such as AND) cannot be solved. In PyTorch nn.Linear includes it by default (bias=True).
  • Training the perceptron without scaling the inputs. As we saw with distance (5-400) against hours (2-72), convergence suffers or never arrives. Always standardise.
  • Expecting the perceptron to converge on non-separable data. If the classes overlap (the norm in real data, such as returns), the rule oscillates forever. That is why the step was abandoned in favour of smooth activations and a loss to be minimised (05-03).
  • Stacking layers without a non-linear activation. A classic mistake when starting with PyTorch: nn.Sequential(nn.Linear(2, 2), nn.Linear(2, 1)) is mathematically a single linear neuron and does not solve XOR no matter how many layers you add.
  • Confusing parameters with hyperparameters. The weights and biases are learned by the network; the number of layers and neurons is chosen by you (and tuned with validation, 04-06).

Exercises

Exercise 1. Train the perceptron with train_perceptron on the NAND gate (y = [1, 1, 1, 0]). Then build XOR without an AND layer, using the identity XOR = (x₁ OR x₂) AND (x₁ NAND x₂): write a function xor_net_2(x1, x2) with the OR and NAND neurons in the first layer and an AND neuron at the output (weights and biases by hand) and check the truth table.

Exercise 2. Replace the hidden rule for urgent orders with a non-linear one: urgent = ((hours < 24) != (distance > 200)).astype(int) (a "continuous" XOR: urgent if there is little time left or it is far away, but not both). Train the perceptron for 100 epochs on the standardised inputs. Does it converge? What maximum accuracy does it reach? Explain why using section 6.

Exercise 3. A dense network has 21 inputs, one hidden layer of 8 neurons and one output. (a) Work out by hand how many parameters it has. (b) Explain why, if the hidden layer had no activation, the network would be equivalent to a logistic regression with 22 parameters. (c) Check (a) by defining it with nn.Sequential and sum(p.numel() for p in net.parameters()).

Solutions

Solution 1. The perceptron learns NAND in 4 epochs (with rate=0.1 it gets w = (−0.2, −0.1), b = 0.2, which gives z = 0.2, 0.1, 0.0, −0.1 → 1, 1, 1, 0). The network:

def xor_net_2(x1, x2):
    h1 = step(x1 + x2 - 0.5)          # OR
    h2 = step(-x1 - x2 + 1.5)         # NAND: 1 except when both are 1
    return step(h1 + h2 - 1.5)        # AND of the two hidden neurons

Table: (0,0) → h = (0, 1) → 0; (0,1) → (1, 1) → 1; (1,0) → (1, 1) → 1; (1,1) → (1, 0) → 0. Any decomposition of XOR into chained linear functions works; what does not work is any single linear function.

Solution 2. It does not converge: after 100 epochs it is still making between 75 and 90 errors per epoch (out of 200) and the final accuracy hovers around 40-60 % depending on the epoch you stop at, i.e. no better than tossing a coin (urgent orders are 50.5 %). The reason is that of section 6: the four "corners" (little time/near → 0, little time/far → 1, plenty of time/near → 1, plenty of time/far → 0) form an XOR and no straight line separates the two urgent regions from the two non-urgent ones. It would take a hidden layer that, as in the XOR, first transforms the inputs (for instance, two neurons detecting "little time" and "far away") and an output that combines them.

Solution 3. (a) Hidden layer: 21 × 8 weights + 8 biases = 176; output: 8 × 1 + 1 = 9; total 185. (b) Without an activation, the output would be W₂·(W₁·x + b₁) + b₂ = (W₂·W₁)·x + (W₂·b₁ + b₂), i.e. a single weighted sum of the 21 inputs plus a bias: 21 effective weights + 1 bias, followed by the final sigmoid. The 185 parameters would be "wasted" representing a function that only has 22 degrees of freedom. (c) nn.Sequential(nn.Linear(21, 8), nn.Sigmoid(), nn.Linear(8, 1), nn.Sigmoid()) gives 185. This will be, with more neurons and a different activation, the network we will build in 05-02 for the returns predictor.

Conclusion

In this lesson we have learned that the artificial neuron is a weighted sum plus a bias passed through an activation function (step or sigmoid), that it can only draw linear boundaries, and that the logistic regression of 04-04 was already a neuron with a sigmoid: the difference between module 4 and this one is not the building block but how many are stacked. We implemented Rosenblatt's perceptron in numpy with its rule w ← w + η·e·x, watched it converge on AND, OR and the NovaMarket urgent orders (recovering the hidden rule almost exactly), and fail on XOR, the limit that Minsky and Papert pointed out in 1969. The way out is to stack neurons in layers: the hidden layer transforms the inputs into a space where the problem is separable, and that is why a network with non-linear activations is a universal approximator. We closed with the same network written in PyTorch with nn.Sequential, with the weights set by hand.

Two questions remain that this module is going to answer. First: how do you design a real network, with how many layers and neurons, which activations and which output depending on the problem? That is the topic of the next lesson, Neural Network Architecture, where we will build in PyTorch the MLP of the returns predictor on the 21 columns of module 4 and count its parameters. Second: how does the network find its weights without anyone writing them by hand? That is the topic of 05-03.

Fundamentals of Artificial Intelligence (AI)

Module 1: Introduction to Artificial Intelligence

Module 2: Basic Principles of AI

Module 3: Algorithms in AI

Module 4: Machine Learning

Module 5: Neural Networks and Deep Learning

Module 6: Logic and Expert Systems

Module 7: Tools and Programming Languages in AI

Module 8: Projects and Case Studies

Module 9: Exercises and Practice

Module 10: Additional Resources

© Copyright 2026. All rights reserved