In the previous module you met the artificial neuron at a conceptual level through the "purchasing committee" analogy: several inputs, each with its own weight, a bias, and a final decision. Now that your tecnomarket-dl/ environment is ready, it's time to move from intuition to the real mechanics. In this lesson we formalize that neuron under its historical name — the perceptron — see how it learns by adjusting its weights, implement it from scratch in numpy to build TecnoMarket's first fraud filter, discover its great limitation (it can only separate data linearly), and understand why stacking layers — the multilayer perceptron, or MLP — breaks through that barrier. It is the foundation everything else in this course is built on.
Contents
- From the purchasing committee to the formal perceptron
- The perceptron learning rule
- Implementation in numpy: TecnoMarket's fraud filter
- The fatal limitation: linear separability and the XOR problem
- The Multilayer Perceptron (MLP): stacking layers to solve XOR
From the purchasing committee to the formal perceptron
In the basic concepts lesson we saw that an artificial neuron works like a member of a committee: it listens to several opinions (inputs), gives each one more or less importance (weights), has a predisposition of its own (bias), and finally decides. The perceptron, proposed by Frank Rosenblatt in 1958 (we placed it in the history lesson), is exactly that written out in mathematics.
Given an input vector $x = (x_1, x_2, \dots, x_n)$, a weight vector $w = (w_1, w_2, \dots, w_n)$ and a bias $b$:
Step 1 — Weighted sum:
$$z = w_1 x_1 + w_2 x_2 + \dots + w_n x_n + b = w \cdot x + b$$
Step 2 — Step function:
$$\hat{y} = \begin{cases} 1 & \text{if } z \geq 0 \ 0 & \text{if } z < 0 \end{cases}$$
In other words: if the accumulated evidence clears the threshold, the neuron "fires" (output 1); otherwise it stays silent (output 0). The step function is the original activation from the 1950s; in the next lesson we'll see that today we use smoother functions, but for the classic perceptron the step is enough — and true to history.
Note the role of the bias $b$: it shifts the decision threshold. Without a bias, the neuron could only decide "fire if $w \cdot x \geq 0$"; with a bias, it can decide "fire if $w \cdot x \geq -b$" — that is, it can be more demanding or more lenient.
| Element | Committee analogy (01-04) | Symbol | What it does |
|---|---|---|---|
| Inputs | Opinions reaching the committee | $x_i$ | The problem's data |
| Weights | Credibility of each opinion | $w_i$ | How much each input influences |
| Bias | The committee's predisposition | $b$ | Shifts the decision threshold |
| Weighted sum | Balance of opinions | $z$ | Total accumulated evidence |
| Step | Final yes/no vote | $\hat{y}$ | Binary decision (0 or 1) |
The perceptron learning rule
So far the weights were fixed. Rosenblatt's great contribution was a rule to learn them automatically from examples — the first concrete embodiment of the "training = measure the error and correct" principle we saw with the shower faucet analogy.
The rule is surprisingly simple. For each training example $(x, y)$ where $y$ is the correct label (0 or 1):
- Compute the prediction $\hat{y}$ with the current perceptron.
- Compute the error: $e = y - \hat{y}$ (it can be $-1$, $0$ or $+1$).
- Update every weight and the bias:
$$w_i \leftarrow w_i + \eta \cdot e \cdot x_i \qquad b \leftarrow b + \eta \cdot e$$
where $\eta$ (eta) is the learning rate, a small number like 0.1 that controls the size of each correction — exactly the "small turn of the faucet" that keeps us from overshooting from freezing to scalding.
Let's interpret the three possible cases:
- Correct prediction ($e = 0$): nothing gets touched. If it works, don't fix it.
- False negative ($y=1$, $\hat{y}=0$, $e=+1$): the neuron should have fired and didn't. The weights of the active inputs are increased so that next time the sum comes out higher.
- False positive ($y=0$, $\hat{y}=1$, $e=-1$): it fired when it shouldn't have. The weights of the active inputs are decreased.
A classic theorem (the perceptron convergence theorem) guarantees that, if the data can be separated by a straight line (or a plane, in more dimensions), this rule finds a solution in a finite number of steps. Hold on to that "if": we'll come back to it in section 4.
Implementation in numpy: TecnoMarket's fraud filter
Since this case is so small, we'll apply the course methodology in reverse of the usual order: straight onto fictional TecnoMarket data. We pick up the anti-fraud neuron we posed as a conceptual exercise in 01-04, and now we build it and actually train it.
The fraud team describes each order with three binary signals (1 = present, 0 = absent):
x1: the amount exceeds €500x2: the customer account is less than 24 hours oldx3: the shipping address does not match the billing address
And we have 8 historical orders labeled by hand (y = 1 means confirmed fraud):
import numpy as np
# Fictional TecnoMarket data: each row is an order [x1, x2, x3]
X = np.array([
[0, 0, 0], # normal order
[1, 0, 0], # expensive, but long-time customer and consistent addresses
[0, 1, 0], # new account, cheap purchase
[0, 0, 1], # different addresses (a gift, for example)
[1, 1, 0], # expensive + new account -> fraud
[1, 0, 1], # expensive + different addresses -> fraud
[0, 1, 1], # new account + different addresses -> fraud
[1, 1, 1], # all signals -> fraud
])
y = np.array([0, 0, 0, 0, 1, 1, 1, 1])Notice the pattern: one isolated signal is innocent; two or more signals together indicate fraud. It's an "at least 2 out of 3" rule. Let's implement the perceptron:
class Perceptron:
def __init__(self, n_inputs, learning_rate=0.1):
self.w = np.zeros(n_inputs) # initial weights at zero
self.b = 0.0 # initial bias at zero
self.eta = learning_rate
def predict(self, x):
z = np.dot(self.w, x) + self.b # weighted sum: w·x + b
return 1 if z >= 0 else 0 # step function
def train(self, X, y, epochs=10):
for epoch in range(epochs):
errors = 0
for xi, yi in zip(X, y):
e = yi - self.predict(xi) # error: -1, 0 or +1
self.w += self.eta * e * xi # perceptron rule
self.b += self.eta * e
errors += int(e != 0)
print(f"Epoch {epoch+1}: {errors} errors, "
f"w={self.w}, b={self.b:.2f}")
if errors == 0: # convergence: all correct
breakA line-by-line breakdown of what matters:
np.dot(self.w, x)computes the weighted sum $w_1x_1 + w_2x_2 + w_3x_3$ in a single operation. It's the same vector multiplication we verified in the environment check script in 01-05.e = yi - self.predict(xi)measures the error on this particular order.self.w += self.eta * e * xiapplies the correction: only the weights of the signals present in the order get modified (ifxi[j] == 0, that weight doesn't change).- An epoch (vocabulary from 01-04) is one complete pass over the 8 orders.
We train and test:
p = Perceptron(n_inputs=3)
p.train(X, y, epochs=10)
# A new order: expensive and from a freshly created account
suspicious_order = np.array([1, 1, 0])
print("Fraud?", p.predict(suspicious_order)) # -> 1Typical output (the exact values may vary slightly depending on the order of the data):
Epoch 1: 5 errors, w=[0.1 0.1 0.1], b=-0.10
Epoch 2: 3 errors, w=[0.1 0.1 0.2], b=-0.20
Epoch 3: 2 errors, w=[0.2 0.1 0.2], b=-0.20
Epoch 4: 0 errors, w=[0.2 0.1 0.2], b=-0.20
Fraud? 1The perceptron has learned, on its own, a rule equivalent to "at least 2 signals": with weights around 0.1–0.2 and a bias around $-0.2$, a single signal doesn't reach the threshold, but two do. Nobody programmed the rule: it emerged from the data. That is the qualitative leap over the conceptual exercise in 01-04, where we chose the weights by hand.
The fatal limitation: linear separability and the XOR problem
The equation $w \cdot x + b = 0$ geometrically defines a line (with 2 inputs), a plane (with 3) or a hyperplane (with more). The perceptron classifies according to which side of that boundary each point falls on. It can therefore only solve linearly separable problems: those where a single straight line is enough to separate the two classes.
Our fraud filter worked because "at least 2 out of 3 signals" is separable by a plane. But in 1969 Minsky and Papert pointed out a devastatingly simple counterexample: the XOR (exclusive or) function, which returns 1 when the inputs are different:
| $x_1$ | $x_2$ | XOR | TecnoMarket example |
|---|---|---|---|
| 0 | 0 | 0 | Regular customer, normal hours: OK |
| 0 | 1 | 1 | Just one anomaly: review |
| 1 | 0 | 1 | Just the other anomaly: review |
| 1 | 1 | 0 | Both at once: that's the pattern of a known wholesale customer, OK |
Picture the four points on a plane: the ones worth 1 sit in opposite corners ((0,1) and (1,0)), and the ones worth 0 in the other two corners ((0,0) and (1,1)). No straight line exists that leaves the ones on one side and the zeros on the other. You can verify it empirically: train the perceptron above on these 4 points and you'll see it never reaches 0 errors, no matter how many epochs you give it — the convergence theorem only applied to separable data.
This seemingly minor result contributed to the first "AI winter" we saw in the history lesson: if the perceptron can't handle XOR, how is it ever going to recognize a face?
The Multilayer Perceptron (MLP): stacking layers to solve XOR
You already know the way out of the problem from 01-04: hidden layers. A Multilayer Perceptron (MLP) chains layers of neurons together: each layer's output is the next layer's input, exactly as in the 3-4-2-1 network with 29 parameters that we counted by hand.
The key idea: each hidden neuron draws its own line, and the output neuron combines those lines into boundaries that are no longer straight. For XOR, a hidden layer of 2 neurons is enough:
graph LR
x1((x1)) --> h1((h1: OR))
x1 --> h2((h2: AND))
x2((x2)) --> h1
x2 --> h2
h1 --> s((output: h1 AND NOT h2))
h2 --> s
- Hidden neuron h1 learns "at least one input active" (OR): a separable line, a perceptron can do it.
- Hidden neuron h2 learns "both inputs active" (AND): also separable.
- The output combines them: "h1 yes, but h2 no" — that is, some input active but not both. That's XOR!
Let's verify it in numpy using hand-picked weights (we don't yet know how to train them: that requires backpropagation, the subject of the forward and backward propagation lesson):
import numpy as np
def step(z):
return (z >= 0).astype(int)
# Hidden layer: 2 neurons. Each COLUMN of W1 holds one neuron's weights.
W1 = np.array([[1.0, 1.0], # weights from x1 to h1 and h2
[1.0, 1.0]]) # weights from x2 to h1 and h2
b1 = np.array([-0.5, -1.5]) # h1 fires when sum>=0.5 (OR); h2 when >=1.5 (AND)
# Output layer: combines h1 (positive) and h2 (very negative)
W2 = np.array([[1.0], [-2.0]])
b2 = np.array([-0.5])
def mlp_xor(x):
h = step(x @ W1 + b1) # hidden layer forward pass
return step(h @ W2 + b2) # output layer forward pass
for x in [[0,0], [0,1], [1,0], [1,1]]:
print(x, "->", mlp_xor(np.array(x))[0])It works. Notice two implementation details that will keep reappearing:
x @ W1 + b1processes all the neurons in a layer at once with a matrix multiplication: each column ofW1holds one neuron's weights. It's the vectorized form of the weighted sum, and the reason GPUs accelerate deep learning so dramatically (history lesson and environment setup).- The structure is identical to the 3-4-2-1 network from 01-04, just in a 2-2-1 version. Count the parameters: hidden layer $2\times2+2 = 6$, output layer $2\times1+1 = 3$; a total of 9 parameters, using the same counting method we applied there.
One enormous question remains open: here we set the weights by hand, and the perceptron learning rule doesn't work for hidden layers (what "error" does a hidden neuron have, when nobody tells it what its correct output was?). Solving that took nearly 20 years and is called backpropagation; we'll get to it in two lessons. First we need one ingredient: replacing the step with differentiable activation functions, the subject of the next lesson.
Common Mistakes and Tips
- Forgetting the bias. Without $b$, the decision boundary is forced through the origin and many separable problems become impossible. Always include it.
- Mistaking the step output for a probability. The classic perceptron returns a bare 0 or 1, with no shade of confidence. To get probabilities you'll need the sigmoid (next lesson).
- Expecting convergence on non-separable data. If the perceptron oscillates and never reaches 0 errors, your data probably isn't linearly separable; it's not a bug in your code. Stop training on a maximum number of epochs, not only on convergence.
- Using a huge learning rate "to go faster". With the pure perceptron the effect is less severe than in modern networks, but get used to small values now (0.01–0.1): it's the gentle turn of the faucet.
- Mixing up rows and columns in the weight matrices. In the convention we'll use,
Xhas one row per example andWone column per neuron;X @ Wthen comes out with one row per example and one column per neuron. Always check dimensions with.shapebefore hunting for more exotic bugs.
Exercises
- Logic gates. Train the
Perceptronclass on the 2-input AND and OR functions (4 examples each). Verify that it converges in a few epochs and note down the final weights. Then try it with XOR and describe what you observe in the per-epoch error count. - Extended fraud filter. Add a fourth signal to the fraud dataset:
x4 = "the order uses a previously declined card". Generate all 16 possible orders, label as fraud those with 2 or more active signals, and train the perceptron. Does it converge? What weights does it learn and how do you interpret them? - Hand-built MLP for signal-NAND. Design (without training — picking weights by hand as we did with XOR) a 2-2-1 MLP with a step function that implements XNOR (the negation of XOR: it returns 1 when the inputs are equal). Hint: you can reuse the OR and AND neurons from the example and change only the output layer.
Solutions
Exercise 1. For AND: X = [[0,0],[0,1],[1,0],[1,1]], y = [0,0,0,1]; for OR, y = [0,1,1,1]. Both converge (they're separable); typical solutions starting from zeros with $\eta=0.1$: AND ends with weights close to w=[0.2, 0.1], b=-0.2 (it demands both inputs) and OR with w=[0.1, 0.1], b=-0.1 (one is enough). With XOR the per-epoch error count never settles down to 0: it oscillates (for example 2, 3, 2, 3, ...) indefinitely, confirming non-separability.
Exercise 2. The "2 or more signals out of 4" problem is still linearly separable (it's a threshold function). Generating the 16 cases with itertools.product([0,1], repeat=4) and labeling with y = (X.sum(axis=1) >= 2).astype(int), the perceptron converges; it learns weights approximately equal to one another (all signals carry the same weight, because the rule treats them symmetrically) and a negative bias such that two signals are needed to clear the threshold, for example w≈[0.2, 0.2, 0.2, 0.2], b≈-0.3.
Exercise 3. XNOR is 1 at (0,0) and (1,1). Reuse h1 = OR and h2 = AND with the same W1, b1 from the example. The output must be 1 when "none is active" (h1=0) or "both are" (h2=1): W2 = [[-1.0], [2.0]], b2 = [0.5] does the job. Verification: (0,0) → h=(0,0), z=0.5 → 1; (0,1) → h=(1,0), z=-0.5 → 0; (1,0) → same → 0; (1,1) → h=(1,1), z=1.5 → 1. Correct.
Conclusion
In this lesson we formalized the purchasing-committee neuron as a perceptron (weighted sum + step), saw the first learning rule in history, and used it to train a fraud filter for TecnoMarket from scratch in numpy. We also ran into its limit — it only separates what is linearly separable, and XOR is not — and confirmed that stacking layers (the MLP) breaks that limit, though for now by choosing the weights by hand.
Two loose ends remain: the step is an "all or nothing" function that admits neither nuance nor derivatives, and we don't know how to train hidden layers. The first is resolved in the next lesson with modern activation functions (sigmoid, tanh, ReLU, softmax); the second, right after that, with backpropagation. Step by step, we are building every piece of the first complete network you'll train at the end of this module.
Deep Learning Course
Module 1: Introduction to Deep Learning
- What is Deep Learning?
- History and evolution of Deep Learning
- Applications of Deep Learning
- Basic concepts of neural networks
- Setting up the work environment
Module 2: Neural Network Fundamentals
- Perceptron and Multilayer Perceptron
- Activation functions
- Forward and backward propagation
- Optimization and loss functions
- Your first complete neural network
Module 3: Convolutional Neural Networks (CNN)
- Introduction to CNNs
- Convolutional and pooling layers
- Popular CNN architectures
- CNN applications in image recognition
Module 4: Recurrent Neural Networks (RNN)
- Introduction to RNNs
- LSTM and GRU
- RNN applications in natural language processing
- Sequences and time series
Module 5: Advanced Deep Learning Techniques
- Generative Adversarial Networks (GAN)
- Autoencoders
- Transfer Learning
- Regularization and improvement techniques
- Attention mechanisms and Transformers
Module 6: Tools and Frameworks
- Introduction to TensorFlow
- Introduction to PyTorch
- Framework comparison
- Development environments and additional resources
- Saving, loading and deploying models
Module 7: Hands-On Projects
- Image classification with CNNs
- Text generation with RNNs
- Anomaly detection with Autoencoders
- Building a GAN for image generation
- Fine-tuning a pretrained model
