At the end of the previous module we trained a dense 784-128-64-10 network that reached ~97.7% accuracy on MNIST, and we closed with a warning: to get there we had to flatten each 28×28-pixel image into a vector of 784 numbers, destroying all the information about which pixel sat next to which. With tiny black-and-white digits we could get away with it; with the real product photos that TecnoMarket sellers upload, we can't. In this lesson we'll see, with actual numbers, why dense networks don't scale to realistic images, and we'll meet the architecture that has dominated computer vision since 2012: the convolutional neural network (CNN). You'll understand its three fundamental ideas, work through a convolution by hand on a small matrix, and see the overall anatomy of a complete CNN.
Contents
- Why dense networks fail on images
- The three key ideas behind CNNs
- The convolution operation step by step
- General anatomy of a CNN
Why dense networks fail on images
Problem 1: parameter explosion
Recall the parameter-counting method from module 2: a dense layer with n inputs and m neurons has n × m weights plus m biases.
In MNIST, the image was 28×28 pixels in grayscale:
- Flattened input: 28 × 28 = 784 values.
- First dense layer with 128 neurons: 784 × 128 + 128 = 100,480 parameters.
Manageable. But a TecnoMarket product photo, even shrunk to the standard size many vision models use, is 224×224 pixels with 3 color channels (red, green, blue):
- Flattened input: 224 × 224 × 3 = 150,528 values.
- First dense layer with 128 neurons: 150,528 × 128 + 128 = 19,267,712 parameters.
Over 19 million parameters in the first layer alone, almost 200 times more than for MNIST. And 224×224 is a small image: the photo of a refrigerator uploaded by a seller might be 2000×1500 pixels. Let's check it with Python:
# Parameter count of the first dense layer for different image sizes
sizes = [
("MNIST 28x28x1", 28 * 28 * 1),
("Product 224x224x3", 224 * 224 * 3),
("Seller photo 1024x768x3", 1024 * 768 * 3),
]
neurons = 128
for name, inputs in sizes:
params = inputs * neurons + neurons
print(f"{name:28s} -> input: {inputs:>9,} "
f"-> 1st dense layer: {params:>13,} params")Output:
MNIST 28x28x1 -> input: 784 -> 1st dense layer: 100,480 params Product 224x224x3 -> input: 150,528 -> 1st dense layer: 19,267,712 params Seller photo 1024x768x3 -> input: 2,359,296 -> 1st dense layer: 301,990,016 params
With a seller's photo, 302 million parameters in a single layer. More parameters means more memory, more compute, more data needed to train without overfitting (remember the mild overfitting we already spotted in 02-05 with just 100,000 parameters), and slower everything. This road does not scale.
Problem 2: loss of spatial structure
The second problem is subtler but just as serious. When you flatten the image, pixel (0, 0) and pixel (0, 1) — immediate neighbors — end up at positions 0 and 1 of the vector, but pixel (1, 0), which sits directly below the first one, ends up at position 28 (or 224, or 1024...). The dense network has no idea those pixels were neighbors: to it, the input vector is a list of numbers with no geometry.
Practical consequences for TecnoMarket:
- A vertical edge (the side of a laptop) is a local pattern of neighboring pixels; the dense network cannot exploit that directly.
- If the product appears shifted 10 pixels to the right in the photo, to the dense network it is a completely different input: it would have to relearn the same pattern at every possible position.
- Every neuron in the first layer looks at all the pixels at once, when most useful patterns (edges, corners, textures) occupy tiny regions.
We need an architecture that respects the 2D structure of the image and recognizes a pattern regardless of where it appears. That is exactly what a CNN does.
The three key ideas behind CNNs
CNNs rest on three principles that attack the two problems above head-on.
- Local connectivity
Instead of connecting every neuron to every pixel, each neuron in a convolutional layer looks only at a small window of the image (for example, 3×3 pixels), called its receptive field. It's the same logic you would use yourself: to decide whether there's an edge in one corner of the photo, you don't need to look at the whole photo, just that corner.
Going back to the purchasing committee analogy from module 1: instead of a committee where every member reads the entire catalog, now each member specializes in inspecting one small area of the product.
- Shared weights
Here's the big efficiency trick: the same window of weights (the filter or kernel) slides across the whole image. If a 3×3 filter has learned to detect vertical edges, it will detect vertical edges in the top-left corner, in the center, and at the bottom right, with the same 9 weights.
- A 3×3 filter on a grayscale image: 9 weights + 1 bias = 10 parameters, no matter whether the image is 28×28 or 1024×768.
- Compare that with the 19 million of the dense layer on the 224×224×3 image.
On top of that, this gives the network an extremely valuable property called translation equivariance: if the microwave shifts within the photo, its activation pattern shifts with it, but it still gets detected. Nothing has to be relearned per position.
- Hierarchy of features
In 01-01 we defined deep learning as learning a hierarchy of representations, from the simple to the abstract. CNNs are the clearest embodiment of that idea. By stacking convolutional layers:
| Level | What it detects | Example in a product photo |
|---|---|---|
| Early layers | Edges and colors | Vertical/horizontal lines, light-dark transitions |
| Intermediate layers | Textures and motifs | Metal grille, plastic surface, wood |
| Upper-middle layers | Object parts | A screen, a handle, a keyboard, a refrigerator door |
| Final layers | Whole objects | "This is a laptop", "this is a coffee maker" |
Each layer combines the detections of the previous one: edges combine into textures, textures into parts, parts into objects. Nobody programs this hierarchy; it emerges from training, just as in the 3-8-1 network for the fictitious orders the weights learned on their own which combinations gave away a fraud.
The convolution operation step by step
Let's look at the exact mechanics of how a filter "slides" across an image. We'll work with a 5×5 toy image (imagine a tiny grayscale crop of a product photo, with dark background on the left and the bright product on the right) and a 3×3 filter.
Input image (5×5):
There is a vertical edge between the column of zeros (dark) and the columns of ones (bright).
Vertical edge detector filter (3×3):
This filter responds strongly when its left half sees different values than its right half, that is, when there is a vertical transition. (It's a version of the classic Prewitt filter; the beauty of CNNs is that these values are not hand-designed but learned through backpropagation like any other weight.)
The procedure: we place the filter over the top-left corner of the image, multiply each weight by the pixel underneath it, add up the 9 products, and write down the result. Then we slide the filter one position to the right and repeat; when we reach the end of the row, we move down one position.
Position (0,0), the filter covers the first 3 rows and columns:
Position (0,1), we slide over one column:
Position (0,2):
Since every row of the image is identical, the result repeats as we move down. The resulting feature map is 3×3 (a 3×3 filter over a 5×5 image fits in 3×3 positions; in 03-02 we'll see the general formula and how to control this size):
How to read it: the values that are large in magnitude (here -3) mark where the vertical edge is; the 0 marks a uniform area. The feature map is literally a "map of where the pattern this filter is looking for appears". The sign depends on the orientation of the contrast (dark→bright here); after the convolution an activation such as ReLU (02-02) is applied, and filters with opposite signs capture each orientation.
Let's verify it with numpy:
import numpy as np
image = np.array([
[0, 0, 1, 1, 1],
[0, 0, 1, 1, 1],
[0, 0, 1, 1, 1],
[0, 0, 1, 1, 1],
[0, 0, 1, 1, 1],
], dtype=float)
kernel = np.array([
[1, 0, -1],
[1, 0, -1],
[1, 0, -1],
], dtype=float)
def convolve2d(image, kernel):
"""'Valid' convolution: the filter is only placed where it fits entirely."""
rows, cols = image.shape # image rows and columns
k_rows, k_cols = kernel.shape # filter rows and columns
output = np.zeros((rows - k_rows + 1, cols - k_cols + 1))
for i in range(output.shape[0]):
for j in range(output.shape[1]):
window = image[i:i + k_rows, j:j + k_cols] # 3x3 crop of the image
output[i, j] = np.sum(window * kernel) # element-wise product and sum
return output
print(convolve2d(image, kernel))Output:
It matches our hand calculation. Look closely at the loop: it is exactly "place, multiply, sum, slide". Every convolutional layer, no matter how large, does this very thing (in vectorized form and with many filters at once).
Key points worth pinning down:
- One filter = one pattern detector. A real convolutional layer has many filters (32, 64...), each looking for a different pattern, and produces one feature map per filter.
- The filter values are trainable weights: the network adjusts them with the same backpropagation and gradient descent machinery you already mastered in 02-03 and 02-04.
- After the sum, a bias is added and an activation (typically ReLU) is applied, just as in a dense neuron. A conv layer is, at heart, the same old neuron but with local connectivity and shared weights.
General anatomy of a CNN
A complete CNN alternates two kinds of blocks: a convolutional stage that extracts increasingly abstract features, and a final dense stage that uses those features to classify. Between convolutions, pooling layers are interleaved to shrink the maps (we'll study them in depth in 03-02; for now it's enough to know they "summarize" regions to condense the information).
flowchart LR
A["Product photo<br/>224x224x3"] --> B["Conv + ReLU<br/>detects edges"]
B --> C["Pooling<br/>shrinks size"]
C --> D["Conv + ReLU<br/>detects textures/parts"]
D --> E["Pooling<br/>shrinks size"]
E --> F["Flatten"]
F --> G["Dense + ReLU"]
G --> H["Dense softmax<br/>category: laptop,<br/>coffee maker, monitor..."]
Note the order of ideas:
- Conv + pooling blocks (repeated 2, 3, or many more times): they build the hierarchy edges → textures → parts → objects, preserving the spatial structure the whole way.
- Flatten: only once the maps are small and contain abstract features ("there's a screen", "there are keys") do we flatten. Flattening here is no longer a crime: the exact position matters little by now — what matters is what has been detected.
- Final dense layers: they combine those detections to produce the classification, exactly like the 784-128-64-10 network from 02-05, but receiving meaningful features instead of raw pixels.
This division of labor is the key: the convolutional part acts as a learned feature extractor and the dense part as the classifier. For the flagship project of this module — automatically classifying the photos that TecnoMarket sellers upload — this will be the mental template; we'll assemble the technique piece by piece over the coming lessons, and the full guided project arrives in 07-01.
Common Mistakes and Tips
- Thinking the filters are designed by hand. The vertical edge filter we used is illustrative; in a real CNN the filter values are initialized at random and learned during training. The network decides on its own which detectors it needs.
- Confusing the filter with the feature map. The filter is the weights (what the network learns); the feature map is the output of applying that filter to one specific image (it changes with every image).
- Believing CNNs get rid of dense layers. They don't: almost every classification CNN ends in dense layers. Convolutions extract features; dense layers classify.
- Forgetting that convolution is the same old neuron. Multiply by weights, sum, add a bias, activate: it's the perceptron from 02-01 with two clever restrictions (locality and shared weights). If backprop made sense to you in 02-03, you already know how a CNN is trained.
- Tip: whenever you meet a new CNN architecture, always ask yourself "where does the feature extractor end and where does the classifier begin?". That reading will serve you throughout the course, especially in transfer learning (05-03).
Exercises
Exercise 1: the dense layer's bill
TecnoMarket sellers upload photos that the system resizes to 128×128 in color (3 channels). If we connected that flattened image to a first dense layer of 256 neurons, how many parameters would that layer have? And how many parameters does a 5×5 convolutional filter applied to that same image in grayscale (1 channel) have? Work both out by hand and compare.
Exercise 2: convolution by hand
Apply by hand (and then check with the lesson's convolve2d function) the horizontal edge detector filter to this 4×4 image, which has a bright band on top and a dark one below:
What size is the resulting feature map? At which positions does it respond most strongly, and why?
Exercise 3: classify the three ideas
For each situation, say which of the three key CNN ideas (local connectivity, shared weights, hierarchy of features) explains it best:
- The same "metal corner" detector works whether the toaster is centered or at the edge of the photo.
- A first-layer neuron only processes a 3×3-pixel area.
- Layer 8 of the network responds to "refrigerator doors" by combining "straight edge", "smooth surface", and "handle" detections from earlier layers.
Solutions
Exercise 1:
- Dense layer: input = 128 × 128 × 3 = 49,152 values. Parameters = 49,152 × 256 + 256 = 12,583,168 (about 12.6 million).
- 5×5 conv filter on 1 channel: 5 × 5 = 25 weights + 1 bias = 26 parameters.
- Comparison: the dense layer needs ~484,000 times more parameters than one filter. Even if a real conv layer uses 32 or 64 filters (26 × 64 = 1,664 parameters), the difference is still four orders of magnitude. Moreover, the filter's parameter count does not depend on the image size.
Exercise 2:
The resulting map is (4−3+1) × (4−3+1) = 2×2. Computing each position:
- Position (0,0): window = rows 0-2, columns 0-2 → (1+1+1) + 0 − (0+0+0) = 3
- Position (0,1): same structure → 3
- Position (1,0): window = rows 1-3 → (1+1+1) + 0 − (0+0+0) = 3
- Position (1,1): → 3
Every position responds strongly because every possible 3×3 window contains the bright→dark transition: the horizontal band crosses the entire image, and with only 4 rows, any 3-row window includes it. The filter's top row lands on bright values and its bottom row on dark ones (or on the transition), and that difference is exactly what the filter measures. In a larger image, with uniform areas far from the edge, we would see zeros away from the band and high values only along the edge.
Exercise 3:
- Shared weights: the same filter is applied at every position, so the pattern is detected wherever it is (translation equivariance).
- Local connectivity: the neuron's receptive field is a small window, not the whole image.
- Hierarchy of features: deep layers compose detections from earlier layers to represent increasingly abstract concepts.
Conclusion
We have closed the loop we opened at the end of module 2: dense networks don't scale to real images because they explode in parameters (19 million in a single layer for a modest 224×224×3 photo) and because flattening destroys the image's geometry. CNNs solve both problems with three ideas: local connectivity (each neuron looks at a small window), shared weights (the same filter slides across the whole image, with only tens of parameters), and hierarchy of features (edges → textures → parts → objects, the "hierarchy of representations" from 01-01 turned into architecture). You've run the convolution operation by hand and in numpy, and you already know how to read the anatomy of a CNN: conv+pooling blocks that extract features and dense layers that classify.
But we've deliberately left loose ends: why did the feature map come out 3×3 and not 5×5? Can we keep the image from shrinking? What exactly does pooling do, and how many parameters does a convolutional layer with many filters and color channels have? All of that — stride, padding, channels, size formulas, and the Keras Conv2D and MaxPooling2D layers — is the territory of the next lesson: convolutional and pooling layers.
Deep Learning Course
Module 1: Introduction to Deep Learning
- What is Deep Learning?
- History and evolution of Deep Learning
- Applications of Deep Learning
- Basic concepts of neural networks
- Setting up the work environment
Module 2: Neural Network Fundamentals
- Perceptron and Multilayer Perceptron
- Activation functions
- Forward and backward propagation
- Optimization and loss functions
- Your first complete neural network
Module 3: Convolutional Neural Networks (CNN)
- Introduction to CNNs
- Convolutional and pooling layers
- Popular CNN architectures
- CNN applications in image recognition
Module 4: Recurrent Neural Networks (RNN)
- Introduction to RNNs
- LSTM and GRU
- RNN applications in natural language processing
- Sequences and time series
Module 5: Advanced Deep Learning Techniques
- Generative Adversarial Networks (GAN)
- Autoencoders
- Transfer Learning
- Regularization and improvement techniques
- Attention mechanisms and Transformers
Module 6: Tools and Frameworks
- Introduction to TensorFlow
- Introduction to PyTorch
- Framework comparison
- Development environments and additional resources
- Saving, loading and deploying models
Module 7: Hands-On Projects
- Image classification with CNNs
- Text generation with RNNs
- Anomaly detection with Autoencoders
- Building a GAN for image generation
- Fine-tuning a pretrained model
