In the previous lesson you performed a convolution by hand and saw that a 3×3 filter over a 5×5 image produced a 3×3 map, but we left several questions open: why exactly that size? How do you convolve a color image with 3 channels? What does the pooling in the anatomy diagram actually do? This lesson answers all of that: it is the "engineering" lesson of CNNs, where you'll learn to configure a convolutional layer (filters, stride, padding, channels), to predict the size of its output, and to count its parameters — essential skills for reading and designing any architecture, including the famous ones we'll cover in 03-03. We'll close by building real Conv2D and MaxPooling2D layers in Keras, interpreting their summary() line by line, and visualizing what the filters "see" on a product-style image.

Contents

  1. The convolutional layer in detail: filters, stride, and padding
  2. Input and output channels
  3. Output size and parameter counting
  4. Pooling: max and average
  5. Conv2D and MaxPooling2D in Keras: summary() line by line
  6. Visualizing feature maps with matplotlib

The convolutional layer in detail: filters, stride, and padding

A convolutional layer is defined by a handful of design decisions (its hyperparameters):

  • Number of filters: how many distinct pattern detectors the layer has. Each filter produces its own feature map.
  • Filter size (kernel size): the sliding window, typically 3×3 or 5×5. Small windows capture fine local patterns; nowadays 3×3 is the standard (in 03-03 you'll see why VGG made it canonical).
  • Stride: how many pixels the filter moves at each step.
  • Padding: whether or not a border of zeros is added around the image before convolving.

Stride: the step size

In 03-01 we slid the filter one position at a time: stride = 1. With stride = 2, the filter jumps two at a time, evaluating half the positions per dimension, so the output map comes out roughly half as wide and half as tall (a quarter of the positions in total).

Stride 1 over 5 columns:        Stride 2 over 5 columns:
positions 0,1,2  (3 columns)    positions 0,2  (2 columns)
  • Stride 1: maximum analysis resolution; this is the usual choice in convolutional layers.
  • Stride 2: reduces the spatial size (sometimes used instead of pooling to "shrink" the maps).

Padding: valid vs. same

Convolving with no extras makes the map shrink (from 5×5 we went to 3×3) and, worse still, edge pixels take part in fewer windows than central ones: the top-left corner appears in only 1 window, while a central pixel appears in 9. If the product logo sits right against the edge of the photo, it gets analyzed "less".

The solution is padding: surrounding the image with a frame of zeros before convolving.

Mode What it does Output size (stride 1) When to use it
valid No padding: the filter is only placed where it fits entirely Shrinks: n - f + 1 When losing the border doesn't matter or you want to reduce size
same Pads with exactly as many zeros as needed Same as the input: n When you want to stack many layers without the image evaporating

Example with our 5×5 image and 3×3 filter: with same, a 1-pixel frame of zeros is added (the image becomes 7×7) and the output is 5×5 again. Without padding, after 2 layers of 3×3 the image would have gone from 5×5 to 1×1: with small images and deep networks, valid eats up the image in no time.

Input and output channels

So far we've convolved grayscale images (1 channel). A real product photo has 3 channels (RGB), that is, it's a block of height × width × 3. How does the filter adapt?

Golden rule: the filter always has as many channels as its input. A "3×3 filter" over an RGB input is actually a 3×3×3 block of weights (27 weights): it has one 3×3 slice for the red channel, another for the green, and another for the blue. At each position, the 27 products are multiplied and summed (plus the bias) and a single number comes out. That's why:

  • 1 filter → 1 2D feature map, regardless of how many channels the input has.
  • A layer with 32 filters → output with 32 channels (a height × width × 32 block).

And here comes the chaining that makes CNNs deep: the next convolutional layer receives that 32-channel block, so its filters will be 3×3×32. The "channels" stop being colors and become maps of detected patterns: channel 7 might be "vertical edges", channel 21 "metallic texture". Each layer combines the patterns of the previous one — the hierarchy of features from 03-01 in action.

flowchart LR
    A["RGB input<br/>64x64x3"] -->|"32 filters of 3x3x3"| B["Block<br/>64x64x32"]
    B -->|"64 filters of 3x3x32"| C["Block<br/>64x64x64"]

Output size and parameter counting

The output size formula

With input of size n × n, filter f × f, padding p (padding pixels per side), and stride s:

output size = ⌊(n + 2p − f) / s⌋ + 1

(⌊·⌋ = round down.) Common checks and examples:

Input (n) Filter (f) Padding Stride (s) Calculation Output
5 3 valid (p=0) 1 (5+0−3)/1 + 1 3
5 3 same (p=1) 1 (5+2−3)/1 + 1 5
28 3 valid (p=0) 1 (28+0−3)/1 + 1 26
28 3 same (p=1) 1 (28+2−3)/1 + 1 28
64 3 same (p=1) 2 ⌊(64+2−3)/2⌋ + 1 32
224 7 p=3 2 ⌊(224+6−7)/2⌋ + 1 112

The first row is exactly our hand-worked example from 03-01: from 5×5 to 3×3. The last one is the first layer of ResNet, which you'll meet in 03-03: you now know where its "112" comes from.

For same with stride 1, the required padding is p = (f − 1) / 2: with a 3×3 filter, p=1; with 5×5, p=2 (which is why filters are usually odd-sized).

Parameters of a convolutional layer

parameters = (f × f × input_channels + 1) × number_of_filters

The +1 is each filter's bias. Example: a Conv2D with 32 filters of 3×3 over an RGB input:

  • (3 × 3 × 3 + 1) × 32 = 28 × 32 = 896 parameters.

Let's compare with an "equivalent" dense layer connecting the same input to the same output, for a 64×64×3 image and 64×64×32 output (same padding):

Layer Calculation Parameters
Conv2D, 32 filters of 3×3 (3·3·3 + 1) × 32 896
Dense from 64·64·3 = 12,288 inputs to 64·64·32 = 131,072 outputs 12,288 × 131,072 + 131,072 ~1,610,743,808

896 versus 1,610 million: a factor of almost two million. It's the direct consequence of the local connectivity and shared weights from 03-01. Also note that the conv layer's parameters do not depend on the image size, only on the filter and the channels: the same 896-parameter layer processes 64×64 photos or 1024×768 ones.

Pooling: max and average

Pooling is the "summarizing" operation we saw in the diagram in 03-01. It splits each feature map into windows (typically 2×2, with stride 2, no overlap) and replaces each window with a single value:

  • Max pooling: keeps the maximum of the window. Interpretation: "was the pattern detected anywhere in this area?" — it preserves the strongest detection and discards the rest.
  • Average pooling: the mean of the window. It summarizes the overall activation of the area; it's smoother, and its global variant (averaging the entire map) is used at the end of modern architectures as a replacement for Flatten (you'll see it in 03-03).

Hand-worked example of 2×2 max pooling with stride 2 on a 4×4 map:

Feature map:                  Max pooling 2x2:
 1  3 | 2  0
 4  2 | 1  1        ->         4  2
------+------                  6  8
 0  6 | 5  8
 1  2 | 3  0

Each 2×2 block is reduced to its maximum: max(1,3,4,2)=4, max(2,0,1,1)=2, max(0,6,1,2)=6, max(5,8,3,0)=8.

Effects of 2×2 pooling:

  • Halves the width and height → the next layer processes 4 times fewer positions → less compute and memory. Channels are unchanged.
  • Has 0 parameters: there is nothing to learn; it's a fixed operation.
  • Contributes a small invariance to local translations: if the detected edge shifts 1 pixel within the window, the maximum doesn't change. The CNN doesn't care whether the coffee maker's handle sits 2 pixels further left in the seller's photo.
  • By shrinking the map, each pixel in later layers "sees" a larger area of the original image (the receptive field grows), which helps move from detecting edges to detecting parts and objects.

The classic pattern we already previewed in 03-01 is to alternate: conv (extract) → pool (summarize and shrink) → conv → pool... usually doubling the number of filters after each pooling (16 → 32 → 64...): as we lose spatial resolution, we gain richness of patterns.

Conv2D and MaxPooling2D in Keras: summary() line by line

Let's build in Keras the convolutional stage for TecnoMarket product photos resized to 64×64 RGB, with 4 output categories (laptop, coffee maker, monitor, toaster). We're not going to train yet (the full guided project is 07-01); the goal here is to read the model.

from tensorflow import keras
from tensorflow.keras import layers

model = keras.Sequential([
    keras.Input(shape=(64, 64, 3)),                    # 64x64 RGB photo
    layers.Conv2D(32, (3, 3), padding="same", activation="relu"),
    layers.MaxPooling2D((2, 2)),
    layers.Conv2D(64, (3, 3), padding="same", activation="relu"),
    layers.MaxPooling2D((2, 2)),
    layers.Conv2D(128, (3, 3), padding="same", activation="relu"),
    layers.MaxPooling2D((2, 2)),
    layers.Flatten(),
    layers.Dense(64, activation="relu"),
    layers.Dense(4, activation="softmax"),             # 4 product categories
])

model.summary()

Breaking down the arguments of Conv2D(32, (3, 3), padding="same", activation="relu"):

  • 32: number of filters → the output will have 32 channels.
  • (3, 3): filter size.
  • padding="same": the output keeps the input's width and height (the default stride is 1; you would change it with strides=2).
  • activation="relu": the activation after each convolution + bias, as we recommended for hidden layers in 02-02.

MaxPooling2D((2, 2)) uses a 2×2 window and, by default, a stride equal to the window (2): it halves the size.

The summary() produces (formatted):

Layer (type)                 Output Shape        Param #
=========================================================
conv2d (Conv2D)              (None, 64, 64, 32)      896
max_pooling2d (MaxPooling2D) (None, 32, 32, 32)        0
conv2d_1 (Conv2D)            (None, 32, 32, 64)   18,496
max_pooling2d_1 (MaxPooling) (None, 16, 16, 64)        0
conv2d_2 (Conv2D)            (None, 16, 16, 128)  73,856
max_pooling2d_2 (MaxPooling) (None, 8, 8, 128)         0
flatten (Flatten)            (None, 8192)              0
dense (Dense)                (None, 64)          524,352
dense_1 (Dense)              (None, 4)               260
=========================================================
Total params: 617,860

Let's verify each line with our formulas (the None is the batch dimension, as in 02-05):

  1. conv2d: same padding → still 64×64; 32 filters → 32 channels. Parameters: (3·3·3 + 1) × 32 = 896. ✔ (our earlier example)
  2. max_pooling2d: 64/2 = 32 → (32, 32, 32). 0 parameters, like every pooling layer.
  3. conv2d_1: input with 32 channels → 3×3×32 filters. (3·3·32 + 1) × 64 = 289 × 64 = 18,496. ✔
  4. conv2d_2: (3·3·64 + 1) × 128 = 577 × 128 = 73,856. ✔
  5. flatten: 8 × 8 × 128 = 8,192 values. Flattening here is reasonable: these are 8,192 abstract features, not raw pixels.
  6. dense: 8,192 × 64 + 64 = 524,352. Notice that the dense layer after the Flatten concentrates 85% of the model's parameters — the reason modern architectures replace it with global average pooling (03-03).
  7. dense_1: 64 × 4 + 4 = 260, with softmax for the 4 categories (02-02).

Total: ~618,000 parameters to process 64×64×3 images. The dense network from 02-05 needed ~100,000 just for 28×28×1 MNIST; an equivalent dense network for 64×64×3 would balloon into the millions. And of those 618,000, only 93,248 sit in the convolutional part: the feature extractor is dirt cheap.

Visualizing feature maps with matplotlib

To build intuition for what each filter does, let's apply a convolution + ReLU and max pooling to a synthetic "product photo" (a bright box on a dark background, like the packaging of a TecnoMarket product) and plot the resulting maps. We use hand-set filters so the result is interpretable; on a trained model you'd do the same with the learned filters.

import numpy as np
import matplotlib.pyplot as plt

# 1) Synthetic 32x32 image: a bright "product box" on a dark background
image = np.zeros((32, 32))
image[8:24, 10:22] = 1.0           # bright rectangle (the product)

# 2) Two edge-detection filters
vertical_kernel = np.array([[1, 0, -1]] * 3, dtype=float)      # vertical edges
horizontal_kernel = vertical_kernel.T                          # horizontal edges

def convolve2d(image, kernel):
    rows, cols = image.shape
    k_rows, k_cols = kernel.shape
    output = np.zeros((rows - k_rows + 1, cols - k_cols + 1))
    for i in range(output.shape[0]):
        for j in range(output.shape[1]):
            output[i, j] = np.sum(image[i:i + k_rows, j:j + k_cols] * kernel)
    return output

def relu(x):
    return np.maximum(0, x)        # the ReLU from 02-02

def max_pooling(fmap, size=2):
    rows, cols = fmap.shape
    output = np.zeros((rows // size, cols // size))
    for i in range(output.shape[0]):
        for j in range(output.shape[1]):
            output[i, j] = np.max(fmap[i*size:(i+1)*size, j*size:(j+1)*size])
    return output

# 3) Conv -> ReLU -> pooling pipeline for each filter
maps = {
    "Vertical edges": max_pooling(relu(convolve2d(image, vertical_kernel))),
    "Horizontal edges": max_pooling(relu(convolve2d(image, horizontal_kernel))),
}

# 4) Visualization
fig, axes = plt.subplots(1, 3, figsize=(12, 4))
axes[0].imshow(image, cmap="gray")
axes[0].set_title("Original 32x32 image")
for ax, (name, fmap) in zip(axes[1:], maps.items()):
    ax.imshow(fmap, cmap="gray")
    ax.set_title(f"{name}\n{fmap.shape[0]}x{fmap.shape[1]} after pooling")
plt.tight_layout()
plt.show()

What you'll see when you run it:

  • The vertical edges map lights up only along the left side of the box (a dark→bright transition for this filter; the right side yields negative values that ReLU zeroes out — a second filter with opposite signs would capture that side, which is why layers carry many filters).
  • The horizontal edges map lights up along the top side of the box.
  • Both maps are 15×15: the valid convolution leaves 30×30 and the 2×2 pooling halves it.

Each filter produces a "map of where its pattern is", and pooling condenses those maps while preserving the detections. With the filters of a trained model, this same technique (extracting the output of an intermediate layer and plotting it) is a standard tool for inspecting what a CNN has learned.

Common Mistakes and Tips

  • Forgetting the channels when counting parameters. The most frequent mistake: computing (3·3 + 1) × 64 instead of (3·3·32 + 1) × 64. The filter always spans all the channels of its input.
  • Believing that pooling learns something. It has 0 parameters; it's a fixed operation. It reduces size — it doesn't "extract features" on its own.
  • Mixing up the effect of padding and stride. Padding decides whether size is lost at the borders (same prevents it); stride decides how often the filter is evaluated (stride 2 halves the size). They are independent controls.
  • Stacking valid convs on small images. Each 3×3 valid layer shaves off 2 pixels per dimension; ten layers on a 28×28 image leave it at 8×8 before you've even started pooling. To stack in depth, same is your friend.
  • Ignoring where the parameters live. Always look at the summary(): if a dense layer after the Flatten concentrates most of the parameters, there's your prime candidate for overfitting (remember 02-05) and for bloating the model.
  • Tip: for any summary(), reproduce by hand the Output Shape and Param # columns of the first two layers. It's the exercise that consolidates this lesson the fastest.

Exercises

Exercise 1: predicting shapes and parameters

Without running anything, work out the Output Shape and Param # of each layer of this model for 32×32 RGB TecnoMarket photos:

model = keras.Sequential([
    keras.Input(shape=(32, 32, 3)),
    layers.Conv2D(16, (5, 5), padding="valid", activation="relu"),
    layers.MaxPooling2D((2, 2)),
    layers.Conv2D(32, (3, 3), padding="same", activation="relu"),
    layers.MaxPooling2D((2, 2)),
])

Exercise 2: pooling by hand

Apply max pooling and average pooling (2×2 window, stride 2) to this feature map, and explain which of the two would better preserve the detection of a weak but real pattern that activates only one pixel:

2  0  1  1
0  0  1  3
8  1  0  0
2  1  0  0

Exercise 3: designing the reduction

Product photos arrive at 128×128×3. You want a convolutional stage that ends in 16×16 maps using [Conv2D 3×3 same + MaxPooling2D 2×2] blocks. How many blocks do you need? If you start with 16 filters and double them in each block, how many channels will the final output have, and how many parameters will the last conv layer have?

Solutions

Exercise 1:

Layer Output Shape Parameter calculation Param #
Conv2D 16 filters 5×5 valid (None, 28, 28, 16) — (32−5)/1+1 = 28 (5·5·3 + 1) × 16 = 76 × 16 1,216
MaxPooling2D 2×2 (None, 14, 14, 16) — 0
Conv2D 32 filters 3×3 same (None, 14, 14, 32) (3·3·16 + 1) × 32 = 145 × 32 4,640
MaxPooling2D 2×2 (None, 7, 7, 32) — 0

Total: 5,856 parameters.

Exercise 2:

Max pooling:     Average pooling:
2  3              0.5   1.5
8  0              3.0   0.0

(Maxima: max(2,0,0,0)=2, max(1,1,1,3)=3, max(8,1,2,1)=8, max(0,0,0,0)=0. Means: 2/4=0.5, 6/4=1.5, 12/4=3.0, 0/4=0.0.)

Max pooling better preserves a weak but real detection that activates a single pixel: the maximum keeps it intact (the 8 survives as 8), while the mean dilutes it among the inactive neighbors (8 → 3.0). That's why max pooling is the default choice in intermediate stages: what matters is whether the pattern showed up, not how much the area activated on average.

Exercise 3:

Each block halves the size: 128 → 64 → 32 → 16. You need 3 blocks. The filters: 16 → 32 → 64, so the final output is 16×16 with 64 channels. The last conv layer receives 32 channels and has 64 filters of 3×3: (3·3·32 + 1) × 64 = 289 × 64 = 18,496 parameters.

Conclusion

You now command the full mechanics of the two blocks that make up a CNN's extraction stage. For the convolutional layer, you know how to configure its four decisions (filters, kernel size, stride, padding), you know the filters span all the channels of their input and that the layer produces one channel per filter, and you can predict its output with the formula ⌊(n + 2p − f)/s⌋ + 1 and count its parameters with (f·f·channels + 1) × filters — 896 parameters where an equivalent dense layer would need 1,600 million. For pooling, you know it summarizes windows (max preserves detections, average smooths them), that it cuts compute to a quarter per block without adding parameters, and that it grants local invariance for free. And you've translated all of this into Keras, verifying every line of a summary() by hand and visualizing feature maps with matplotlib.

With these pieces you can now read the blueprints of any CNN. And that is precisely what we'll do in the next lesson: tour the popular architectures that made the field's history — LeNet, AlexNet, VGG, Inception, ResNet — understanding what idea each one contributed and confirming, summary() in hand, that you can already decipher them.

© Copyright 2026. All rights reserved