In 03-03 we catalogued the great CNN architectures — VGG, ResNet, MobileNet, EfficientNet — and closed with a promise: you don't need to train them from scratch to use them. This lesson keeps that promise. Transfer learning means taking a model trained on a large, generic problem and reusing its knowledge on your small, specific one. It is, by far, the technique with the best effort-to-result ratio in practical deep learning: it is the reason TecnoMarket can have a professional-grade product-category classifier with a few dozen photos per category, instead of needing millions.
Contents
- Why training from scratch is a luxury
- The intuition: generic features transfer
- Two strategies: feature extraction and fine-tuning
- When to use each strategy
- Full example in Keras: MobileNetV2 for the TecnoMarket catalog
- Fine-tuning: unfreeze with care
- Transfer learning beyond images
Why training from scratch is a luxury
Let's put concrete numbers on the problem. The architectures from 03-03 were trained on ImageNet: ~1.28 million images labeled across 1,000 classes. Training one of those networks from scratch involves:
- Data: gathering and labeling on the order of a million images. Hand-labeling, even at a few cents per image, costs tens of thousands of euros and months of work.
- Compute: the original training of a ResNet-50 on ImageNet took on the order of weeks on a GPU of the day; today, with modern hardware, it is still hours or days of expensive GPU time (hundreds or thousands of euros in the cloud per experiment — and it rarely works on the first experiment).
- Expertise: tuning the hyperparameters of a training run like that is a craft in itself.
Now the reality at TecnoMarket: to classify product photos into, say, 8 categories, the team has about 200 photos per category (~1,600 images). With so little data, a CNN trained from scratch will memorize the training set — the overfitting that already surfaced in 02-05, but in its severe form.
The way out: someone (Google, in MobileNet's case) already paid for the luxury. The resulting weights are published and Keras downloads them with a single argument (weights="imagenet"). The question is how to take advantage of them.
The intuition: generic features transfer
Recall the hierarchy of representations from 03-01: a CNN learns edges → textures → parts → objects, from its first layers to its last. The key observation behind transfer learning is that this hierarchy becomes specific only at the end:
- The early layers detect edges, corners, color gradients, basic textures. These features are universal: they are just as useful for telling cats apart as for telling coffee makers from toasters. A network trained on ImageNet already has them, and they are essentially the same ones it would learn from your data... if you had enough of it.
- The middle layers detect motifs and parts (grilles, handles, metallic surfaces, screens) — still fairly reusable for product photos.
- The final layers combine everything into concepts specific to ImageNet's 1,000 classes ("Siberian husky", "rotary dial telephone"). These are the least transferable: TecnoMarket doesn't sell huskies.
Hence the general recipe: keep the convolutional base (the generic visual knowledge, hugely expensive to obtain) and replace the classification head (the part specific to the original problem) with a new one tailored to your classes.
flowchart LR
subgraph Pretrained on ImageNet
A[Early layers<br/>edges and textures<br/>HIGHLY transferable] --> B[Middle layers<br/>parts and motifs<br/>transferable]
end
B --> C[Original head<br/>1000 ImageNet classes<br/>discarded]
B --> D[NEW head<br/>8 TecnoMarket categories<br/>trained]
style C stroke-dasharray: 5 5
Two strategies
There are two ways to reuse the base, from most conservative to most ambitious:
1. Feature extraction (frozen base). The convolutional base is frozen (trainable = False): its weights are untouched. It acts as a fixed function that turns each image into a feature vector — an embedding, exactly the concept from 03-04 — and only the new head on top is trained. Advantages: extremely fast, impossible to damage the pretrained knowledge, works with very little data (the head has few parameters to fit).
2. Fine-tuning. After training the head, you unfreeze the last layers of the base and keep training everything together with a very low learning rate. The base adapts its high-level features to the new domain (product photos have white backgrounds and frontal framing that ImageNet doesn't emphasize). Advantage: more potential accuracy. Risk: with little data or a careless learning rate, you can destroy the pretrained features (a phenomenon known as catastrophic forgetting) and overfit.
When to use each strategy
The decision hinges on two axes: how much data you have and how similar your domain is to ImageNet.
| Small dataset (hundreds–a few thousand) | Large dataset (tens of thousands+) | |
|---|---|---|
| Domain similar to ImageNet (natural photos: products, animals, scenes) | Feature extraction. The features already work and there isn't data for more | Fine-tuning of the last layers (or more layers): there is enough data to refine without breaking things |
| Different domain (X-rays, satellite imagery, spectrograms) | The hard case: extract features from intermediate layers (the last ones are too "ImageNet") and consider aggressive data augmentation (05-04) | Deep fine-tuning, even of nearly the whole network: the high-level features must be relearned |
TecnoMarket sits in the top-left-to-middle cell: natural product photos (similar domain) with ~1,600 images (small). We will start with feature extraction and apply gentle fine-tuning afterwards.
Full example in Keras: MobileNetV2 for the catalog
We choose MobileNetV2 (from 03-03: the family designed to be light and fast, ideal if the classifier will run on TecnoMarket's modest servers or on the warehouse staff's phones).
Assume the images are organized in folders by category (data/coffee_makers/, data/toasters/, ...), the format image_dataset_from_directory reads directly:
from tensorflow import keras
from tensorflow.keras import layers
IMG_SIZE = (160, 160) # MobileNetV2 accepts several sizes; 160x160 is light
N_CLASSES = 8
# 1. Load the data with a train/validation split
train_ds = keras.utils.image_dataset_from_directory(
"tecnomarket-dl/data/products", validation_split=0.2, subset="training",
seed=42, image_size=IMG_SIZE, batch_size=32)
val_ds = keras.utils.image_dataset_from_directory(
"tecnomarket-dl/data/products", validation_split=0.2, subset="validation",
seed=42, image_size=IMG_SIZE, batch_size=32)
# 2. Pretrained base, WITHOUT its 1000-class head
base = keras.applications.MobileNetV2(
input_shape=IMG_SIZE + (3,),
include_top=False, # discard the ImageNet head
weights="imagenet", # download the pretrained knowledge
)
base.trainable = False # FROZEN: feature extraction
# 3. Full model: preprocessing + base + new head
inputs = keras.Input(shape=IMG_SIZE + (3,))
# Each architecture expects its own specific normalization; ALWAYS use it
x = keras.applications.mobilenet_v2.preprocess_input(inputs) # to [-1, 1]
x = base(x, training=False) # training=False: also pins the BatchNorm layers
x = layers.GlobalAveragePooling2D()(x) # 5x5x1280 map -> vector of 1280
x = layers.Dropout(0.2)(x) # a bit of regularization (05-04)
outputs = layers.Dense(N_CLASSES, activation="softmax")(x)
model = keras.Model(inputs, outputs)
model.compile(optimizer=keras.optimizers.Adam(1e-3),
loss="sparse_categorical_crossentropy",
metrics=["accuracy"])
# 4. Train ONLY the head (the base is frozen)
history = model.fit(train_ds, validation_data=val_ds, epochs=10)Breaking down the decisions:
include_top=False: discards the final ImageNet-specific layers; we keep the convolutional base, which emits a 5×5×1280 feature map.GlobalAveragePooling2D: averages each of the 1,280 maps down to a single number → a vector of 1,280. It is the "parameter-free" version of flatten + dense, and it reduces the head's risk of overfitting.preprocess_input: each architecture was trained with its own normalization (MobileNetV2 expects pixels in [-1, 1]). Skipping it is the most common silent mistake: the network "works" but performs poorly.base(x, training=False): besides freezing weights, it pins the statistics of the base's BatchNormalization layers (you will see what those are in 05-04); it is the correct way to use a frozen base in Keras.- If you run
model.summary()(the habit you picked up in 03-02) you will see something revealing: ~2.26 million non-trainable parameters (the base) and only ~10,000 trainable ones (the head). You are training 0.5% of the network.
With this setup, it is common to jump from 50-60% accuracy (a small CNN from scratch with 1,600 images) to 85-90%+ in a few epochs. That leap is ImageNet's knowledge working for free.
Fine-tuning: unfreeze with care
Second phase, optional but common: unfreeze the top of the base to adapt it to the "product photo" domain.
# 1. Unfreeze ONLY the final stretch of the base
base.trainable = True
for layer in base.layers[:-30]: # all but the last ~30 layers
layer.trainable = False # stay frozen
# 2. RECOMPILE with a VERY low learning rate (100x lower)
model.compile(optimizer=keras.optimizers.Adam(1e-5), # was 1e-3
loss="sparse_categorical_crossentropy",
metrics=["accuracy"])
# 3. A few more epochs, watching the validation curve
history_ft = model.fit(train_ds, validation_data=val_ds, epochs=5)The three golden rules of fine-tuning:
- Head first, base second. If you unfreeze while the head is still random, its large, chaotic gradients — the blame assignment from 02-03 in its most destructive form — wipe out the pretrained weights.
- A very low learning rate (typically 10⁻⁵, two orders of magnitude below the usual). You are retouching valuable knowledge, not learning from scratch: small steps.
- Unfreeze from the top down. The last layers (the most specific) first; the early ones (universal edges and textures) are almost never worth touching.
If after fine-tuning the validation accuracy drops while training accuracy climbs, you have slipped into overfitting: go back to the previous checkpoint and stick with the frozen base, or add regularization (next lesson).
The full guided project — with fine-grained evaluation, a per-category confusion matrix and error analysis — is lesson 07-05; there we will take this same setup all the way to a production-defensible model.
Transfer learning beyond images
The idea is not exclusive to vision:
- Pretrained word embeddings: in 04-03 we trained the
Embeddinglayer from scratch; you can also load vectors already trained on billions of words (Word2Vec, GloVe) — that is transfer learning for the first layer of a text model. - Pretrained language models: the modern and far more powerful version — models trained to "understand the language" in general which are then fine-tuned for sentiment, review classification or search. These are BERT and company, which we will meet in the next lesson (05-05), where this idea of transfer meets the Transformer architecture.
The pattern is always the same: expensive generic knowledge, trained once by whoever can afford it; cheap specific adaptation, within anyone's reach. It is probably the single most important concept for applying deep learning in a real company.
Common Mistakes and Tips
- Forgetting the chosen architecture's
preprocess_input. Each family expects its own range/normalization. With the wrong normalization there is no runtime error, just mediocre, inexplicable metrics. It is the first suspect when a pretrained model performs oddly. - Unfreezing everything and training with the usual learning rate. Typical result: within two epochs the network forgets ImageNet and performs worse than frozen. Frozen base first, lr ~10⁻⁵ afterwards.
- Forgetting to recompile after changing
trainable. In Keras, changingtrainablehas no effect until you callcompile()again. If your curves don't change after "unfreezing", this is almost certainly why. - Evaluating only global accuracy. With imbalanced categories (many phone photos, few air-fryer photos) accuracy deceives; look at per-class performance — we will do it systematically in 07-05.
- Picking a giant base by default. If the classifier must respond in milliseconds in production, EfficientNet-B7 is overkill; MobileNetV2 or EfficientNet-B0 (remember the relative sizes from 03-03) are usually the sweet spot.
- Ignoring the license and provenance of the weights. For commercial use at TecnoMarket, verify that the pretrained weights allow it (those in
keras.applicationsdo, but not everything published on the internet does).
Exercises
Exercise 1. The TecnoMarket team wants to classify package X-ray images from the warehouse security scanner into "normal" / "review", with only 800 labeled images. Using the 2×2 table, reason out which transfer learning strategy you would apply and what special precaution you would take regarding which layers of the base to use.
Exercise 2. A colleague runs this code and complains that fine-tuning "does nothing" (the curves are identical to those of the frozen phase). Find the two mistakes:
base.trainable = True
for layer in base.layers[:-30]:
layer.trainable = False
history_ft = model.fit(train_ds, validation_data=val_ds, epochs=5)
# (the model was compiled before touching trainable, with Adam(1e-3))Exercise 3. Explain why GlobalAveragePooling2D + Dense(8) produces a head with far fewer parameters than Flatten() + Dense(8) over a 5×5×1280 output map, roughly calculating the dense layer's parameters in each case, and why that matters with a small dataset.
Solutions
Solution 1.
This is the small dataset + different domain cell (X-rays don't resemble ImageNet's natural photos): the hard case. Strategy: feature extraction, but with the precaution of not trusting the base's final layers, which encode very "ImageNet" concepts (animals, everyday objects) of little use for X-rays; better to extract features from intermediate layers (textures, shapes, contrasts, which are indeed universal) and train the head on those. Sensible complements: aggressive data augmentation (05-04) and, if more images are collected over time, migrating to deep fine-tuning. Aggressive fine-tuning from the start, with 800 images, would overfit.
Solution 2.
- The model was not recompiled after changing
trainable: in Keras the change takes no effect untilcompile()is called again, so the base is still effectively frozen. That is why it "does nothing". - When recompiling, the learning rate must come down: recompiling with the original
Adam(1e-3)would do the opposite of "nothing" — steps too large that would degrade the pretrained weights. The right move:model.compile(optimizer=keras.optimizers.Adam(1e-5), ...)and thenfit.
Solution 3.
- With
Flatten(): the 5×5×1280 map becomes a vector of 5·5·1280 = 32,000 values; the 8-output dense layer needs 32,000·8 + 8 ≈ 256,000 parameters. - With
GlobalAveragePooling2D: each of the 1,280 maps is averaged to one number → a vector of 1,280; the dense layer needs 1,280·8 + 8 ≈ 10,000 parameters (25 times fewer).
With only ~1,600 images, a 256,000-parameter head has more than enough capacity to memorize the training set (overfitting); the 10,000-parameter one is forced to lean on the base's features, which is exactly what we want. Global pooling also makes the head independent of the spatial size of the incoming map.
Conclusion
Transfer learning turns deep learning's biggest obstacle — the cost of data and compute — into a problem already solved by others: the pretrained convolutional base contributes the generic features (edges, textures, parts: the hierarchy from 03-01) and you only pay for the specific adaptation, choosing between a frozen base (little data) and fine-tuning with a low learning rate (more data, more accuracy). With MobileNetV2 and about 1,600 photos, TecnoMarket's catalog classifier jumped into another league, and in 07-05 we will take it all the way to production-level evaluation.
But that model, like all the previous ones, shares an enemy we have been dodging since the curves of 02-05: overfitting. The next lesson tackles it head-on, with the full arsenal — L1/L2, Dropout, Batch Normalization, Early Stopping and data augmentation — to improve any network in this course.
Deep Learning Course
Module 1: Introduction to Deep Learning
- What is Deep Learning?
- History and evolution of Deep Learning
- Applications of Deep Learning
- Basic concepts of neural networks
- Setting up the work environment
Module 2: Neural Network Fundamentals
- Perceptron and Multilayer Perceptron
- Activation functions
- Forward and backward propagation
- Optimization and loss functions
- Your first complete neural network
Module 3: Convolutional Neural Networks (CNN)
- Introduction to CNNs
- Convolutional and pooling layers
- Popular CNN architectures
- CNN applications in image recognition
Module 4: Recurrent Neural Networks (RNN)
- Introduction to RNNs
- LSTM and GRU
- RNN applications in natural language processing
- Sequences and time series
Module 5: Advanced Deep Learning Techniques
- Generative Adversarial Networks (GAN)
- Autoencoders
- Transfer Learning
- Regularization and improvement techniques
- Attention mechanisms and Transformers
Module 6: Tools and Frameworks
- Introduction to TensorFlow
- Introduction to PyTorch
- Framework comparison
- Development environments and additional resources
- Saving, loading and deploying models
Module 7: Hands-On Projects
- Image classification with CNNs
- Text generation with RNNs
- Anomaly detection with Autoencoders
- Building a GAN for image generation
- Fine-tuning a pretrained model
