Closing module 6 left us with a complete toolbox: we know how to train with Keras, control training with callbacks, build pipelines with tf.data, and package models for production. Time to keep the first promise pending since module 3: building, end to end, TecnoMarket's product image classifier. In this lesson you will develop a full project — data exploration, pipeline, model, training, evaluation, error analysis, iterative improvement and delivery — following the course methodology: prototype on a public dataset (CIFAR-10) and carry it over to the fictional TecnoMarket case.
Contents
- Project brief and work plan
- Phase 1: exploring the dataset
- Phase 2: data pipeline with
tf.dataand data augmentation - Phase 3: the model — a mini-VGG with BN and dropout
- Phase 4: training with callbacks and TensorBoard
- Phase 5: evaluation — a per-class confusion matrix
- Phase 6: visual error analysis and iterative improvement
- Phase 7: delivery — saving and wiring up the review queue
Project brief and work plan
Business context. TecnoMarket sellers upload photos when listing an item and currently pick the category by hand; 12% of new listings end up misclassified and hurt the search engine. We want a model that proposes the category automatically from the photo, with the safeguard we designed in 03-04: if the confidence doesn't clear a threshold, the listing goes to the human review queue.
Stand-in. Since we don't (yet) have TecnoMarket's real dataset, we prototype with CIFAR-10: 60,000 32×32 images across 10 classes (airplane, automobile, bird, cat, deer, dog, frog, horse, ship, truck). Ten classes stand in for ten product categories; the whole pipeline will be transplantable once the real photos arrive.
Phased plan (the same skeleton we will use in every project in this module):
| Phase | Deliverable |
|---|---|
| Exploration | Know the sizes, classes, balance |
| Pipeline | tf.data with data augmentation built in |
| Model | Our own CNN (the mini-VGG from 03-03 + BN/dropout from 05-04) |
| Training | Callbacks (06-01), curves, TensorBoard |
| Evaluation | Test accuracy + an interpreted confusion matrix |
| Improvement | Iterations guided by diagnosis |
| Delivery | Model+preprocessing saved (06-05) + confidence threshold (03-04) |
Phase 1: exploring the dataset
Never train on data you haven't looked at. We work in tecnomarket-dl/notebooks/07_01_classifier.ipynb (the project structure from 06-04):
import tensorflow as tf
import numpy as np
import matplotlib.pyplot as plt
tf.random.set_seed(42) # reproducibility, as we established in 06-04
np.random.seed(42)
(x_train, y_train), (x_test, y_test) = tf.keras.datasets.cifar10.load_data()
CLASS_NAMES = ["airplane", "automobile", "bird", "cat", "deer",
"dog", "frog", "horse", "ship", "truck"]
print(x_train.shape, x_test.shape) # (50000, 32, 32, 3) (10000, 32, 32, 3)
print(np.bincount(y_train.flatten())) # 5000 per class: a balanced datasetWhat we verify and why it matters:
- Shape: 32×32×3 — color images, and very small. Our CNN must be sized accordingly (the output-size formula from 03-02 will tell us how many blocks fit before we run out of pixels).
- Balance: 5,000 images per class. That legitimizes using accuracy as the main metric; in 07-03 we'll see what happens when there is no balance.
- Visual inspection: plot a grid of 25 images with their labels. You'll see that at 32×32 even a human struggles to tell cat from dog — calibrate your expectations.
plt.figure(figsize=(8, 8))
for i in range(25):
plt.subplot(5, 5, i + 1)
plt.imshow(x_train[i])
plt.title(CLASS_NAMES[int(y_train[i])], fontsize=8)
plt.axis("off")
plt.tight_layout()Phase 2: data pipeline with tf.data and data augmentation
We apply the map → shuffle → batch → prefetch pattern from 06-01, and add the data augmentation with Keras layers we saw in 05-04. We also set aside a validation set separate from the test set:
# Split off validation (last 5000 of train) — the test set stays frozen (06-05)
x_val, y_val = x_train[45000:], y_train[45000:]
x_train, y_train = x_train[:45000], y_train[:45000]
BATCH = 64
AUTOTUNE = tf.data.AUTOTUNE
# Data augmentation as layers (05-04): only active during training
augmentation = tf.keras.Sequential([
tf.keras.layers.RandomFlip("horizontal"),
tf.keras.layers.RandomTranslation(0.1, 0.1),
tf.keras.layers.RandomZoom(0.1),
], name="augmentation")
def preprocess(x, y):
return tf.cast(x, tf.float32) / 255.0, y # normalize to [0, 1]
train_ds = (tf.data.Dataset.from_tensor_slices((x_train, y_train))
.map(preprocess, num_parallel_calls=AUTOTUNE)
.shuffle(10_000)
.batch(BATCH)
.map(lambda x, y: (augmentation(x, training=True), y),
num_parallel_calls=AUTOTUNE)
.prefetch(AUTOTUNE))
val_ds = (tf.data.Dataset.from_tensor_slices((x_val, y_val))
.map(preprocess).batch(BATCH).prefetch(AUTOTUNE))
test_ds = (tf.data.Dataset.from_tensor_slices((x_test, y_test))
.map(preprocess).batch(BATCH).prefetch(AUTOTUNE))Details that make the difference:
- The augmentation comes after
batchand with an explicittraining=True: it is applied to whole batches and only in the training pipeline. Validation and test see the original images, as the honest validation of 06-05 demands. - Horizontal flip yes, vertical no: an upside-down car is not a realistic catalog photo. Augmentation must generate plausible variations (the criterion from 05-04).
- Normalization (
/255) lives inpreprocess; in the delivery phase we will move it inside the model so it travels with it (06-05).
Phase 3: the model — a mini-VGG with BN and dropout
We pick up the mini-VGG we built in 03-03 (Conv-Conv-Pool blocks with doubling filters) and fold in the improvements from 05-04: BatchNormalization after each convolution and increasing Dropout:
from tensorflow.keras import layers, models
def conv_block(x, filters, dropout):
for _ in range(2):
x = layers.Conv2D(filters, 3, padding="same", use_bias=False)(x)
x = layers.BatchNormalization()(x)
x = layers.Activation("relu")(x)
x = layers.MaxPooling2D()(x)
return layers.Dropout(dropout)(x)
inputs = layers.Input(shape=(32, 32, 3))
x = conv_block(inputs, 32, 0.2) # 32x32 -> 16x16
x = conv_block(x, 64, 0.3) # 16x16 -> 8x8
x = conv_block(x, 128, 0.4) # 8x8 -> 4x4
x = layers.Flatten()(x)
x = layers.Dense(128, activation="relu")(x)
x = layers.Dropout(0.5)(x)
outputs = layers.Dense(10, activation="softmax")(x)
model = models.Model(inputs, outputs, name="tecnomarket_cnn_v1")
model.summary() # ~1.2M parametersThe decisions, explained:
- Functional API (introduced in 03-03): it lets us factor the block into a function and makes future surgery easier (extracting embeddings as in 03-04, for instance).
use_bias=Falsebefore BN: BatchNormalization already contributes its own shift; the convolution's bias would be redundant.- Increasing dropout (0.2 → 0.5): the deeper layers, with more parameters, need more regularization (05-04).
- Three pooling blocks take us from 32×32 down to 4×4 — the formula from 03-02 confirms that a fourth block would leave 2×2 feature maps, too impoverished.
Phase 4: training with callbacks and TensorBoard
We compile with Adam and cross-entropy (02-04) and train with the callbacks from 06-01:
model.compile(optimizer=tf.keras.optimizers.Adam(1e-3),
loss="sparse_categorical_crossentropy",
metrics=["accuracy"])
callbacks = [
tf.keras.callbacks.EarlyStopping(monitor="val_accuracy", patience=8,
restore_best_weights=True),
tf.keras.callbacks.ModelCheckpoint("models/cnn_v1_best.keras",
monitor="val_accuracy",
save_best_only=True),
tf.keras.callbacks.ReduceLROnPlateau(monitor="val_loss", factor=0.5,
patience=3, min_lr=1e-5),
tf.keras.callbacks.TensorBoard(log_dir="logs/cnn_v1"),
]
history = model.fit(train_ds, validation_data=val_ds,
epochs=60, callbacks=callbacks)While it trains (on a modest GPU, ~15-25 minutes; on CPU, expect over an hour), open TensorBoard as in 06-04 (tensorboard --logdir logs) and keep an eye on:
- Train/val curves rising together: thanks to data augmentation and dropout, the gap between training and validation stays small (without them, this same network shoots up to 99% on train and stalls at ~72% on validation — textbook overfitting, 05-04).
- Steps when the learning rate drops:
ReduceLROnPlateauusually produces a visible 1-2 point improvement when it kicks in.
Honest typical result: validation plateaus between 82% and 85% around epoch 40-50. Don't expect more from a homemade CNN trained from scratch on 32×32; that's fine.
Phase 5: evaluation — a per-class confusion matrix
Overall accuracy isn't enough for a business decision: we need to know what the model confuses.
from sklearn.metrics import confusion_matrix, classification_report
test_loss, test_acc = model.evaluate(test_ds)
print(f"Test accuracy: {test_acc:.3f}") # typical: 0.82 - 0.85
y_prob = model.predict(test_ds)
y_pred = np.argmax(y_prob, axis=1)
print(classification_report(y_test.flatten(), y_pred, target_names=CLASS_NAMES))
cm = confusion_matrix(y_test.flatten(), y_pred)A typical reading of the matrix (your exact numbers will vary):
| Frequent confusion | Reading | TecnoMarket analogue |
|---|---|---|
| cat ↔ dog | Small animals, similar textures at 32×32 | "earbuds" ↔ "gaming headsets" |
| airplane ↔ ship | Blue backgrounds dominate the image | products shot against the same backdrop |
| automobile ↔ truck | Same semantic family | "phone" ↔ "tablet" |
| frog, ship | Classes above 90% accuracy | visually distinctive categories |
The business lesson: errors are not spread uniformly. The confusable pairs are exactly where the confidence threshold from 03-04 will have to work hardest.
Phase 6: visual error analysis and iterative improvement
Before touching anything, look at the failures with your own eyes:
errors = np.where(y_pred != y_test.flatten())[0]
plt.figure(figsize=(10, 6))
for i, idx in enumerate(errors[:15]):
plt.subplot(3, 5, i + 1)
plt.imshow(x_test[idx])
plt.title(f"actual: {CLASS_NAMES[int(y_test[idx])]}\npred: {CLASS_NAMES[y_pred[idx]]}"
f" ({y_prob[idx].max():.2f})", fontsize=7)
plt.axis("off")
plt.tight_layout()You'll see three families of errors: genuinely ambiguous images (you wouldn't get them right either), low-confidence errors (the threshold will catch them) and high-confidence errors (the dangerous ones). With that diagnosis, iterate in this order — from cheapest to most expensive:
- More epochs / patience: if the curves were still climbing when training stopped, this is free improvement.
- Tune the data augmentation: if there's a train-val gap, intensify it; if train never reaches a decent level, ease it off (excessive augmentation hurts too, 05-04).
- Tune the regularization: move dropout ±0.1 depending on the gap.
- Capacity: a fourth convolutional block or more filters — only if train has fallen short.
- Change the architecture: the door we'll open in 07-05.
The module's golden rule: one change per iteration, with a fixed seed (06-04), or you won't know what caused what.
Phase 7: delivery — saving and wiring up the review queue
Following 06-05, we package model + preprocessing as one unit, so the API can never drift out of sync with training:
# Inference model: normalization inside + probabilities outside
raw_inputs = layers.Input(shape=(32, 32, 3), dtype=tf.uint8)
x = layers.Rescaling(1.0 / 255)(tf.cast(raw_inputs, tf.float32))
prob = model(x, training=False)
serving_model = models.Model(raw_inputs, prob)
serving_model.save("models/category_classifier_v1.keras")And we apply the 03-04 policy with the confidence threshold:
THRESHOLD = 0.80 # calibrated on validation, never on test
def classify_listing(image_uint8):
p = serving_model.predict(image_uint8[np.newaxis, ...], verbose=0)[0]
if p.max() >= THRESHOLD:
return {"category": CLASS_NAMES[int(p.argmax())], "auto": True}
return {"category": None, "auto": False, "destination": "review_queue",
"suggestions": [CLASS_NAMES[i] for i in p.argsort()[-3:][::-1]]}With the threshold at 0.80, a typical outcome is that the model automatically resolves around 70% of new listings with an accuracy on that subset near 92-94%, and routes the rest to human review with the three most likely categories as suggestions — humans and model collaborating, just as we designed in 03-04. The service would be exposed with the same FastAPI template from 06-05.
Common Mistakes and Tips
- Applying data augmentation to validation/test as well: it artificially inflates the difficulty and leads you to bad decisions. Augmentation lives only in
train_ds. - Picking the confidence threshold by looking at the test set: the test set is the frozen set (06-05); if you use it for calibration, it no longer measures anything. Threshold and tuning, always on validation.
- Changing three things at once between iterations: you won't know what worked. One change, one fixed seed, one comparison.
- Getting frustrated at not passing 85%: that's the reasonable ceiling for this approach on this dataset. The big improvement won't come from insisting, but from changing strategy (07-05).
- Forgetting
restore_best_weights=True: without it, EarlyStopping leaves you with the last epoch's weights, not the best ones.
Exercises
- Add a fourth convolutional block with 256 filters (adjusting the dropout) and compare accuracy and training time against the three-block version. Is it worth it?
- Produce the list of the 20 wrong predictions with the highest confidence. Which class pairs show up? Propose a data augmentation change targeted at them.
- For thresholds 0.6, 0.7, 0.8 and 0.9 on the validation set, compute what percentage of images gets automated and what accuracy that subset achieves. Present the table and choose a threshold, justifying the trade-off.
Solutions
- Replace the last call with
x = conv_block(x, 256, 0.4)after the 128-filter block (the feature maps go from 4×4 to 2×2, right at the limit). Typical result: +0.5-1 accuracy point in exchange for ~1.5× time per epoch. Worth it only if that point matters; document both figures. idx = errors[np.argsort(-y_prob[errors].max(axis=1))[:20]]and plot as in phase 6. Cat↔dog and automobile↔truck usually dominate. Targeted change:RandomContrast(0.2)or random crops (RandomCropwith padding) that force the network to attend to shape rather than background.- For each threshold
u:mask = p_val.max(axis=1) >= u; coverage =mask.mean(); automated accuracy =(pred_val[mask] == y_val_flat[mask]).mean(). Typical table: 0.6 → 85%/89%; 0.7 → 78%/91%; 0.8 → 70%/93%; 0.9 → 55%/96%. The choice depends on the cost of an error versus the cost of a review: for catalog listings, 0.8 is usually a good compromise.
Conclusion
You have delivered TecnoMarket's first complete project: a category classifier with a tf.data pipeline, data augmentation, a regularized mini-VGG, training controlled by callbacks, evaluation with a confusion matrix, error analysis and a package ready to serve behind a confidence threshold with a review queue. The result — 82-85% on test — is honest for a CNN trained from scratch, and in 07-05 we will beat it with transfer learning while barely writing more code. First, though, we switch modalities: in the next lesson we keep the promise from 04-04 and build a character-by-character text generator for TecnoMarket's product description drafts.
Deep Learning Course
Module 1: Introduction to Deep Learning
- What is Deep Learning?
- History and evolution of Deep Learning
- Applications of Deep Learning
- Basic concepts of neural networks
- Setting up the work environment
Module 2: Neural Network Fundamentals
- Perceptron and Multilayer Perceptron
- Activation functions
- Forward and backward propagation
- Optimization and loss functions
- Your first complete neural network
Module 3: Convolutional Neural Networks (CNN)
- Introduction to CNNs
- Convolutional and pooling layers
- Popular CNN architectures
- CNN applications in image recognition
Module 4: Recurrent Neural Networks (RNN)
- Introduction to RNNs
- LSTM and GRU
- RNN applications in natural language processing
- Sequences and time series
Module 5: Advanced Deep Learning Techniques
- Generative Adversarial Networks (GAN)
- Autoencoders
- Transfer Learning
- Regularization and improvement techniques
- Attention mechanisms and Transformers
Module 6: Tools and Frameworks
- Introduction to TensorFlow
- Introduction to PyTorch
- Framework comparison
- Development environments and additional resources
- Saving, loading and deploying models
Module 7: Hands-On Projects
- Image classification with CNNs
- Text generation with RNNs
- Anomaly detection with Autoencoders
- Building a GAN for image generation
- Fine-tuning a pretrained model
