Closing module 6 left us with a complete toolbox: we know how to train with Keras, control training with callbacks, build pipelines with tf.data, and package models for production. Time to keep the first promise pending since module 3: building, end to end, TecnoMarket's product image classifier. In this lesson you will develop a full project — data exploration, pipeline, model, training, evaluation, error analysis, iterative improvement and delivery — following the course methodology: prototype on a public dataset (CIFAR-10) and carry it over to the fictional TecnoMarket case.

Contents

  1. Project brief and work plan
  2. Phase 1: exploring the dataset
  3. Phase 2: data pipeline with tf.data and data augmentation
  4. Phase 3: the model — a mini-VGG with BN and dropout
  5. Phase 4: training with callbacks and TensorBoard
  6. Phase 5: evaluation — a per-class confusion matrix
  7. Phase 6: visual error analysis and iterative improvement
  8. Phase 7: delivery — saving and wiring up the review queue

Project brief and work plan

Business context. TecnoMarket sellers upload photos when listing an item and currently pick the category by hand; 12% of new listings end up misclassified and hurt the search engine. We want a model that proposes the category automatically from the photo, with the safeguard we designed in 03-04: if the confidence doesn't clear a threshold, the listing goes to the human review queue.

Stand-in. Since we don't (yet) have TecnoMarket's real dataset, we prototype with CIFAR-10: 60,000 32×32 images across 10 classes (airplane, automobile, bird, cat, deer, dog, frog, horse, ship, truck). Ten classes stand in for ten product categories; the whole pipeline will be transplantable once the real photos arrive.

Phased plan (the same skeleton we will use in every project in this module):

Phase Deliverable
Exploration Know the sizes, classes, balance
Pipeline tf.data with data augmentation built in
Model Our own CNN (the mini-VGG from 03-03 + BN/dropout from 05-04)
Training Callbacks (06-01), curves, TensorBoard
Evaluation Test accuracy + an interpreted confusion matrix
Improvement Iterations guided by diagnosis
Delivery Model+preprocessing saved (06-05) + confidence threshold (03-04)

Phase 1: exploring the dataset

Never train on data you haven't looked at. We work in tecnomarket-dl/notebooks/07_01_classifier.ipynb (the project structure from 06-04):

import tensorflow as tf
import numpy as np
import matplotlib.pyplot as plt

tf.random.set_seed(42)  # reproducibility, as we established in 06-04
np.random.seed(42)

(x_train, y_train), (x_test, y_test) = tf.keras.datasets.cifar10.load_data()

CLASS_NAMES = ["airplane", "automobile", "bird", "cat", "deer",
               "dog", "frog", "horse", "ship", "truck"]

print(x_train.shape, x_test.shape)   # (50000, 32, 32, 3) (10000, 32, 32, 3)
print(np.bincount(y_train.flatten())) # 5000 per class: a balanced dataset

What we verify and why it matters:

  • Shape: 32×32×3 — color images, and very small. Our CNN must be sized accordingly (the output-size formula from 03-02 will tell us how many blocks fit before we run out of pixels).
  • Balance: 5,000 images per class. That legitimizes using accuracy as the main metric; in 07-03 we'll see what happens when there is no balance.
  • Visual inspection: plot a grid of 25 images with their labels. You'll see that at 32×32 even a human struggles to tell cat from dog — calibrate your expectations.
plt.figure(figsize=(8, 8))
for i in range(25):
    plt.subplot(5, 5, i + 1)
    plt.imshow(x_train[i])
    plt.title(CLASS_NAMES[int(y_train[i])], fontsize=8)
    plt.axis("off")
plt.tight_layout()

Phase 2: data pipeline with tf.data and data augmentation

We apply the map → shuffle → batch → prefetch pattern from 06-01, and add the data augmentation with Keras layers we saw in 05-04. We also set aside a validation set separate from the test set:

# Split off validation (last 5000 of train) — the test set stays frozen (06-05)
x_val, y_val = x_train[45000:], y_train[45000:]
x_train, y_train = x_train[:45000], y_train[:45000]

BATCH = 64
AUTOTUNE = tf.data.AUTOTUNE

# Data augmentation as layers (05-04): only active during training
augmentation = tf.keras.Sequential([
    tf.keras.layers.RandomFlip("horizontal"),
    tf.keras.layers.RandomTranslation(0.1, 0.1),
    tf.keras.layers.RandomZoom(0.1),
], name="augmentation")

def preprocess(x, y):
    return tf.cast(x, tf.float32) / 255.0, y   # normalize to [0, 1]

train_ds = (tf.data.Dataset.from_tensor_slices((x_train, y_train))
            .map(preprocess, num_parallel_calls=AUTOTUNE)
            .shuffle(10_000)
            .batch(BATCH)
            .map(lambda x, y: (augmentation(x, training=True), y),
                 num_parallel_calls=AUTOTUNE)
            .prefetch(AUTOTUNE))

val_ds = (tf.data.Dataset.from_tensor_slices((x_val, y_val))
          .map(preprocess).batch(BATCH).prefetch(AUTOTUNE))

test_ds = (tf.data.Dataset.from_tensor_slices((x_test, y_test))
           .map(preprocess).batch(BATCH).prefetch(AUTOTUNE))

Details that make the difference:

  • The augmentation comes after batch and with an explicit training=True: it is applied to whole batches and only in the training pipeline. Validation and test see the original images, as the honest validation of 06-05 demands.
  • Horizontal flip yes, vertical no: an upside-down car is not a realistic catalog photo. Augmentation must generate plausible variations (the criterion from 05-04).
  • Normalization (/255) lives in preprocess; in the delivery phase we will move it inside the model so it travels with it (06-05).

Phase 3: the model — a mini-VGG with BN and dropout

We pick up the mini-VGG we built in 03-03 (Conv-Conv-Pool blocks with doubling filters) and fold in the improvements from 05-04: BatchNormalization after each convolution and increasing Dropout:

from tensorflow.keras import layers, models

def conv_block(x, filters, dropout):
    for _ in range(2):
        x = layers.Conv2D(filters, 3, padding="same", use_bias=False)(x)
        x = layers.BatchNormalization()(x)
        x = layers.Activation("relu")(x)
    x = layers.MaxPooling2D()(x)
    return layers.Dropout(dropout)(x)

inputs = layers.Input(shape=(32, 32, 3))
x = conv_block(inputs, 32, 0.2)       # 32x32 -> 16x16
x = conv_block(x, 64, 0.3)            # 16x16 -> 8x8
x = conv_block(x, 128, 0.4)           # 8x8   -> 4x4
x = layers.Flatten()(x)
x = layers.Dense(128, activation="relu")(x)
x = layers.Dropout(0.5)(x)
outputs = layers.Dense(10, activation="softmax")(x)

model = models.Model(inputs, outputs, name="tecnomarket_cnn_v1")
model.summary()   # ~1.2M parameters

The decisions, explained:

  • Functional API (introduced in 03-03): it lets us factor the block into a function and makes future surgery easier (extracting embeddings as in 03-04, for instance).
  • use_bias=False before BN: BatchNormalization already contributes its own shift; the convolution's bias would be redundant.
  • Increasing dropout (0.2 → 0.5): the deeper layers, with more parameters, need more regularization (05-04).
  • Three pooling blocks take us from 32×32 down to 4×4 — the formula from 03-02 confirms that a fourth block would leave 2×2 feature maps, too impoverished.

Phase 4: training with callbacks and TensorBoard

We compile with Adam and cross-entropy (02-04) and train with the callbacks from 06-01:

model.compile(optimizer=tf.keras.optimizers.Adam(1e-3),
              loss="sparse_categorical_crossentropy",
              metrics=["accuracy"])

callbacks = [
    tf.keras.callbacks.EarlyStopping(monitor="val_accuracy", patience=8,
                                     restore_best_weights=True),
    tf.keras.callbacks.ModelCheckpoint("models/cnn_v1_best.keras",
                                       monitor="val_accuracy",
                                       save_best_only=True),
    tf.keras.callbacks.ReduceLROnPlateau(monitor="val_loss", factor=0.5,
                                         patience=3, min_lr=1e-5),
    tf.keras.callbacks.TensorBoard(log_dir="logs/cnn_v1"),
]

history = model.fit(train_ds, validation_data=val_ds,
                    epochs=60, callbacks=callbacks)

While it trains (on a modest GPU, ~15-25 minutes; on CPU, expect over an hour), open TensorBoard as in 06-04 (tensorboard --logdir logs) and keep an eye on:

  • Train/val curves rising together: thanks to data augmentation and dropout, the gap between training and validation stays small (without them, this same network shoots up to 99% on train and stalls at ~72% on validation — textbook overfitting, 05-04).
  • Steps when the learning rate drops: ReduceLROnPlateau usually produces a visible 1-2 point improvement when it kicks in.

Honest typical result: validation plateaus between 82% and 85% around epoch 40-50. Don't expect more from a homemade CNN trained from scratch on 32×32; that's fine.

Phase 5: evaluation — a per-class confusion matrix

Overall accuracy isn't enough for a business decision: we need to know what the model confuses.

from sklearn.metrics import confusion_matrix, classification_report

test_loss, test_acc = model.evaluate(test_ds)
print(f"Test accuracy: {test_acc:.3f}")   # typical: 0.82 - 0.85

y_prob = model.predict(test_ds)
y_pred = np.argmax(y_prob, axis=1)

print(classification_report(y_test.flatten(), y_pred, target_names=CLASS_NAMES))
cm = confusion_matrix(y_test.flatten(), y_pred)

A typical reading of the matrix (your exact numbers will vary):

Frequent confusion Reading TecnoMarket analogue
cat ↔ dog Small animals, similar textures at 32×32 "earbuds" ↔ "gaming headsets"
airplane ↔ ship Blue backgrounds dominate the image products shot against the same backdrop
automobile ↔ truck Same semantic family "phone" ↔ "tablet"
frog, ship Classes above 90% accuracy visually distinctive categories

The business lesson: errors are not spread uniformly. The confusable pairs are exactly where the confidence threshold from 03-04 will have to work hardest.

Phase 6: visual error analysis and iterative improvement

Before touching anything, look at the failures with your own eyes:

errors = np.where(y_pred != y_test.flatten())[0]
plt.figure(figsize=(10, 6))
for i, idx in enumerate(errors[:15]):
    plt.subplot(3, 5, i + 1)
    plt.imshow(x_test[idx])
    plt.title(f"actual: {CLASS_NAMES[int(y_test[idx])]}\npred: {CLASS_NAMES[y_pred[idx]]}"
              f" ({y_prob[idx].max():.2f})", fontsize=7)
    plt.axis("off")
plt.tight_layout()

You'll see three families of errors: genuinely ambiguous images (you wouldn't get them right either), low-confidence errors (the threshold will catch them) and high-confidence errors (the dangerous ones). With that diagnosis, iterate in this order — from cheapest to most expensive:

  1. More epochs / patience: if the curves were still climbing when training stopped, this is free improvement.
  2. Tune the data augmentation: if there's a train-val gap, intensify it; if train never reaches a decent level, ease it off (excessive augmentation hurts too, 05-04).
  3. Tune the regularization: move dropout ±0.1 depending on the gap.
  4. Capacity: a fourth convolutional block or more filters — only if train has fallen short.
  5. Change the architecture: the door we'll open in 07-05.

The module's golden rule: one change per iteration, with a fixed seed (06-04), or you won't know what caused what.

Phase 7: delivery — saving and wiring up the review queue

Following 06-05, we package model + preprocessing as one unit, so the API can never drift out of sync with training:

# Inference model: normalization inside + probabilities outside
raw_inputs = layers.Input(shape=(32, 32, 3), dtype=tf.uint8)
x = layers.Rescaling(1.0 / 255)(tf.cast(raw_inputs, tf.float32))
prob = model(x, training=False)
serving_model = models.Model(raw_inputs, prob)
serving_model.save("models/category_classifier_v1.keras")

And we apply the 03-04 policy with the confidence threshold:

THRESHOLD = 0.80  # calibrated on validation, never on test

def classify_listing(image_uint8):
    p = serving_model.predict(image_uint8[np.newaxis, ...], verbose=0)[0]
    if p.max() >= THRESHOLD:
        return {"category": CLASS_NAMES[int(p.argmax())], "auto": True}
    return {"category": None, "auto": False, "destination": "review_queue",
            "suggestions": [CLASS_NAMES[i] for i in p.argsort()[-3:][::-1]]}

With the threshold at 0.80, a typical outcome is that the model automatically resolves around 70% of new listings with an accuracy on that subset near 92-94%, and routes the rest to human review with the three most likely categories as suggestions — humans and model collaborating, just as we designed in 03-04. The service would be exposed with the same FastAPI template from 06-05.

Common Mistakes and Tips

  • Applying data augmentation to validation/test as well: it artificially inflates the difficulty and leads you to bad decisions. Augmentation lives only in train_ds.
  • Picking the confidence threshold by looking at the test set: the test set is the frozen set (06-05); if you use it for calibration, it no longer measures anything. Threshold and tuning, always on validation.
  • Changing three things at once between iterations: you won't know what worked. One change, one fixed seed, one comparison.
  • Getting frustrated at not passing 85%: that's the reasonable ceiling for this approach on this dataset. The big improvement won't come from insisting, but from changing strategy (07-05).
  • Forgetting restore_best_weights=True: without it, EarlyStopping leaves you with the last epoch's weights, not the best ones.

Exercises

  1. Add a fourth convolutional block with 256 filters (adjusting the dropout) and compare accuracy and training time against the three-block version. Is it worth it?
  2. Produce the list of the 20 wrong predictions with the highest confidence. Which class pairs show up? Propose a data augmentation change targeted at them.
  3. For thresholds 0.6, 0.7, 0.8 and 0.9 on the validation set, compute what percentage of images gets automated and what accuracy that subset achieves. Present the table and choose a threshold, justifying the trade-off.

Solutions

  1. Replace the last call with x = conv_block(x, 256, 0.4) after the 128-filter block (the feature maps go from 4×4 to 2×2, right at the limit). Typical result: +0.5-1 accuracy point in exchange for ~1.5× time per epoch. Worth it only if that point matters; document both figures.
  2. idx = errors[np.argsort(-y_prob[errors].max(axis=1))[:20]] and plot as in phase 6. Cat↔dog and automobile↔truck usually dominate. Targeted change: RandomContrast(0.2) or random crops (RandomCrop with padding) that force the network to attend to shape rather than background.
  3. For each threshold u: mask = p_val.max(axis=1) >= u; coverage = mask.mean(); automated accuracy = (pred_val[mask] == y_val_flat[mask]).mean(). Typical table: 0.6 → 85%/89%; 0.7 → 78%/91%; 0.8 → 70%/93%; 0.9 → 55%/96%. The choice depends on the cost of an error versus the cost of a review: for catalog listings, 0.8 is usually a good compromise.

Conclusion

You have delivered TecnoMarket's first complete project: a category classifier with a tf.data pipeline, data augmentation, a regularized mini-VGG, training controlled by callbacks, evaluation with a confusion matrix, error analysis and a package ready to serve behind a confidence threshold with a review queue. The result — 82-85% on test — is honest for a CNN trained from scratch, and in 07-05 we will beat it with transfer learning while barely writing more code. First, though, we switch modalities: in the next lesson we keep the promise from 04-04 and build a character-by-character text generator for TecnoMarket's product description drafts.

© Copyright 2026. All rights reserved