As we closed the previous module, we said that you already know the techniques — CNNs, RNNs, GANs, autoencoders, transfer learning, attention — and that the time had come to consolidate the tools. We start with the one you have been using since lesson 02-05 without ever looking at it head-on: TensorFlow. Every time you wrote keras.Sequential, model.fit() or model.save(), TensorFlow was underneath doing the heavy lifting: representing data as tensors, computing gradients automatically, and running the operations on CPU or GPU. In this lesson we open the box: you will understand what TensorFlow really is, how its tensors work, how it computes gradients with GradientTape (we will reproduce the manual gradient descent from 02-03), the three ways to build models, and how to feed your networks with efficient data pipelines. For the TecnoMarket team, this is the difference between "following recipes" and "mastering the tool".

Contents

  1. What TensorFlow is and its ecosystem
  2. Tensors: constants, variables, shapes and dtypes
  3. Broadcasting: operating on different shapes
  4. Automatic differentiation with tf.GradientTape
  5. The three APIs for building models
  6. Efficient data pipelines with tf.data
  7. Useful callbacks during training

What TensorFlow is and its ecosystem

TensorFlow is an open-source machine learning platform created by Google and released in 2015. Its name says it all: it makes tensors (multidimensional arrays) flow through graphs of mathematical operations.

It helps to distinguish the layers of the ecosystem:

Layer What it is Example usage
TensorFlow (core) Engine for tensors, operations and automatic differentiation tf.matmul, tf.GradientTape
Keras High-level API for defining and training models keras.Sequential, model.fit()
tf.data Building data pipelines Dataset.from_tensor_slices(...)
TensorBoard Training visualization Loss curves, histograms
TFLite / TF.js / TF Serving Deployment to mobile, browser and servers Covered in 06-05

The key point for you: Keras IS TensorFlow's official high-level API. You have not been using "something else" throughout the course; you have been using TensorFlow through its most convenient interface. The relationship is like an automatic car (Keras) and its engine (TensorFlow): until now you were driving; today we open the hood.

import tensorflow as tf
from tensorflow import keras

print(tf.__version__)          # e.g. 2.16.x
print(keras.__name__)          # keras lives inside the TF ecosystem
  • If you ran pip install tensorflow in 01-05, you already have everything this lesson needs.
  • TensorFlow detects the GPU automatically if one is available (in Colab: Runtime → Change runtime type → GPU).
# Which devices can TensorFlow see?
print(tf.config.list_physical_devices())
# [PhysicalDevice(name='/physical_device:CPU:0', ...), maybe a GPU]

Tensors: constants, variables, shapes and dtypes

A tensor is a multidimensional array with a uniform data type (dtype). You already know them conceptually from the course: a TecnoMarket product image is a (height, width, 3) tensor, a batch of vectorized reviews is a (batch, length) tensor.

tf.constant: immutable tensors

import tensorflow as tf

scalar = tf.constant(4.99)                      # price of a USB cable
vector = tf.constant([120.0, 15.5, 899.0])      # prices of 3 products
matrix = tf.constant([[1, 2], [3, 4]])          # 2D tensor

print(scalar.shape)   # ()        -> 0 dimensions
print(vector.shape)   # (3,)      -> 1 dimension with 3 elements
print(matrix.shape)   # (2, 2)
print(matrix.dtype)   # <dtype: 'int32'>
print(vector.dtype)   # <dtype: 'float32'>

Points you should internalize:

  • shape: the tensor's shape. Shape errors are by far the most frequent in deep learning; learning to read a message like Incompatible shapes: (32, 10) vs (32, 1) properly will save you hours.
  • dtype: the data type. By default floats are float32 (a balance between precision and speed) and integers are int32. You cannot mix dtypes in one operation without converting first with tf.cast.
  • A tf.constant is immutable: you cannot change its values once created.

tf.Variable: tensors that learn

The weights of a network are variables: tensors whose value must be updatable at every training step. When in 02-03 we talked about the "blame assignment" that adjusts each weight, those weights are internally tf.Variable objects.

w = tf.Variable(3.0)           # a weight initialized to 3.0
b = tf.Variable(0.0)           # a bias

w.assign(2.5)                  # change its value (a constant can't do this!)
w.assign_add(0.1)              # w = w + 0.1
print(w.numpy())               # 2.6  -> .numpy() converts to numpy
tf.constant tf.Variable
Mutable No Yes (assign, assign_add)
Typical use Input data, hyperparameters Model weights and biases
Tracked by GradientTape Only if explicitly requested Yes, automatically

Basic operations

a = tf.constant([[1.0, 2.0], [3.0, 4.0]])
b = tf.constant([[10.0, 20.0], [30.0, 40.0]])

print(a + b)              # element-wise addition
print(a * b)              # element-wise product (NOT matrix multiplication!)
print(tf.matmul(a, b))    # matrix product (the one dense layers use)
print(tf.reduce_mean(a))  # mean of all elements -> 2.5
print(tf.reshape(a, (4, 1)))  # change the shape without changing the data

Remember from 02-01: a dense layer is exactly tf.matmul(inputs, weights) + bias followed by an activation. Now you can write it by hand.

Broadcasting: operating on different shapes

Broadcasting lets you operate on tensors of different shapes: TensorFlow virtually "stretches" the small tensor so it fits the large one, without copying data. It is the same rule numpy uses.

prices = tf.constant([[120.0], [15.5], [899.0]])    # shape (3, 1): 3 products
vat = tf.constant(1.21)                             # scalar, shape ()

with_vat = prices * vat          # the scalar applies to all 3 -> shape (3, 1)

discounts = tf.constant([0.95, 0.90, 0.80])         # shape (3,): 3 campaigns
table = prices * discounts       # (3,1) x (3,) -> broadcasting -> (3, 3)
print(table.shape)               # (3, 3): every product with every discount

Practical rules (shapes are compared right to left):

  1. Two dimensions are compatible if they are equal or if one of them is 1.
  2. Missing dimensions are treated as 1.
  3. If no rule applies, you get an Incompatible shapes error.

Broadcasting is extremely powerful, but also a source of silent bugs: in the example above, maybe you wanted 3 prices with 3 discounts (result (3,)) and got a (3, 3) table without any error. Always check the shape of the result.

Automatic differentiation with tf.GradientTape

Here is the crown jewel. In 02-03 we computed gradients by hand with numpy, applying the chain rule step by step to distribute the blame for the error. TensorFlow does exactly that, but automatically: tf.GradientTape is a "recording tape" that logs every operation you perform inside its context, then plays it backwards to compute derivatives.

Let's reproduce the classic example: minimizing a simple loss function, first the gradient and then the full descent.

import tensorflow as tf

# We want to find the w that minimizes the "loss" L(w) = (w - 4)^2
# From calculus we know the minimum is at w = 4 and that dL/dw = 2*(w - 4)

w = tf.Variable(0.0)   # we start far from the minimum

with tf.GradientTape() as tape:      # 1. the tape starts recording
    loss = (w - 4.0) ** 2            # 2. recorded operations

grad = tape.gradient(loss, w)        # 3. play it backwards
print(grad.numpy())                  # -8.0  (= 2*(0-4), the exact derivative)

Line-by-line explanation:

  • with tf.GradientTape() as tape: opens the recording context. Everything that happens inside and involves variables gets logged.
  • loss = (w - 4.0) ** 2 does not just compute the value (16.0): the tape notes "subtract, then square", which is what it needs to apply the chain rule.
  • tape.gradient(loss, w) asks: "how much does loss change if I nudge w a little?". It is the same question backpropagation asked in 02-03.

Now the complete gradient descent, the same loop we wrote with numpy:

w = tf.Variable(0.0)
learning_rate = 0.1

for step in range(30):
    with tf.GradientTape() as tape:
        loss = (w - 4.0) ** 2
    grad = tape.gradient(loss, w)
    w.assign_sub(learning_rate * grad)   # w = w - lr * gradient

print(w.numpy())   # ~3.995: it has converged to 4, just like in 02-03

Comparison with the numpy version from 02-03:

numpy (02-03) TensorFlow (GradientTape)
Derivative You write it by hand: 2*(w-4) The tape computes it automatically
Risk of error High (one wrong derivative and everything breaks) Practically zero
Deep networks Unfeasible by hand (thousands of derivatives) Same effort: one line
Educational value You understand the mechanism You scale the mechanism

When you call model.fit() in Keras, this is exactly what happens internally: a GradientTape records the forward pass, computes the gradients of the loss with respect to all the model's variables, and the optimizer (SGD, Adam...) updates them. fit() is this loop, packaged.

The three APIs for building models

TensorFlow/Keras offers three ways to define a model. You already master two of them:

  1. Sequential (you have known it since 02-05)

For linear stacks of layers: input → layer → layer → output.

model = keras.Sequential([
    keras.layers.Dense(128, activation="relu"),
    keras.layers.Dense(64, activation="relu"),
    keras.layers.Dense(10, activation="softmax"),
])

  1. Functional API (you used it in 03-03)

For graphs with branches, multiple inputs/outputs or skip connections, like the residual block we built:

inputs = keras.Input(shape=(784,))
x = keras.layers.Dense(128, activation="relu")(inputs)
x = keras.layers.Dense(64, activation="relu")(x)
outputs = keras.layers.Dense(10, activation="softmax")(x)
model = keras.Model(inputs, outputs)

  1. Subclassing (the new one)

You define a class inheriting from keras.Model and write the forward pass yourself in call(). Maximum flexibility: you can use loops, conditionals, arbitrary logic.

class TecnoMarketClassifier(keras.Model):
    def __init__(self):
        super().__init__()
        self.dense1 = keras.layers.Dense(128, activation="relu")
        self.dense2 = keras.layers.Dense(64, activation="relu")
        self.output_layer = keras.layers.Dense(10, activation="softmax")

    def call(self, x):
        x = self.dense1(x)      # you decide the data flow
        x = self.dense2(x)
        return self.output_layer(x)

model = TecnoMarketClassifier()
  • In __init__ you declare the layers (the parts).
  • In call you connect the parts (the flow). It runs on every forward pass.
API When to use it Flexibility Already seen in
Sequential Simple linear stack Low 02-05
Functional Graphs with branches and skips Medium-high 03-03
Subclassing Arbitrary logic, research Maximum This lesson

Advice: always use the simplest API that solves your problem. Subclassing will feel familiar when you meet PyTorch in the next lesson: there it is the standard way of working.

Efficient data pipelines with tf.data

Until now we passed numpy arrays directly to fit(). That works with MNIST, but TecnoMarket's real catalog has hundreds of thousands of photos and reviews that do not fit in memory. tf.data builds pipelines: pipes that load, transform and serve the data in batches, in parallel with training.

import tensorflow as tf

# Simulate already-vectorized TecnoMarket reviews (as in 04-03) and their labels
import numpy as np
reviews = np.random.rand(10000, 200).astype("float32")   # 10k reviews
labels = np.random.randint(0, 2, size=(10000,))          # 0=negative, 1=positive

dataset = tf.data.Dataset.from_tensor_slices((reviews, labels))

dataset = (dataset
    .shuffle(buffer_size=10000)   # 1. shuffle (keeps the model from memorizing the order)
    .batch(32)                    # 2. group into batches of 32
    .prefetch(tf.data.AUTOTUNE))  # 3. prepare the next batch while training

# It is used directly in fit, without batch_size (the dataset already carries it):
# model.fit(dataset, epochs=5)

What each step does and why it matters:

  • shuffle(buffer_size): keeps a buffer of elements and draws at random. If the reviews CSV is sorted by product, without shuffle each batch would be single-topic and training would oscillate. Ideally buffer_size ≥ dataset size (if it fits).
  • batch(32): the mini-batch gradient descent from 02-04 needs batches; this is where they are formed.
  • prefetch(AUTOTUNE): while the GPU trains on batch N, the CPU prepares batch N+1. Without prefetch, the GPU sits idle waiting for data; with it, the pipe never runs dry. AUTOTUNE lets TensorFlow decide how many batches to prepare ahead.

For the product photos, the usual pattern adds map to decode and transform in parallel:

def load_photo(path, label):
    img = tf.io.read_file(path)                    # read bytes from disk
    img = tf.image.decode_jpeg(img, channels=3)    # bytes -> tensor (h, w, 3)
    img = tf.image.resize(img, [224, 224]) / 255.0 # fixed size and normalize
    return img, label

photo_ds = (tf.data.Dataset.from_tensor_slices((paths, photo_labels))
    .map(load_photo, num_parallel_calls=tf.data.AUTOTUNE)  # in parallel
    .shuffle(1000)
    .batch(32)
    .prefetch(tf.data.AUTOTUNE))

This is how the TecnoMarket team would feed the MobileNetV2 from 05-03 with their full catalog without exhausting memory: images are read from disk exactly when they are needed.

Recommended order: map (transform) → shuffle → batch (group) → prefetch (anticipate). Shuffling before batching guarantees different batches every epoch.

Useful callbacks during training

Callbacks are objects that Keras invokes at key moments of training (end of each epoch, of each batch...). You already used EarlyStopping in 05-04; let's complete the basic kit:

callbacks = [
    # You know this one from 05-04: stop when validation stops improving
    keras.callbacks.EarlyStopping(patience=3, restore_best_weights=True),

    # NEW: automatically save the best model seen so far
    keras.callbacks.ModelCheckpoint(
        "tecnomarket-dl/models/best_model.keras",
        monitor="val_loss",        # which metric to watch
        save_best_only=True),      # only overwrite when it improves

    # NEW: log metrics to visualize them with TensorBoard
    keras.callbacks.TensorBoard(log_dir="tecnomarket-dl/logs"),
]

# model.fit(dataset, validation_data=val_ds, epochs=50, callbacks=callbacks)
  • ModelCheckpoint with save_best_only=True is your safety net: even if training is cut short halfway through (Colab disconnects, the power goes out), the best model is safe on disk. Compare with the manual model.save() from 02-05: that one saved at the end; this one saves the best, as you go.
  • TensorBoard writes logs that are later displayed as interactive curves; we will see how to use it in practice in 06-04.
  • All three callbacks coexist happily in the same list: they watch, save and log at the same time.

Common Mistakes and Tips

  • Mixing dtypes: tf.constant(1) + tf.constant(1.0) fails (int32 + float32). Fix: tf.cast(x, tf.float32). Tip: write your numbers with a decimal point (1.0) from the start.
  • Forgetting that constants are not tracked: if you define a weight as a tf.constant, tape.gradient will return None. Trainable parameters must be tf.Variable.
  • Using the tape twice: by default a GradientTape is consumed when you call gradient(). If you need several gradients from the same recording, create the tape with persistent=True (and delete it afterwards with del tape).
  • Silent broadcasting: an operation between (n, 1) and (n,) produces (n, n) without warning. Print result.shape whenever in doubt; it is free and saves hours of debugging.
  • shuffle with a tiny buffer: shuffle(10) over 100,000 reviews barely shuffles. Use a buffer the size of the dataset or, if it does not fit, as large as possible.
  • Tip: tensor.numpy() converts any tensor to numpy so you can inspect or plot it. It is your bridge back to familiar territory.

Exercises

Exercise 1: tensors and broadcasting

TecnoMarket has 4 products priced [120.0, 15.5, 899.0, 45.0] and wants to apply three discount scenarios: 5%, 15% and 30%. Using broadcasting (no loops), build a table of shape (4, 3) where each row is a product and each column a scenario with the final price. Check the shape of the result.

Exercise 2: gradient with GradientTape

Use tf.GradientTape to compute the gradient of L(w, b) = (3*w + b - 10)**2 with respect to w and b, starting from w=1.0, b=0.0. Then write a gradient descent loop (learning rate 0.01, 200 steps) and verify that the final loss is nearly zero. Is there a single possible solution (w, b)?

Exercise 3: tf.data pipeline

Create a tf.data.Dataset from 1,000 synthetic samples (np.random.rand(1000, 20) as features and random binary labels) with shuffle, batches of 64 and prefetch. Iterate over it and print the shape of the first two batches. Why can the last batch of an epoch have fewer than 64 elements, and how would you avoid it?

Solutions

Solution 1:

import tensorflow as tf

prices = tf.constant([[120.0], [15.5], [899.0], [45.0]])   # shape (4, 1)
factors = tf.constant([0.95, 0.85, 0.70])                  # shape (3,)

table = prices * factors        # broadcasting: (4,1) x (3,) -> (4,3)
print(table.shape)              # (4, 3)
print(table.numpy().round(2))

The key is giving the prices shape (4, 1) (a column): that way each size-1 dimension gets "stretched" against the other.

Solution 2:

w = tf.Variable(1.0)
b = tf.Variable(0.0)

with tf.GradientTape() as tape:
    loss = (3.0 * w + b - 10.0) ** 2
grads = tape.gradient(loss, [w, b])
print([g.numpy() for g in grads])   # [-42.0, -14.0]
# dL/dw = 2*(3w+b-10)*3 = 2*(-7)*3 = -42 ; dL/db = 2*(-7)*1 = -14

lr = 0.01
for _ in range(200):
    with tf.GradientTape() as tape:
        loss = (3.0 * w + b - 10.0) ** 2
    gw, gb = tape.gradient(loss, [w, b])
    w.assign_sub(lr * gw)
    b.assign_sub(lr * gb)

print(loss.numpy())   # ~0.0
print(w.numpy(), b.numpy())

There is no unique solution: any pair with 3w + b = 10 zeroes the loss (a whole line of solutions). Gradient descent finds one of them, the one closest to the starting point. This foreshadows why in real networks different initializations yield different final weights with similar performance.

Solution 3:

import numpy as np, tensorflow as tf

X = np.random.rand(1000, 20).astype("float32")
y = np.random.randint(0, 2, size=(1000,))

ds = (tf.data.Dataset.from_tensor_slices((X, y))
      .shuffle(1000)
      .batch(64)
      .prefetch(tf.data.AUTOTUNE))

for i, (xb, yb) in enumerate(ds.take(2)):
    print(f"Batch {i}: X {xb.shape}, y {yb.shape}")   # (64, 20) and (64,)

1000 / 64 = 15 full batches with a remainder of 40: the last batch has 40 elements. If your code requires fixed-size batches, use batch(64, drop_remainder=True) (at the cost of discarding those 40 samples every epoch).

Conclusion

You have opened the hood of the tool you had been using all course long: TensorFlow represents data as tensors (constants for data, variables for weights), computes gradients with GradientTape — the same backpropagation from 02-03, automated —, offers three building APIs (Sequential, functional and subclassing, from least to most flexible), feeds models with tf.data pipelines (shuffle → batch → prefetch) and watches over training with callbacks like ModelCheckpoint. With this, model.fit() has stopped being magic: it is a GradientTape loop with extras.

In the next lesson you will meet the other giant: PyTorch, the framework where that training loop is not packaged — you write it yourself, line by line. You will see that everything you learned today — tensors, autograd, subclassing — has a direct twin over there, and we will rebuild the MNIST network from 02-05 to prove it.

© Copyright 2026. All rights reserved