As we closed the previous module, we said that you already know the techniques — CNNs, RNNs, GANs, autoencoders, transfer learning, attention — and that the time had come to consolidate the tools. We start with the one you have been using since lesson 02-05 without ever looking at it head-on: TensorFlow. Every time you wrote keras.Sequential, model.fit() or model.save(), TensorFlow was underneath doing the heavy lifting: representing data as tensors, computing gradients automatically, and running the operations on CPU or GPU. In this lesson we open the box: you will understand what TensorFlow really is, how its tensors work, how it computes gradients with GradientTape (we will reproduce the manual gradient descent from 02-03), the three ways to build models, and how to feed your networks with efficient data pipelines. For the TecnoMarket team, this is the difference between "following recipes" and "mastering the tool".
Contents
- What TensorFlow is and its ecosystem
- Tensors: constants, variables, shapes and dtypes
- Broadcasting: operating on different shapes
- Automatic differentiation with tf.GradientTape
- The three APIs for building models
- Efficient data pipelines with tf.data
- Useful callbacks during training
What TensorFlow is and its ecosystem
TensorFlow is an open-source machine learning platform created by Google and released in 2015. Its name says it all: it makes tensors (multidimensional arrays) flow through graphs of mathematical operations.
It helps to distinguish the layers of the ecosystem:
| Layer | What it is | Example usage |
|---|---|---|
| TensorFlow (core) | Engine for tensors, operations and automatic differentiation | tf.matmul, tf.GradientTape |
| Keras | High-level API for defining and training models | keras.Sequential, model.fit() |
| tf.data | Building data pipelines | Dataset.from_tensor_slices(...) |
| TensorBoard | Training visualization | Loss curves, histograms |
| TFLite / TF.js / TF Serving | Deployment to mobile, browser and servers | Covered in 06-05 |
The key point for you: Keras IS TensorFlow's official high-level API. You have not been using "something else" throughout the course; you have been using TensorFlow through its most convenient interface. The relationship is like an automatic car (Keras) and its engine (TensorFlow): until now you were driving; today we open the hood.
import tensorflow as tf
from tensorflow import keras
print(tf.__version__) # e.g. 2.16.x
print(keras.__name__) # keras lives inside the TF ecosystem- If you ran
pip install tensorflowin 01-05, you already have everything this lesson needs. - TensorFlow detects the GPU automatically if one is available (in Colab: Runtime → Change runtime type → GPU).
# Which devices can TensorFlow see?
print(tf.config.list_physical_devices())
# [PhysicalDevice(name='/physical_device:CPU:0', ...), maybe a GPU]Tensors: constants, variables, shapes and dtypes
A tensor is a multidimensional array with a uniform data type (dtype). You already know them conceptually from the course: a TecnoMarket product image is a (height, width, 3) tensor, a batch of vectorized reviews is a (batch, length) tensor.
tf.constant: immutable tensors
import tensorflow as tf
scalar = tf.constant(4.99) # price of a USB cable
vector = tf.constant([120.0, 15.5, 899.0]) # prices of 3 products
matrix = tf.constant([[1, 2], [3, 4]]) # 2D tensor
print(scalar.shape) # () -> 0 dimensions
print(vector.shape) # (3,) -> 1 dimension with 3 elements
print(matrix.shape) # (2, 2)
print(matrix.dtype) # <dtype: 'int32'>
print(vector.dtype) # <dtype: 'float32'>Points you should internalize:
- shape: the tensor's shape. Shape errors are by far the most frequent in deep learning; learning to read a message like
Incompatible shapes: (32, 10) vs (32, 1)properly will save you hours. - dtype: the data type. By default floats are
float32(a balance between precision and speed) and integers areint32. You cannot mix dtypes in one operation without converting first withtf.cast. - A
tf.constantis immutable: you cannot change its values once created.
tf.Variable: tensors that learn
The weights of a network are variables: tensors whose value must be updatable at every training step. When in 02-03 we talked about the "blame assignment" that adjusts each weight, those weights are internally tf.Variable objects.
w = tf.Variable(3.0) # a weight initialized to 3.0
b = tf.Variable(0.0) # a bias
w.assign(2.5) # change its value (a constant can't do this!)
w.assign_add(0.1) # w = w + 0.1
print(w.numpy()) # 2.6 -> .numpy() converts to numpytf.constant |
tf.Variable |
|
|---|---|---|
| Mutable | No | Yes (assign, assign_add) |
| Typical use | Input data, hyperparameters | Model weights and biases |
| Tracked by GradientTape | Only if explicitly requested | Yes, automatically |
Basic operations
a = tf.constant([[1.0, 2.0], [3.0, 4.0]])
b = tf.constant([[10.0, 20.0], [30.0, 40.0]])
print(a + b) # element-wise addition
print(a * b) # element-wise product (NOT matrix multiplication!)
print(tf.matmul(a, b)) # matrix product (the one dense layers use)
print(tf.reduce_mean(a)) # mean of all elements -> 2.5
print(tf.reshape(a, (4, 1))) # change the shape without changing the dataRemember from 02-01: a dense layer is exactly tf.matmul(inputs, weights) + bias followed by an activation. Now you can write it by hand.
Broadcasting: operating on different shapes
Broadcasting lets you operate on tensors of different shapes: TensorFlow virtually "stretches" the small tensor so it fits the large one, without copying data. It is the same rule numpy uses.
prices = tf.constant([[120.0], [15.5], [899.0]]) # shape (3, 1): 3 products
vat = tf.constant(1.21) # scalar, shape ()
with_vat = prices * vat # the scalar applies to all 3 -> shape (3, 1)
discounts = tf.constant([0.95, 0.90, 0.80]) # shape (3,): 3 campaigns
table = prices * discounts # (3,1) x (3,) -> broadcasting -> (3, 3)
print(table.shape) # (3, 3): every product with every discountPractical rules (shapes are compared right to left):
- Two dimensions are compatible if they are equal or if one of them is 1.
- Missing dimensions are treated as 1.
- If no rule applies, you get an
Incompatible shapeserror.
Broadcasting is extremely powerful, but also a source of silent bugs: in the example above, maybe you wanted 3 prices with 3 discounts (result (3,)) and got a (3, 3) table without any error. Always check the shape of the result.
Automatic differentiation with tf.GradientTape
Here is the crown jewel. In 02-03 we computed gradients by hand with numpy, applying the chain rule step by step to distribute the blame for the error. TensorFlow does exactly that, but automatically: tf.GradientTape is a "recording tape" that logs every operation you perform inside its context, then plays it backwards to compute derivatives.
Let's reproduce the classic example: minimizing a simple loss function, first the gradient and then the full descent.
import tensorflow as tf
# We want to find the w that minimizes the "loss" L(w) = (w - 4)^2
# From calculus we know the minimum is at w = 4 and that dL/dw = 2*(w - 4)
w = tf.Variable(0.0) # we start far from the minimum
with tf.GradientTape() as tape: # 1. the tape starts recording
loss = (w - 4.0) ** 2 # 2. recorded operations
grad = tape.gradient(loss, w) # 3. play it backwards
print(grad.numpy()) # -8.0 (= 2*(0-4), the exact derivative)Line-by-line explanation:
with tf.GradientTape() as tape:opens the recording context. Everything that happens inside and involves variables gets logged.loss = (w - 4.0) ** 2does not just compute the value (16.0): the tape notes "subtract, then square", which is what it needs to apply the chain rule.tape.gradient(loss, w)asks: "how much doeslosschange if I nudgewa little?". It is the same question backpropagation asked in 02-03.
Now the complete gradient descent, the same loop we wrote with numpy:
w = tf.Variable(0.0)
learning_rate = 0.1
for step in range(30):
with tf.GradientTape() as tape:
loss = (w - 4.0) ** 2
grad = tape.gradient(loss, w)
w.assign_sub(learning_rate * grad) # w = w - lr * gradient
print(w.numpy()) # ~3.995: it has converged to 4, just like in 02-03Comparison with the numpy version from 02-03:
| numpy (02-03) | TensorFlow (GradientTape) | |
|---|---|---|
| Derivative | You write it by hand: 2*(w-4) |
The tape computes it automatically |
| Risk of error | High (one wrong derivative and everything breaks) | Practically zero |
| Deep networks | Unfeasible by hand (thousands of derivatives) | Same effort: one line |
| Educational value | You understand the mechanism | You scale the mechanism |
When you call model.fit() in Keras, this is exactly what happens internally: a GradientTape records the forward pass, computes the gradients of the loss with respect to all the model's variables, and the optimizer (SGD, Adam...) updates them. fit() is this loop, packaged.
The three APIs for building models
TensorFlow/Keras offers three ways to define a model. You already master two of them:
- Sequential (you have known it since 02-05)
For linear stacks of layers: input → layer → layer → output.
model = keras.Sequential([
keras.layers.Dense(128, activation="relu"),
keras.layers.Dense(64, activation="relu"),
keras.layers.Dense(10, activation="softmax"),
])
- Functional API (you used it in 03-03)
For graphs with branches, multiple inputs/outputs or skip connections, like the residual block we built:
inputs = keras.Input(shape=(784,))
x = keras.layers.Dense(128, activation="relu")(inputs)
x = keras.layers.Dense(64, activation="relu")(x)
outputs = keras.layers.Dense(10, activation="softmax")(x)
model = keras.Model(inputs, outputs)
- Subclassing (the new one)
You define a class inheriting from keras.Model and write the forward pass yourself in call(). Maximum flexibility: you can use loops, conditionals, arbitrary logic.
class TecnoMarketClassifier(keras.Model):
def __init__(self):
super().__init__()
self.dense1 = keras.layers.Dense(128, activation="relu")
self.dense2 = keras.layers.Dense(64, activation="relu")
self.output_layer = keras.layers.Dense(10, activation="softmax")
def call(self, x):
x = self.dense1(x) # you decide the data flow
x = self.dense2(x)
return self.output_layer(x)
model = TecnoMarketClassifier()- In
__init__you declare the layers (the parts). - In
callyou connect the parts (the flow). It runs on every forward pass.
| API | When to use it | Flexibility | Already seen in |
|---|---|---|---|
| Sequential | Simple linear stack | Low | 02-05 |
| Functional | Graphs with branches and skips | Medium-high | 03-03 |
| Subclassing | Arbitrary logic, research | Maximum | This lesson |
Advice: always use the simplest API that solves your problem. Subclassing will feel familiar when you meet PyTorch in the next lesson: there it is the standard way of working.
Efficient data pipelines with tf.data
Until now we passed numpy arrays directly to fit(). That works with MNIST, but TecnoMarket's real catalog has hundreds of thousands of photos and reviews that do not fit in memory. tf.data builds pipelines: pipes that load, transform and serve the data in batches, in parallel with training.
import tensorflow as tf
# Simulate already-vectorized TecnoMarket reviews (as in 04-03) and their labels
import numpy as np
reviews = np.random.rand(10000, 200).astype("float32") # 10k reviews
labels = np.random.randint(0, 2, size=(10000,)) # 0=negative, 1=positive
dataset = tf.data.Dataset.from_tensor_slices((reviews, labels))
dataset = (dataset
.shuffle(buffer_size=10000) # 1. shuffle (keeps the model from memorizing the order)
.batch(32) # 2. group into batches of 32
.prefetch(tf.data.AUTOTUNE)) # 3. prepare the next batch while training
# It is used directly in fit, without batch_size (the dataset already carries it):
# model.fit(dataset, epochs=5)What each step does and why it matters:
shuffle(buffer_size): keeps a buffer of elements and draws at random. If the reviews CSV is sorted by product, without shuffle each batch would be single-topic and training would oscillate. Ideallybuffer_size≥ dataset size (if it fits).batch(32): the mini-batch gradient descent from 02-04 needs batches; this is where they are formed.prefetch(AUTOTUNE): while the GPU trains on batch N, the CPU prepares batch N+1. Without prefetch, the GPU sits idle waiting for data; with it, the pipe never runs dry.AUTOTUNElets TensorFlow decide how many batches to prepare ahead.
For the product photos, the usual pattern adds map to decode and transform in parallel:
def load_photo(path, label):
img = tf.io.read_file(path) # read bytes from disk
img = tf.image.decode_jpeg(img, channels=3) # bytes -> tensor (h, w, 3)
img = tf.image.resize(img, [224, 224]) / 255.0 # fixed size and normalize
return img, label
photo_ds = (tf.data.Dataset.from_tensor_slices((paths, photo_labels))
.map(load_photo, num_parallel_calls=tf.data.AUTOTUNE) # in parallel
.shuffle(1000)
.batch(32)
.prefetch(tf.data.AUTOTUNE))This is how the TecnoMarket team would feed the MobileNetV2 from 05-03 with their full catalog without exhausting memory: images are read from disk exactly when they are needed.
Recommended order: map (transform) → shuffle → batch (group) → prefetch (anticipate). Shuffling before batching guarantees different batches every epoch.
Useful callbacks during training
Callbacks are objects that Keras invokes at key moments of training (end of each epoch, of each batch...). You already used EarlyStopping in 05-04; let's complete the basic kit:
callbacks = [
# You know this one from 05-04: stop when validation stops improving
keras.callbacks.EarlyStopping(patience=3, restore_best_weights=True),
# NEW: automatically save the best model seen so far
keras.callbacks.ModelCheckpoint(
"tecnomarket-dl/models/best_model.keras",
monitor="val_loss", # which metric to watch
save_best_only=True), # only overwrite when it improves
# NEW: log metrics to visualize them with TensorBoard
keras.callbacks.TensorBoard(log_dir="tecnomarket-dl/logs"),
]
# model.fit(dataset, validation_data=val_ds, epochs=50, callbacks=callbacks)ModelCheckpointwithsave_best_only=Trueis your safety net: even if training is cut short halfway through (Colab disconnects, the power goes out), the best model is safe on disk. Compare with the manualmodel.save()from 02-05: that one saved at the end; this one saves the best, as you go.TensorBoardwrites logs that are later displayed as interactive curves; we will see how to use it in practice in 06-04.- All three callbacks coexist happily in the same list: they watch, save and log at the same time.
Common Mistakes and Tips
- Mixing dtypes:
tf.constant(1) + tf.constant(1.0)fails (int32+float32). Fix:tf.cast(x, tf.float32). Tip: write your numbers with a decimal point (1.0) from the start. - Forgetting that constants are not tracked: if you define a weight as a
tf.constant,tape.gradientwill returnNone. Trainable parameters must betf.Variable. - Using the tape twice: by default a
GradientTapeis consumed when you callgradient(). If you need several gradients from the same recording, create the tape withpersistent=True(and delete it afterwards withdel tape). - Silent broadcasting: an operation between
(n, 1)and(n,)produces(n, n)without warning. Printresult.shapewhenever in doubt; it is free and saves hours of debugging. shufflewith a tiny buffer:shuffle(10)over 100,000 reviews barely shuffles. Use a buffer the size of the dataset or, if it does not fit, as large as possible.- Tip:
tensor.numpy()converts any tensor to numpy so you can inspect or plot it. It is your bridge back to familiar territory.
Exercises
Exercise 1: tensors and broadcasting
TecnoMarket has 4 products priced [120.0, 15.5, 899.0, 45.0] and wants to apply three discount scenarios: 5%, 15% and 30%. Using broadcasting (no loops), build a table of shape (4, 3) where each row is a product and each column a scenario with the final price. Check the shape of the result.
Exercise 2: gradient with GradientTape
Use tf.GradientTape to compute the gradient of L(w, b) = (3*w + b - 10)**2 with respect to w and b, starting from w=1.0, b=0.0. Then write a gradient descent loop (learning rate 0.01, 200 steps) and verify that the final loss is nearly zero. Is there a single possible solution (w, b)?
Exercise 3: tf.data pipeline
Create a tf.data.Dataset from 1,000 synthetic samples (np.random.rand(1000, 20) as features and random binary labels) with shuffle, batches of 64 and prefetch. Iterate over it and print the shape of the first two batches. Why can the last batch of an epoch have fewer than 64 elements, and how would you avoid it?
Solutions
Solution 1:
import tensorflow as tf
prices = tf.constant([[120.0], [15.5], [899.0], [45.0]]) # shape (4, 1)
factors = tf.constant([0.95, 0.85, 0.70]) # shape (3,)
table = prices * factors # broadcasting: (4,1) x (3,) -> (4,3)
print(table.shape) # (4, 3)
print(table.numpy().round(2))The key is giving the prices shape (4, 1) (a column): that way each size-1 dimension gets "stretched" against the other.
Solution 2:
w = tf.Variable(1.0)
b = tf.Variable(0.0)
with tf.GradientTape() as tape:
loss = (3.0 * w + b - 10.0) ** 2
grads = tape.gradient(loss, [w, b])
print([g.numpy() for g in grads]) # [-42.0, -14.0]
# dL/dw = 2*(3w+b-10)*3 = 2*(-7)*3 = -42 ; dL/db = 2*(-7)*1 = -14
lr = 0.01
for _ in range(200):
with tf.GradientTape() as tape:
loss = (3.0 * w + b - 10.0) ** 2
gw, gb = tape.gradient(loss, [w, b])
w.assign_sub(lr * gw)
b.assign_sub(lr * gb)
print(loss.numpy()) # ~0.0
print(w.numpy(), b.numpy())There is no unique solution: any pair with 3w + b = 10 zeroes the loss (a whole line of solutions). Gradient descent finds one of them, the one closest to the starting point. This foreshadows why in real networks different initializations yield different final weights with similar performance.
Solution 3:
import numpy as np, tensorflow as tf
X = np.random.rand(1000, 20).astype("float32")
y = np.random.randint(0, 2, size=(1000,))
ds = (tf.data.Dataset.from_tensor_slices((X, y))
.shuffle(1000)
.batch(64)
.prefetch(tf.data.AUTOTUNE))
for i, (xb, yb) in enumerate(ds.take(2)):
print(f"Batch {i}: X {xb.shape}, y {yb.shape}") # (64, 20) and (64,)1000 / 64 = 15 full batches with a remainder of 40: the last batch has 40 elements. If your code requires fixed-size batches, use batch(64, drop_remainder=True) (at the cost of discarding those 40 samples every epoch).
Conclusion
You have opened the hood of the tool you had been using all course long: TensorFlow represents data as tensors (constants for data, variables for weights), computes gradients with GradientTape — the same backpropagation from 02-03, automated —, offers three building APIs (Sequential, functional and subclassing, from least to most flexible), feeds models with tf.data pipelines (shuffle → batch → prefetch) and watches over training with callbacks like ModelCheckpoint. With this, model.fit() has stopped being magic: it is a GradientTape loop with extras.
In the next lesson you will meet the other giant: PyTorch, the framework where that training loop is not packaged — you write it yourself, line by line. You will see that everything you learned today — tensors, autograd, subclassing — has a direct twin over there, and we will rebuild the MNIST network from 02-05 to prove it.
Deep Learning Course
Module 1: Introduction to Deep Learning
- What is Deep Learning?
- History and evolution of Deep Learning
- Applications of Deep Learning
- Basic concepts of neural networks
- Setting up the work environment
Module 2: Neural Network Fundamentals
- Perceptron and Multilayer Perceptron
- Activation functions
- Forward and backward propagation
- Optimization and loss functions
- Your first complete neural network
Module 3: Convolutional Neural Networks (CNN)
- Introduction to CNNs
- Convolutional and pooling layers
- Popular CNN architectures
- CNN applications in image recognition
Module 4: Recurrent Neural Networks (RNN)
- Introduction to RNNs
- LSTM and GRU
- RNN applications in natural language processing
- Sequences and time series
Module 5: Advanced Deep Learning Techniques
- Generative Adversarial Networks (GAN)
- Autoencoders
- Transfer Learning
- Regularization and improvement techniques
- Attention mechanisms and Transformers
Module 6: Tools and Frameworks
- Introduction to TensorFlow
- Introduction to PyTorch
- Framework comparison
- Development environments and additional resources
- Saving, loading and deploying models
Module 7: Hands-On Projects
- Image classification with CNNs
- Text generation with RNNs
- Anomaly detection with Autoencoders
- Building a GAN for image generation
- Fine-tuning a pretrained model
