You now master the two big frameworks and know when to use each. But there is an uncomfortable truth the TecnoMarket team discovered when reviewing their work from the last few modules: almost everything lives in loose Colab tabs, with cells run out of order, results nobody can reproduce and models saved under names like final_model_v2_GOOD.keras. In 01-05 you set up the basic environment; this lesson makes it professional: you will understand the real hardware and platform options (Colab and its limits, alternatives, a local GPU), how to go from notebook to a reproducible project with scripts and random seeds, how to apply version control to an ML project (and what NOT to push to git), how to monitor training runs with TensorBoard, and where to keep learning once the course ends.
Contents
- Colab in depth: what it gives and what it limits
- Alternatives: Kaggle Notebooks and other platforms
- Working locally with a GPU: CUDA and cuDNN
- From notebook to project: structure and scripts
- Reproducibility: random seeds and configuration
- Version control for ML
- Monitoring training runs with TensorBoard
- Where to keep learning
- TecnoMarket's working environment
Colab in depth: what it gives and what it limits
Google Colab, your companion since 01-05, is a notebook service with a free GPU. It pays to know its rules of the game:
| Aspect | Free Colab | Colab Pro / Pro+ (paid) |
|---|---|---|
| GPU | Yes, subject to availability (often a T4) | Priority and better GPUs (depending on plan) |
| Session length | Limited (hours); disconnects on inactivity | Longer sessions |
| RAM / disk | Limited | Expanded |
| Persistence | None: the disk is wiped on close | Same: the disk remains ephemeral |
| Cost | $0 | Monthly subscription |
The two limitations that hurt most in practice:
- The disk is ephemeral. Anything you do not save elsewhere (Google Drive, manual download) disappears when the session ends. That is why the
ModelCheckpointfrom 06-01 pointing at Drive is a lifesaver: if Colab disconnects at epoch 40 of 50, the best model is safe. - Sessions get cut off. Multi-hour training runs are not reliably feasible on the free plan. Rule of thumb: free Colab for learning and prototyping (this whole course fits there); Pro for serious personal projects; your own hardware or the cloud for sustained professional work.
# Essential pattern in Colab: mount Drive to persist models and logs
from google.colab import drive
drive.mount("/content/drive")
PROJECT_PATH = "/content/drive/MyDrive/tecnomarket-dl"
# From here on, checkpoints and logs are saved to PROJECT_PATH/models, /logs...Alternatives: Kaggle Notebooks and other platforms
- Kaggle Notebooks: the most direct free alternative to Colab. It offers a GPU with a weekly quota of hours, plus two advantages of its own: the platform's datasets mount with one click (no downloads), and the public notebooks from competitions are a goldmine of real applied techniques — reading winning solutions is among the most instructive things you can do.
- Papers with Code: not an execution environment but an index that links every paper with its implementation and its benchmarks. When we talked about ResNet in 03-03 or transformers in 05-05, that is where you would find each architecture's reference code and the state of the art of each task.
- Pay-as-you-go cloud (the big providers and GPU-specialized platforms): you rent a GPU machine by the hour. It is the natural step when a serious training run no longer fits in Colab but does not justify buying hardware. Keep it on your radar; you do not need it for this course.
Working locally with a GPU: CUDA and cuDNN
If you have (or are considering buying) a PC with an NVIDIA GPU, you can train locally without session limits. Understand the layers conceptually, because installation errors are almost always a mismatch between them:
graph TB
A["Your code (Keras / PyTorch)"] --> B["Framework (TensorFlow / torch)"]
B --> C["cuDNN: optimized deep learning primitives<br/>(convolutions, LSTM...)"]
C --> D["CUDA: general-purpose GPU computing platform"]
D --> E["NVIDIA driver"]
E --> F["Physical GPU"]
- CUDA is NVIDIA's platform for running general-purpose computation on the GPU.
- cuDNN is a library on top of CUDA with the deep learning operations (the convolutions from 03-02, the LSTMs from 04-02) ultra-optimized. Both frameworks use it underneath: that is why their performance ties (06-03).
- Version compatibility (driver ↔ CUDA ↔ cuDNN ↔ framework) is the classic source of headaches. Good news: current versions of PyTorch (and TensorFlow via pip) bundle the necessary CUDA libraries; usually a reasonably up-to-date NVIDIA driver plus the install command shown on each framework's official website is enough.
- Without an NVIDIA GPU there is no CUDA: on Apple Silicon Macs the frameworks use the built-in GPU by another route (Metal/MPS), and on any machine the CPU always works, just more slowly (for the dense networks in this course, perfectly viable).
Verification in both frameworks (always your first step in a fresh environment):
import tensorflow as tf
print(tf.config.list_physical_devices("GPU")) # [] if no GPU is visible
import torch
print(torch.cuda.is_available()) # True / False
print(torch.cuda.get_device_name(0) if torch.cuda.is_available() else "CPU")From notebook to project: structure and scripts
Notebooks are perfect for exploring but terrible as a final product: cells executable in any order, hidden state in memory, hard to version and to automate. The professional transition is moving the stable code into scripts inside a project structure. The tecnomarket-dl/ folder you created in 01-05 grows like this:
tecnomarket-dl/ ├── data/ # raw and processed data (NOT pushed to git) ├── models/ # trained weights (NOT pushed to git; in use since 02-05) ├── logs/ # TensorBoard logs ├── notebooks/ # exploration: the .ipynb files live here, and only exploration ├── src/ # stable source code │ ├── data.py # loading and preprocessing (the tf.data pipeline from 06-01) │ ├── model.py # architecture definition │ └── train.py # training script ├── config.yaml # hyperparameters and paths ├── requirements.txt # dependencies (you created it in 01-05) └── README.md # what this is and how to run it
The recommended workflow, which is the one TecnoMarket adopts:
- Explore in a notebook (
notebooks/): try ideas, visualize data, iterate fast. - Consolidate into
src/: when something works, it becomes functions inside modules. - Train by script:
python src/train.pyalways produces the same process end to end, runnable by anyone on the team (or by a machine, every night).
A minimal train.py with external configuration:
# src/train.py
import yaml
from data import load_datasets # your src/ modules
from model import create_model
with open("config.yaml") as f: # hyperparameters are NOT hardcoded
cfg = yaml.safe_load(f)
train_ds, val_ds = load_datasets(cfg["data_path"], cfg["batch_size"])
model = create_model(cfg["layers"], cfg["dropout"])
model.fit(train_ds, validation_data=val_ds, epochs=cfg["epochs"])
model.save(f"models/{cfg['experiment_name']}.keras")# config.yaml
experiment_name: reviews_v3
data_path: data/reviews.csv
batch_size: 32
epochs: 20
layers: [64, 32]
dropout: 0.3Why separate the configuration? Because every experiment ends up described by a file: changing the dropout is no longer editing code, it is editing a value; and comparing two experiments is comparing two configs. This is the handcrafted precursor of the tracking tools we will meet shortly.
Reproducibility: random seeds and configuration
Two "identical" training runs give different results: random weight initialization (02-01), batch shuffling, dropout (05-04)... To truly compare experiments, fix the random seeds at the top of the script:
import random, numpy as np
SEED = 42
random.seed(SEED)
np.random.seed(SEED)
import tensorflow as tf
tf.random.set_seed(SEED) # TensorFlow/Keras
import torch
torch.manual_seed(SEED) # PyTorch (CPU and GPU)Honest nuances:
- With fixed seeds, two runs on the same machine and versions will (almost always) be identical.
- Some GPU operations are non-deterministic by design (for performance); full determinism requires extra options and a speed cost. For everyday work, fixing seeds is enough: it turns "results that dance around" into "comparable results".
- Complete reproducibility is seeds + library versions (your
requirements.txtfrom 01-05, ideally with pinned versions:tensorflow==2.16.1) + configuration (theconfig.yaml) + the same data.
Version control for ML
Git lets you keep the project's history, go back to any point and collaborate without stepping on each other. The basic cycle applied to tecnomarket-dl/:
cd tecnomarket-dl
git init # once: turn the folder into a repository
git add src/ config.yaml requirements.txt README.md
git commit -m "Project structure and reviews_v3 training"
# ...days later, after changing the model...
git add src/model.py config.yaml
git commit -m "Add dropout 0.3 to the review classifier"
git log --oneline # change historyWhat does NOT get committed
Git is designed for text (code, configs). Large binary files degrade it:
| Do not push | Why | Where it goes instead |
|---|---|---|
data/ (datasets) |
Large, sometimes containing personal customer data | Your own storage + a download script/instructions |
models/ (.keras, .pt weights) |
MB/GB binaries that change on every training run | Artifact storage; in git only the how to reproduce them |
TensorBoard logs/ |
Regenerable, bulky | They stay local or in their tool |
| Credentials and API keys | Severe, permanent security risk (the history never forgets) | Environment variables, out of the repo |
The .gitignore file at the project root automates it:
The mental rule: git holds whatever lets you REGENERATE the results (code, config, requirements), not the results themselves.
DVC and MLflow: the next level
Two names for your radar (you do not need them yet, but you will run into them):
- DVC (Data Version Control): extends the git idea to data and models — git stores a lightweight pointer, the heavy file lives in external storage, and every commit knows which version of the data it used.
- MLflow: experiment tracking — it automatically records parameters, metrics and artifacts for every training run and shows them in a comparable table. It is the industrial version of your
config.yaml+ spreadsheet of results.
Monitoring training runs with TensorBoard
In 06-01 you added the TensorBoard callback; now we are going to actually use it. TensorBoard reads the logs written during training and turns them into interactive charts: loss and accuracy curves, run comparisons, weight histograms.
import datetime
from tensorflow import keras
# One subdirectory per run, timestamped: that is how runs get compared
log_dir = "logs/reviews_" + datetime.datetime.now().strftime("%Y%m%d-%H%M%S")
tb = keras.callbacks.TensorBoard(log_dir=log_dir)
model.fit(train_ds, validation_data=val_ds, epochs=20, callbacks=[tb])To view it:
# In Colab / Jupyter, embedded in the notebook itself:
%load_ext tensorboard
%tensorboard --logdir logsHow to read what you will see (connecting with what you have learned):
- Scalars tab:
loss/val_losscurves and per-epoch metrics. The divergence between training and validation is the overfitting you learned to fight in 05-04 — here you watch it draw itself live, without waiting for the end. - Comparing runs: each subdirectory of
logs/shows up as a differently colored curve. Train with dropout 0.2 and 0.4 (two configs, two runs) and compare them overlaid: this turns hyperparameter tuning into a visual decision. - PyTorch plays too: with
torch.utils.tensorboard.SummaryWriteryou write the same logs from your explicit loop of 06-02 (writer.add_scalar("loss/train", loss.item(), step)). A single monitoring tool for both of the team's frameworks.
Where to keep learning
Stable resources, with no URLs that expire — all of them are found by searching their name:
- Official documentation for TensorFlow/Keras and for PyTorch: both include excellent guided tutorials and are the final reference for any API question. PyTorch's official tutorials ("60 Minute Blitz") complement lesson 06-02 perfectly.
- Classic courses and books: the deep learning courses on the major educational platforms (Andrew Ng's are the historical standard); Deep Learning by Goodfellow, Bengio and Courville (the reference theory, free online); Hands-On Machine Learning by Géron (practical, with Keras); the fast.ai course and book (a "code-first" approach, on PyTorch).
- Papers: arXiv (preprints of ongoing research) and Papers with Code (paper + implementation + benchmark). Start with the classics you already know by name from the course: ResNet (03-03), Attention Is All You Need (05-05).
- Communities: Kaggle (competitions and public notebooks), Stack Overflow (specific errors), each framework's official forums and Hugging Face (models and examples).
- Deliberate practice: the best post-course recipe is to pick a public dataset that interests you, replicate the course methodology (prototype → application) and document it in your own git repository. One finished project teaches more than ten tutorials — and it is exactly what we will do together in module 7.
TecnoMarket's working environment
Here is how the data team is organized after this lesson:
- Exploration: Colab (free for short experiments; Pro for the long training runs of module 7), always with Drive mounted and
ModelCheckpointpointing at it. - Project: a
tecnomarket-dlgit repository with thesrc/+config.yaml+requirements.txtstructure with pinned versions;data/,models/andlogs/in.gitignore. - Consolidation rule: nothing moves from
notebooks/tosrc/without fixed random seeds and without being runnable withpython src/train.pyend to end. - Monitoring: TensorBoard with one subdirectory per experiment; both frameworks write to the same
logs/. - Pending for when they grow: DVC to version the reviews dataset and MLflow for tracking, noted in the README as the next step.
Common Mistakes and Tips
- Training for hours on Colab without checkpoints on Drive: the disconnection will come, and with an ephemeral disk it takes your model with it.
ModelCheckpoint+ mounted Drive, always. - Notebooks with cells run out of order: hidden state makes it "work in my notebook" and fail on re-run. Before signing anything off: Restart & Run All. If it does not survive that, it is not finished.
- Committing data, weights or (worse) credentials: besides bloating the repo, a key pushed to git stays in the history even if you delete it later. Write the
.gitignorebefore the first commit. - Comparing experiments without fixing seeds: the 0.4% improvement of your new dropout may be pure initialization luck. Fixed seeds first; conclusions second.
- Installing CUDA by hand without needing to: first try the framework's official install command — nowadays it usually includes everything. Only if
torch.cuda.is_available()returnsFalsewith a GPU present is it time to investigate drivers. - Tip: give every logs directory and every config a date and a descriptive name (
reviews_dropout03_20260824). Your three-weeks-from-now self does not remember whattest2_finalwas.
Exercises
Exercise 1: designing the .gitignore
A TecnoMarket teammate's repository contains: src/train.py, config.yaml, data/customer_reviews.csv (400 MB, with real emails), models/reviews_v3.keras (85 MB), logs/ (2 GB), requirements.txt, api_keys.txt and notebooks/exploration.ipynb. State what should be versioned in git, what should not, and why; write the resulting .gitignore and point out the file that is a problem beyond its size.
Exercise 2: reproducibility
A teammate runs the same MNIST network notebook twice, gets 97.68% and 97.41%, and concludes that "the model came out worse the second time". Explain why the conclusion is wrong, which four elements they should fix/record so the experiments are comparable, and write the seed block for a script that uses TensorFlow and numpy.
Exercise 3: comparing experiments with TensorBoard
Write the code outline (Keras) to train the review classifier twice — dropout 0.2 and dropout 0.5 — so that both runs appear as separate, comparable curves in TensorBoard, and explain what you would look at in the curves to decide which of the two values is better.
Solutions
Solution 1:
- Yes to git:
src/train.py,config.yaml,requirements.txt,notebooks/exploration.ipynb(text, lets the work be regenerated). - No to git:
data/(large and containing personal customer data — beyond the size, pushing it could violate data protection),models/(regenerable binary),logs/(regenerable and huge),api_keys.txt. - The serious problem:
api_keys.txt. It is not about size: a committed credential stays in the git history forever (deleting it from the directory does not delete it from old commits) and must be considered compromised — it has to be revoked and regenerated, and the keys moved to environment variables.
Solution 2: The 97.68% vs. 97.41% difference falls within normal random variation: different weight initialization, different batch shuffling, dropout switching off different neurons. There is no "worse model": there are two samples from the same distribution of results. To truly compare: (1) fixed random seeds, (2) library versions pinned in requirements.txt, (3) recorded configuration (config.yaml), (4) the same data and the same train/test split. And run it as a script top to bottom, not cells out of order.
import random, numpy as np
import tensorflow as tf
SEED = 42
random.seed(SEED)
np.random.seed(SEED)
tf.random.set_seed(SEED)Solution 3:
from tensorflow import keras
for dropout in [0.2, 0.5]:
log_dir = f"logs/reviews_dropout{dropout}" # one subdirectory per experiment
model = keras.Sequential([
keras.layers.Dense(64, activation="relu"),
keras.layers.Dropout(dropout),
keras.layers.Dense(32, activation="relu"),
keras.layers.Dense(1, activation="sigmoid"),
])
model.compile(optimizer="adam", loss="binary_crossentropy", metrics=["accuracy"])
model.fit(train_ds, validation_data=val_ds, epochs=20,
callbacks=[keras.callbacks.TensorBoard(log_dir=log_dir)])When you launch tensorboard --logdir logs, each dropout is one curve. What to look at: val_loss (not the training one) — the dropout with the lowest stable validation loss wins; and the gap between each experiment's training and validation curves — if with 0.2 the gap widens epoch after epoch (overfitting, as in 05-04) while with 0.5 they stay together with a similar or better val_loss, 0.5 is the choice.
Conclusion
The TecnoMarket team no longer works in loose tabs: they know what to expect from Colab and when to jump to Kaggle, a local GPU or the cloud; they understand the CUDA/cuDNN layers just enough to diagnose; they have turned their work into a project with src/, configs and random seeds that anyone can reproduce; they version in git whatever regenerates results (and only that, with .gitignore protecting data, weights and credentials); they monitor every experiment in TensorBoard; and they have a map of resources to keep growing. In one word: professionalization.
One last piece of the tools module remains, and it is the one that connects everything to the real world: a model living in models/ is of no use to any TecnoMarket customer. In the next lesson you will learn to save models properly in both frameworks, to package them together with their preprocessing, and to deploy them: from a batch script to a REST API answering live requests.
Deep Learning Course
Module 1: Introduction to Deep Learning
- What is Deep Learning?
- History and evolution of Deep Learning
- Applications of Deep Learning
- Basic concepts of neural networks
- Setting up the work environment
Module 2: Neural Network Fundamentals
- Perceptron and Multilayer Perceptron
- Activation functions
- Forward and backward propagation
- Optimization and loss functions
- Your first complete neural network
Module 3: Convolutional Neural Networks (CNN)
- Introduction to CNNs
- Convolutional and pooling layers
- Popular CNN architectures
- CNN applications in image recognition
Module 4: Recurrent Neural Networks (RNN)
- Introduction to RNNs
- LSTM and GRU
- RNN applications in natural language processing
- Sequences and time series
Module 5: Advanced Deep Learning Techniques
- Generative Adversarial Networks (GAN)
- Autoencoders
- Transfer Learning
- Regularization and improvement techniques
- Attention mechanisms and Transformers
Module 6: Tools and Frameworks
- Introduction to TensorFlow
- Introduction to PyTorch
- Framework comparison
- Development environments and additional resources
- Saving, loading and deploying models
Module 7: Hands-On Projects
- Image classification with CNNs
- Text generation with RNNs
- Anomaly detection with Autoencoders
- Building a GAN for image generation
- Fine-tuning a pretrained model
