Deep learning is the technology behind voice assistants, cars that "see" the road, machine translators, and the systems that recommend which movie to watch tonight. In this first lesson we will define exactly what it is, how it relates to artificial intelligence and machine learning, why it is called "deep" and, above all, when it makes sense to use it and when it does not. We will also introduce the running case study that will accompany us throughout the course: TecnoMarket, an online electronics and home goods store that wants to bring deep learning into its day-to-day operations.

Contents

  1. The TecnoMarket case: the thread running through the course
  2. Defining Deep Learning
  3. AI, Machine Learning and Deep Learning: how they relate
  4. Why "deep"? The hierarchy of representations
  5. The key difference from classical machine learning
  6. When to use Deep Learning (and when not to)

The TecnoMarket case: the thread running through the course

Imagine you work on the data team at TecnoMarket, a fictional online store that sells electronics and home goods. Thousands of sellers upload products every day, customers write reviews, payments are processed, and the marketing department constantly needs graphic assets. Management has decided to bet on deep learning to solve some very specific problems:

  • Automatically classifying the product photos that sellers upload (is it a laptop, a coffee maker, a TV?).
  • Analyzing customer reviews to know whether they are positive or negative without reading them one by one.
  • Forecasting product demand to manage the warehouse better.
  • Detecting fraudulent transactions among thousands of daily payments.
  • Generating promotional images for marketing campaigns.

Throughout the course, every technique you learn will be framed as a step in this project. Before applying each technique to TecnoMarket's (fictional) data, we will try it out on standard public datasets (MNIST, CIFAR-10, IMDB...), which serve as "test benches": they are small, well documented, and let you prototype quickly. This way of working — prototype on public data, then apply to the real problem — is exactly how it is done in industry.

Defining Deep Learning

Deep learning is a branch of machine learning that uses artificial neural networks with many layers to learn, directly from the data, the representations needed to solve a task.

Let's unpack that definition:

  • Branch of machine learning: it is not something separate from machine learning, but a specific family of techniques within it.
  • Artificial neural networks: mathematical models inspired (very loosely) by how neurons in the brain connect to one another. We will look at their components in lesson 01-04.
  • With many layers: hence the "deep". The data passes through several stages of transformation, not just one.
  • Learning directly from the data: we do not program rules by hand; the system adjusts its internal parameters from examples.

An example with TecnoMarket: instead of writing rules like "if the image has a rectangular screen and a keyboard, it's a laptop" (something nearly impossible to program well), we show the system thousands of labeled photos ("this is a laptop", "this is a coffee maker") and it learns on its own which visual patterns distinguish each category.

AI, Machine Learning and Deep Learning: how they relate

It is very common to use these three terms interchangeably, but they are not synonyms. They are concentric circles: each one is contained within the previous one.

graph TD
    A["Artificial Intelligence<br/>(any technique that simulates intelligence)"] --> B["Machine Learning<br/>(systems that learn from data)"]
    B --> C["Deep Learning<br/>(neural networks with many layers)"]
Concept What it is Example at TecnoMarket
Artificial Intelligence (AI) The broad field: any technique that lets a machine perform tasks we associate with human intelligence. It ranges from hand-written rule systems to the most modern models. A chatbot with predefined responses like "if the customer types 'return', show form X".
Machine Learning (ML) A subfield of AI: systems that learn patterns from data instead of following explicitly programmed rules. A model that predicts whether a customer will abandon their cart based on their browsing history, using a decision tree.
Deep Learning (DL) A subfield of ML: machine learning with deep neural networks, capable of learning complex representations. A model that looks at a product photo and says which category it belongs to.

The practical takeaway: all deep learning is machine learning, and all machine learning is AI, but not the other way around. When you read in the news that "an AI does X", these days it almost always refers to a deep learning system — but as a professional, you should use the terms precisely.

Why "deep"? The hierarchy of representations

The word deep does not mean "deep" in the sense of "intelligent" or "philosophical". It refers to something very concrete: the number of layers the data passes through inside the network.

What is interesting is what happens in those layers. Each layer transforms the output of the previous one into a slightly more abstract representation. This is what is called a hierarchy of representations:

  1. Early layers: detect very simple patterns. In an image, for example, edges, corners and patches of color.
  2. Middle layers: combine those simple patterns into more complex shapes: textures, circles, grids, contours.
  3. Final layers: combine the shapes into concepts: "wheel", "screen", "coffee maker handle".
  4. Output layer: uses those concepts to make the final decision: "this is a coffee maker with 97% confidence".
graph LR
    A[Photo pixels] --> B[Edges and colors]
    B --> C[Textures and shapes]
    C --> D["Parts: handle, spout, reservoir"]
    D --> E["Concept: coffee maker"]

The human analogy helps: when you recognize a coffee maker, you do not consciously analyze it pixel by pixel. Your visual system also builds hierarchies: light → contours → shapes → object. Deep learning replicates this idea mathematically.

Key point: nobody programs what each layer should detect. The network discovers on its own which intermediate representations are useful during training. This brings us to the fundamental difference from classical ML.

The key difference from classical machine learning

In classical machine learning (logistic regression, decision trees, SVMs, random forests...), the workflow has a critical manual step: feature engineering. A human expert decides which variables to extract from the raw data to feed the model.

In deep learning, that step disappears: the network receives the raw data (or nearly raw) and learns the relevant features by itself.

Aspect Classical ML Deep Learning
Feature extraction Manual, done by experts Automatic, learned by the network
Typical input data Tables of pre-engineered variables Raw data: images, text, audio
Domain knowledge required High (you need to know what to extract) Lower for the features (though still useful)
Performance with little data Usually better Usually worse
Performance with massive data Plateaus Keeps improving
Interpretability Generally high Generally low ("black box")

Let's see it with TecnoMarket's photo classifier, in conceptual pseudocode (this is not real training code — that comes in module 2):

# CLASSICAL APPROACH: a human designs the features
def extract_features_by_hand(photo):
    features = {
        "width_height_ratio": compute_ratio(photo),
        "dominant_color": most_frequent_color(photo),
        "has_straight_edges": detect_lines(photo),
        "amount_of_metal": estimate_metallic_shine(photo),
        # ... and how do I describe "looks like a coffee maker" numerically?
    }
    return features

classical_model.train(extract_features_by_hand(photos), labels)

# DEEP LEARNING APPROACH: the network learns the features on its own
deep_model.train(raw_photos, labels)

Notice the problem with the classical approach: what numeric features distinguish a coffee maker from an electric kettle? They are extremely hard to write by hand, which is why computer vision advanced slowly for decades. Deep learning removed that bottleneck: you feed it the raw photos and the labels, and the hierarchy of representations we saw earlier emerges on its own from training.

This does not mean classical ML is obsolete. For tabular data (a table of customers with age, monthly spend and number of orders), classical methods are still often the best choice: faster, cheaper and more interpretable.

When to use Deep Learning (and when not to)

Deep learning is a powerful tool, but it is not the answer to everything. As TecnoMarket's future technical leads, your first professional skill is knowing when to apply it.

Deep learning is a good fit when...

  • The data is unstructured: images, free text, audio, video. This is where DL crushes classical ML.
  • You have lots of data (thousands or, better, millions of labeled examples), or you can leverage pretrained models (we will see this in module 5 with transfer learning).
  • The problem is too complex for hand-written rules: nobody knows how to write the rules that define "sarcastic review" or "photo of a coffee maker".
  • Accuracy matters more than explainability: you accept a certain degree of "black box" in exchange for better results.
  • You have compute resources (GPU) or can use them in the cloud.

It is NOT a good fit (or a worse one) when...

  • You have little data and no applicable pretrained model: with 200 examples, a well-tuned classical model usually wins.
  • The data is tabular and the problem is simple: to predict whether a customer will pay on time from 10 variables, a classical gradient boosting model will probably perform as well or better at a fraction of the cost.
  • You need to explain every decision: in regulated contexts (credit approval, diagnoses), interpretability may be a legal requirement.
  • The computational cost is not worth it: training deep networks consumes time, energy and money. A model that takes days to train for a 0.5% improvement may not justify the bill.
  • A simple solution already works: the golden rule of engineering: start simple, add complexity only when needed.

Applying it to TecnoMarket

TecnoMarket problem Deep learning? Why
Classifying product photos Yes Images = unstructured data; impossible with hand-written rules
Analyzing review sentiment Yes Free text, complex linguistic nuances
Forecasting demand from sales history It depends If there are complex temporal patterns and lots of data, yes (module 4); if it is a short, stable series, classical statistical methods may suffice
Detecting fraud It depends The usual combination: rules + classical ML + DL for subtle patterns
Calculating the VAT on an invoice No It is a deterministic formula; there is nothing to "learn"

This table sums up the right mindset: deep learning is an excellent hammer, but not everything is a nail.

Common Mistakes and Tips

  • Mistake: using "AI", "ML" and "deep learning" as synonyms. They are nested sets. Terminological precision = professional credibility.
  • Mistake: believing that "deep" means "smarter". It only refers to the number of layers. A poorly trained deep network is worse than a simple model done well.
  • Mistake: thinking that deep learning requires no human work. It eliminates manual feature engineering, but it demands a great deal of work on data: collecting it, cleaning it, labeling it. In real projects, 70-80% of the time goes into the data.
  • Mistake: starting with deep learning "because it's trendy". Always start with the simplest solution that could work; scale up to DL when the problem justifies it.
  • Tip: when someone proposes a project, ask yourself three questions: is the data unstructured? How many examples do I have? Do I need to explain the decisions? With those three answers, the DL vs. classical ML choice almost makes itself.

Exercises

Exercise 1: Classify the techniques

Indicate whether each of these systems is AI (without ML), classical ML, or deep learning:

  1. A system that blocks orders if the amount exceeds €5,000 and the account is less than 7 days old (a fixed, programmed rule).
  2. A decision tree that predicts whether a customer will return a product using 8 variables from their profile.
  3. A 50-layer neural network that identifies the product shown in a photo.

Exercise 2: Deep learning, yes or no?

For each TecnoMarket need, reason about whether you would recommend deep learning, classical ML, or no learning technique at all, and why:

  1. Automatically transcribing customer service phone calls.
  2. Calculating shipping costs based on weight and destination.
  3. Predicting a customer's annual spend from 12 profile variables, with a history of 3,000 customers.

Exercise 3: The hierarchy of representations

Explain in your own words, in 4 levels (from the simplest to the most abstract), what each "zone" of layers might learn in a deep network that analyzes TecnoMarket's text reviews to decide whether they are positive or negative. (Hint: think of the textual equivalent of "edges → shapes → parts → object".)

Solutions

Solution 1:

  1. AI without ML: it is a hand-written rule system; it simulates an "intelligent" decision but learns nothing from data.
  2. Classical ML: it learns from data, but uses pre-engineered features (profile variables) and a non-neural model.
  3. Deep learning: a neural network with many layers working on raw data (pixels).

Solution 2:

  1. Deep learning: audio is unstructured data and speech transcription is a problem where DL is the state of the art. Moreover, pretrained models exist that avoid training from scratch.
  2. No learning technique: it is a deterministic formula (rate by weight and zone). Using ML here would be a design mistake: it would add uncertainty to something that can be calculated exactly.
  3. Classical ML: tabular data, few variables and a modest dataset (3,000 rows). A classical model will probably be just as accurate or more, faster to train and easier to explain.

Solution 3 (suggested answer):

  1. Level 1 — basic units: characters or individual words ("perfect", "broken", "never").
  2. Level 2 — local combinations: pairs or groups of words with joint meaning ("doesn't work", "very happy", "would buy again").
  3. Level 3 — structures: sentences with nuances such as negations, comparisons or sarcasm ("I expected more from this brand").
  4. Level 4 — global concept: the overall sentiment of the review (positive/negative), integrating everything above.

Your answer does not need to match word for word: what matters is the idea that each level combines the patterns of the previous one into something more abstract.

Conclusion

In this lesson we have laid the conceptual foundations of the course: deep learning is the branch of machine learning that uses neural networks with many layers to learn hierarchies of representations directly from the data, eliminating the manual feature engineering that limits classical ML. We have also seen that it is not a silver bullet: it shines with abundant unstructured data, but with scarce tabular data or deterministic problems there are better and cheaper options. And we have met TecnoMarket, the online store that will serve as our project throughout the course.

Now that you know what deep learning is, the next natural question is where it comes from: why did an idea from the 1950s take more than half a century to change the world? We will find out in the next lesson, devoted to the history and evolution of deep learning.

© Copyright 2026. All rights reserved