We now know what Machine Learning is and where it comes from. The next step is to organize the field: not all learning problems are alike, and the first decision in any project is identifying what type of learning your problem needs. In this lesson we will study the four major paradigms — supervised, unsupervised, semi-supervised and reinforcement learning — paying special attention to the distinction between regression and classification within supervised learning, and between clustering and dimensionality reduction within unsupervised learning. We will use MercaFresh's problems to see that every business question fits naturally into one of these types. This lesson is a map: the specific algorithms for each type are studied in modules 4 and 5.

Contents

  1. The classification criterion: what information do the data carry?
  2. Supervised learning
  3. Regression vs. classification
  4. Unsupervised learning
  5. Semi-supervised learning
  6. Reinforcement learning
  7. Comparison table of the four types
  8. How to identify the problem type: a practical guide

The classification criterion: what information do the data carry?

The types of ML are distinguished by one very simple question: does our training data include the correct answer (the "label") or not?

  • If every example comes with its known answer → supervised learning.
  • If there are no answers, only "raw" data → unsupervised learning.
  • If only a few examples have an answer → semi-supervised learning.
  • If there is no prior data, but rather an environment in which to try actions and receive rewards → reinforcement learning.
flowchart TD
    Q1{Do you have examples with the<br>correct answer known?}
    Q1 -- "Yes, all of them" --> S[Supervised]
    Q1 -- "Only a few" --> SS[Semi-supervised]
    Q1 -- "No, none" --> NS[Unsupervised]
    Q1 -- "No prior data:<br>you learn by trying actions" --> RF[Reinforcement]
    S --> Q2{What do you predict?}
    Q2 -- "A number" --> R[Regression]
    Q2 -- "A category" --> C[Classification]
    NS --> Q3{What are you looking for?}
    Q3 -- "Natural groups" --> CL[Clustering]
    Q3 -- "Compress variables" --> RD[Dimensionality<br>reduction]

Keep this diagram in your head: it is the first step of any ML project.

Supervised learning

This is the most common type in business practice. It is called "supervised" because during training there is a "supervisor" (the labels) telling the algorithm what the correct answer for each example was, just like the fraud / not fraud labels in the example from lesson 01-01.

Ingredients:

  • Features (X): the input data for each example (order amount, customer tenure...).
  • Label (y): the known correct answer for each example (bought / didn't buy; €37.50...).
  • Goal: learn the X → y relationship in order to predict the label for new cases.

At MercaFresh, supervised learning is possible because the past has already given us the answers: we know how much milk sold on each day of last year, and we know which customers stopped buying. That labeled history is gold for training.

# Typical structure of a supervised problem at MercaFresh (illustrative)
# X: features of each customer          y: known label
X = [
    # [months_tenure, orders_last_quarter, avg_basket_eur]
    [24, 12, 45.0],
    [3,   1, 18.5],
    [36,  9, 60.2],
]
y = ["stays", "churns", "stays"]   # what we know happened afterwards

Look at the structure: each row of X describes a customer with numbers, and the list y records what actually happened to them. The model will learn which combinations of features foreshadow a churn. This (X, y) pair is the unmistakable signature of supervised learning.

Regression vs. classification

Within supervised learning there are two subtypes, depending on the nature of the label:

Regression: predicting a number

The label is a continuous numeric value: it can take any value within a range.

  • MercaFresh: how many kilos of oranges will we sell on Saturday? (demand)
  • MercaFresh: how much will this customer spend next month?
  • Others: the price of a home, tomorrow's temperature, the minutes of delay of a flight.

Classification: predicting a category

The label is a discrete class: one option out of a finite set.

  • MercaFresh: will this customer churn or not in the next 3 months? (churn: 2 classes, binary classification)
  • MercaFresh: is this order legitimate or fraudulent?
  • Others: is this email spam? Does this photo show a tomato, an apple or a pepper? (multiclass classification)
Criterion Regression Classification
The label is... A continuous number A discrete category
Typical question How much? How many? Which one? Yes or no?
MercaFresh example Tomorrow's demand: 342 units Customer: churns / stays
Model output 342.7 "churns" (or a probability: 0.82)
An error is measured as... Distance from the true value Hit or miss on the class
Module where it is studied 4 (e.g. linear regression, 04-01) 4 (e.g. logistic regression, 04-02)

A trick to avoid mixing them up: ask whether it makes sense to say a prediction came "close" to the answer. Predicting 340 units when it was 342 is nearly perfect (regression); predicting "stays" when the customer churned is simply a miss (classification). And a terminological warning we will return to in 04-02: logistic regression, despite its name, is a classification algorithm.

Unsupervised learning

Here there are no labels: only data. The goal is no longer to predict a known answer, but to discover hidden structure in the data. It is like being handed a box of pieces with no reference picture: nobody tells you what should come out; you look for which pieces fit together.

Its two main families:

Clustering

Finding natural groups of examples that are similar to each other and different from the rest.

  • MercaFresh: customer segmentation. Nobody has labeled customers as "weekly-shop families", "last-minute urbanites" or "bargain hunters"; those segments don't exist a priori. A clustering algorithm discovers them by analyzing purchase patterns, and the marketing team names them afterwards.
  • Clustering is also useful for anomaly detection: an order that doesn't fit into any of the usual groups deserves review (we'll see it in 05-04 with DBSCAN).

Dimensionality reduction

Compressing many variables into few, while preserving the essential information.

  • MercaFresh: each customer can be described by hundreds of variables (spend per product category, schedules, frequencies...). Reducing them to 2 or 3 "summary dimensions" makes it possible to visualize the customer base on a chart and to speed up other models.
  • It is studied in 05-03 (PCA) and 05-05 (t-SNE and UMAP).
# Typical structure of an UNSUPERVISED problem (illustrative)
X = [
    # [monthly_spend_eur, num_orders_month, pct_bought_on_offer]
    [420.0, 8, 0.05],
    [95.5,  2, 0.60],
    [380.0, 7, 0.10],
]
# There is no 'y': nobody knows a priori which segment each customer belongs to.
# The algorithm will group similar rows together (here, the 1st and 3rd are alike).

The difference jumps out when you compare it with the previous snippet: the y list is gone. That absence is what makes the problem unsupervised.

Semi-supervised learning

An intermediate situation that is very common in the real world: huge amounts of data, but only a few of them labeled, because labeling is expensive (it requires expert human time).

An example at MercaFresh: to train a fraud detector, an analyst can manually review 500 orders and label them as legitimate or fraudulent, but reviewing the 2 million orders in the history is unfeasible. Semi-supervised learning combines the best of both worlds: it uses the 500 labeled orders as a guide and the 2 million unlabeled ones to better understand the overall structure of the data, achieving better results than using only the 500.

A typical strategy (at a high level): train on the few labeled examples, let the model provisionally label the cases it is very confident about, and retrain incorporating them. In this course we treat it as an overview; in professional practice it shows up mostly in computer vision and text processing.

Reinforcement learning

A paradigm different from the previous ones: it does not learn from a historical dataset, but by interacting with an environment. An agent performs actions, the environment responds with a new state and a reward (positive or negative), and the agent learns by trial and error the strategy (policy) that maximizes cumulative reward.

  • It is the paradigm behind AlphaGo (which we saw in the previous lesson): it played millions of games against itself, receiving winning or losing as its reward.
  • Other uses: robotics, autonomous driving, systems optimization (data center cooling).
  • At MercaFresh it could be applied to dynamic pricing: the agent tries small price adjustments (action), observes the resulting sales (state) and the margin earned (reward), and learns a pricing policy. It is an advanced use with real risks (it experiments on real customers!), which is why it stays outside the practical scope of this course.

The classic analogy: this is how you train a dog. Nobody gives it a manual (labeled data); it tries behaviors and receives treats or indifference, and it learns from that.

Comparison table of the four types

Aspect Supervised Unsupervised Semi-supervised Reinforcement
Labeled data? Yes, all of it No Only a small portion No dataset: environment + rewards
Goal Predict the label of new cases Discover hidden structure Predict while leveraging unlabeled data Learn the best action strategy
Analogy Studying with solved exams Sorting a box of pieces with no picture A few solved exams and many unmarked ones Training a dog with treats
MercaFresh example Demand (regression), churn (classification) Customer segmentation, anomalies Fraud detection with few manual reviews Dynamic pricing
Preparation cost High (labeling required) Low Medium High (designing environment and rewards)
Evaluation Direct: compare prediction vs. reality Indirect: are the groups useful/coherent? Direct on the labeled portion Cumulative reward
In this course Module 4 Module 5 Overview (this lesson) Overview (this lesson)

How to identify the problem type: a practical guide

Faced with a new assignment, ask yourself these questions in order:

  1. What business decision do I want to support? (Without this, there is no project; we'll come back to it in lesson 01-05.)
  2. Is there a historical "correct answer" recorded in my data?
    • Yes, for every example → supervised. Go on to question 3.
    • Yes, but only for a few → semi-supervised.
    • No; I want to explore/group → unsupervised.
    • There's no history, but I can try actions and measure results → reinforcement.
  3. If it's supervised: is the answer a number or a category? Number → regression; category → classification.

Let's practice with real MercaFresh assignments:

Business assignment Analysis Type
"I need to know how many baguettes to order from the supplier each day" There's a history of daily sales (labels) and the answer is a number Supervised → regression
"Alert me to which customers are going to cancel the premium service" The history records who canceled; the answer is yes/no Supervised → classification
"I want to understand what types of customer we have so we can design campaigns" No predefined types exist; they have to be discovered Unsupervised → clustering
"We have 200 variables per customer and I want a visual map of the customer base" Compress variables while preserving information Unsupervised → dimensionality reduction
"Detect fraud; we only have 300 confirmed cases among millions of orders" Very few labels, an enormous amount of unlabeled data Semi-supervised

Common Mistakes and Tips

  • Choosing the algorithm before the problem type. "I want to use neural networks" is not a plan; first identify whether your problem is regression, classification or clustering, and later (modules 4-5) you will choose an algorithm.
  • Treating a regression problem as classification (or vice versa). If you turn demand into categories ("high/medium/low") you throw away valuable information for no reason; do it only if the business genuinely decides in bands.
  • Expecting objective answers from clustering. The groups an algorithm finds come with neither names nor a guarantee of usefulness; validating and interpreting them is your job and the business's.
  • Forgetting that the label must exist before you can predict. To predict churn you need historical examples of customers who already left; if the premium service launched last month, you don't have enough labels yet.
  • Tip: memorize the key questions: how much? → regression; which one / yes-or-no? → classification; what groups are there? → clustering; how do I summarize my variables? → dimensionality reduction.

Exercises

Exercise 1

Classify each MercaFresh assignment by its type (and subtype where applicable): (a) predicting the delivery time in minutes for each order; (b) grouping products by joint-purchase patterns to reorganize the website; (c) predicting whether a customer will rate their order 1-2 stars (dissatisfied) or 3-5 (satisfied); (d) learning a stock replenishment policy by trying decisions in a warehouse simulator and measuring profit.

Exercise 2

MercaFresh's customer service team has hand-labeled 800 support conversations as "urgent" / "not urgent", but there are 500,000 historical conversations without labels. Which type of learning fits best and why? What alternative would you have if you decided to use only pure supervised learning?

Exercise 3

Write down (no code, just the structure) what X and y would look like for MercaFresh's bread demand prediction problem: propose at least 4 features for X and say what y would contain and what type it would be (number/category).

Solutions

Solution 1

  • (a) Supervised → regression: the label (actual delivery minutes) is a continuous number and it exists in the history.
  • (b) Unsupervised → clustering: there are no predefined groups; they are discovered from joint purchases.
  • (c) Supervised → classification (binary): the label is one of two categories and the ratings history provides it.
  • (d) Reinforcement learning: an agent tries actions (restock or not) in an environment (simulator) and learns from the reward (profit).

Solution 2

Semi-supervised learning fits best: there is a small labeled portion (800) and an enormous unlabeled mass (500,000) that can contribute structure and improve the model. The pure supervised alternative would be to train only on the 800 labeled ones (risking a poor model due to scarce examples) or to invest in manually labeling many more conversations, which is expensive and slow.

Solution 3

A possible structure:

  • X (features for each historical day): day of the week, whether it is a holiday or its eve, the price of bread that day, units sold on the same weekday of the previous week, whether a promotion was active, the weather forecast.
  • y (label): units of bread sold that day — a continuous number, so the problem is one of regression.

Each row of X would correspond to a day in the history, with its actual sales in y.

Conclusion

You now have the complete map of the territory: supervised learning (with labels; regression if you predict numbers, classification if you predict categories), unsupervised (no labels; clustering to discover groups, dimensionality reduction to compress variables), semi-supervised (few labels, lots of data) and reinforcement (learning by acting and receiving rewards). And you know how to place each MercaFresh problem in its box: demand → regression, churn → classification, segmentation → clustering. In the next lesson we will step outside our supermarket to see the full picture: the applications of Machine Learning across the different sectors of the economy, and the criteria for deciding when using ML makes sense and when it doesn't.

Machine Learning Course

Module 1: Introduction to Machine Learning

Module 2: Foundations of Statistics and Probability

Module 3: Data Preprocessing

Module 4: Supervised Machine Learning Algorithms

Module 5: Unsupervised Machine Learning Algorithms

Module 6: Model Evaluation and Validation

Module 7: Advanced Techniques and Optimization

Module 8: Model Implementation and Deployment

Module 9: Hands-On Projects

Module 10: Additional Resources

© Copyright 2026. All rights reserved