In this first lesson we are going to answer the fundamental question of the course: what exactly is Machine Learning? We will look at both a formal and an intuitive definition, understand how it differs from traditional programming, and get to know the basic concepts we will use throughout the course: model, training, prediction and training data. We will also introduce MercaFresh, a fictional online supermarket whose business problems will serve as the running thread through every lesson. Getting these foundations right matters, because everything else — statistics, preprocessing, algorithms, evaluation — is built on top of them.
Contents
- An intuitive definition
- The formal definition
- Traditional programming vs. Machine Learning
- Key concepts: data, model, training and prediction
- MercaFresh: the case study that runs through the course
- What Machine Learning is (and what it is not)
An intuitive definition
Imagine you want to teach a child to tell apples from oranges. You don't hand them a list of rules like "if the diameter is greater than 7 cm and the roughness exceeds index 0.3, then it's an orange". You simply show them lots of apples and lots of oranges, tell them which is which, and their brain learns from the examples. After a while, the child can recognize a fruit they have never seen before.
Machine Learning works on the same philosophy:
- Instead of writing rules by hand, we show examples to the computer.
- The computer detects patterns in those examples.
- With those patterns, it can generalize: give correct answers for new cases it has never seen.
Put in a single sentence: Machine Learning is the discipline that enables computers to learn from data without being explicitly programmed for each specific case.
The formal definition
The most widely cited definition is by Tom Mitchell (1997):
"A program is said to learn from experience E with respect to a task T and a performance measure P, if its performance on T, as measured by P, improves with experience E."
It sounds abstract, but it becomes clear right away with an example applied to our online supermarket:
| Component | Meaning | Example at MercaFresh |
|---|---|---|
| Task (T) | What we want the program to do | Predict how many units of milk will sell tomorrow |
| Experience (E) | The data it learns from | Daily sales history for the last 3 years |
| Performance (P) | How we measure whether it does it well | Average difference between the prediction and actual sales |
If the program, as it sees more sales history (E), predicts demand (T) with less error (P), then it is learning. That's how concrete machine learning is: it isn't magic, it is measurable improvement from data.
Traditional programming vs. Machine Learning
This is the single most important distinction in the whole lesson. In traditional programming, the programmer writes the rules; in Machine Learning, the rules come out of the data.
flowchart LR
subgraph Traditional["Traditional programming"]
A1[Data] --> P1[Program with hand-written rules]
R1[Rules] --> P1
P1 --> S1[Answers]
end
subgraph ML["Machine Learning"]
A2[Data] --> P2[Learning algorithm]
S2[Known answers] --> P2
P2 --> R2[Model: learned rules]
end
Notice how the flow changes: in the traditional approach, data and rules go in, and answers come out. In Machine Learning, data and known answers go in, and a model comes out (the learned rules), which we will then use to answer new cases.
A code example: flagging suspicious orders
Suppose MercaFresh wants to detect potentially fraudulent orders. With traditional programming we would write rules by hand:
# Traditional approach: rules written by a person
def is_suspicious_order(order_amount, num_items, new_customer):
"""Returns True if the order looks suspicious, based on fixed rules."""
if order_amount > 500 and new_customer:
return True # very large order from a freshly registered customer
if num_items > 100:
return True # abnormally high number of items
return False
print(is_suspicious_order(order_amount=650, num_items=12, new_customer=True))
# Output: TrueLet's walk through the snippet step by step:
- We define a function that takes three characteristics of the order: the amount, the number of items, and whether the customer is new.
- The conditions (
order_amount > 500,num_items > 100) are thresholds decided by a person, based on their intuition. - The program will never improve on its own: if fraudsters change tactics (for instance, placing many small orders), the rules will have to be rewritten by hand.
With Machine Learning, on the other hand, the approach is different. This snippet is only illustrative (we will learn to do it for real in modules 4 and 6), but it conveys the idea:
# Machine Learning approach: the rules are learned from examples
from sklearn.tree import DecisionTreeClassifier
# Historical data: [order_amount, num_items, new_customer (1/0)]
orders = [
[650, 12, 1],
[30, 5, 0],
[420, 80, 1],
[55, 10, 0],
]
# Known labels: 1 = it was fraud, 0 = it was legitimate
labels = [1, 0, 1, 0]
model = DecisionTreeClassifier() # we create an empty model
model.fit(orders, labels) # TRAINING: it learns from the examples
# PREDICTION on a new order the model has never seen
print(model.predict([[580, 15, 1]]))
# Output: [1] -> the model classifies it as suspiciousLet's analyze the key differences:
- We haven't written a single
ifwith thresholds: we gave it historical examples (orders) together with the correct answers (labels). - The call
model.fit(...)is the moment of training: the algorithm examines the examples and works out its own internal rules. - The call
model.predict(...)uses those learned rules to answer a new case. - If fraud changes its pattern, we don't rewrite code: we retrain the model with more recent data.
When does each approach pay off?
| Situation | Recommended approach |
|---|---|
| The rules are few, clear and stable (e.g., calculating the VAT on an order) | Traditional programming |
| The rules are too many or impossible to enumerate (e.g., recognizing a photo of a tomato) | Machine Learning |
| The environment is constantly changing (fraud, customer tastes) | Machine Learning |
| No historical data is available | Traditional programming (for now) |
Key concepts: data, model, training and prediction
These four terms will appear in every lesson of the course, so it's worth pinning them down right away:
- Training data: the set of historical examples the algorithm learns from. In the example above, the
orderslist with itslabels. Each row is an example (or instance) and each column a feature: order amount, number of items, and so on. - Model: the outcome of the learning. It is a mathematical representation of the patterns found in the data. You can think of it as "the learned rules", packaged in an object that knows how to answer questions.
- Training: the process by which the algorithm fits the model to the data. In scikit-learn it corresponds to the
fit()method. - Prediction (or inference): using the already-trained model to produce an answer for new data. In scikit-learn, the
predict()method.
A useful analogy: the training data is the syllabus, training is the studying, the model is the acquired knowledge, and prediction is the exam with new questions.
flowchart LR
D[Training data] --> E[Training]
E --> M[Model]
N[New data] --> M
M --> P[Predictions]
An important nuance we will develop in module 6: what makes a model valuable is not that it gets the data it already saw right, but that it generalizes well to new data. A model that memorizes the past but fails on the present is worthless to the business.
MercaFresh: the case study that runs through the course
MercaFresh is a fictional online supermarket operating in several cities. Every day it logs thousands of orders, and so it accumulates highly valuable data:
- Orders: date, products, quantities, amount, delivery time slot.
- Customers: tenure, purchase frequency, average basket size, city, device used.
- Products: category, price, stock, expiry date, active promotions.
With this data, MercaFresh's management asks itself business questions that are, precisely, Machine Learning problems. We will solve them over the course:
| Business question | ML problem | Where we'll work on it |
|---|---|---|
| How many kilos of fruit will we need on Saturday? | Demand forecasting | Modules 4 and 6 |
| Which customers are about to stop buying from us? | Customer churn | Modules 4, 6 and beyond |
| What natural groups of customers do we have? | Customer segmentation | Module 5 |
| Is this order anomalous or fraudulent? | Anomaly detection | Modules 5 and 9 |
For now you don't need to understand how they are solved; it's enough to see that behind every business question there is data, and behind the data, learnable patterns.
What Machine Learning is (and what it is not)
To close the lesson, let's draw the boundaries of the concept:
- ML is a branch of Artificial Intelligence focused on learning from data. Every ML solution is AI, but not all AI is ML (a system of hand-written expert rules is also AI).
- ML is not magic: if the data is poor, scarce or biased, the model will be poor. The saying "garbage in, garbage out" sums it up.
- ML is not just statistics, although it leans heavily on it (as we'll see in module 2); its emphasis is on automated prediction at scale.
- A model does not understand the business: it finds correlations in the data. Interpretation and decisions remain human.
Common Mistakes and Tips
- Confusing the model with the algorithm. The algorithm is the learning recipe (for example, "decision tree"); the model is the concrete result after training on your data. With the same algorithm and different data you get different models.
- Thinking that more hand-written rules will solve everything. If you catch yourself writing the twentieth
ifto cover a special case, the problem probably calls for a learning approach. - Believing that ML needs no human involvement. Someone has to define the task, choose the data, measure performance and decide what to do with the predictions.
- Expecting certainty. A model gives probable predictions, not truths. At MercaFresh, "this customer has an 80% probability of churning" is extremely valuable information, but it isn't a crystal ball.
- Tip: whenever you read about ML, always translate the terms into Mitchell's definition: what is the task T? What is the experience E? How is P measured? This habit will give you clarity throughout the course.
Exercises
Exercise 1
Identify the T, E and P components (per Mitchell's definition) for this MercaFresh case: "We want the system to automatically suggest products to each customer on the home page, using their purchase history, and to measure the percentage of suggestions the customer clicks on".
Exercise 2
Classify each of these MercaFresh problems as better suited to traditional programming or to Machine Learning, and justify your answer in one sentence:
- Calculating shipping costs based on order weight and postal code (fixed, published rates).
- Estimating how many delivery drivers will be needed over the next long holiday weekend.
- Validating that an email address entered at sign-up has a correct format.
- Detecting product reviews written by bots.
Exercise 3
In the lesson's code example, the orders list had 4 examples with 3 features each. Answer: (a) what does each row represent? (b) what does each column represent? (c) what role does the labels list play? (d) which line of code performs the training and which one the prediction?
Solutions
Solution 1
- T (task): suggest relevant products to each customer on the home page.
- E (experience): the purchase history of the customers (and of similar customers).
- P (performance): the percentage of clicks on the suggestions shown (CTR). If, as it sees more history, the suggestions get more clicks, the system is learning.
Solution 2
- Traditional: the rates are fixed, known and stable rules; there is nothing to learn.
- Machine Learning: it depends on complex historical patterns (holidays, weather, trends); it is a prediction problem based on data.
- Traditional: an email's format is validated with a clear, stable rule.
- Machine Learning: there are no fixed rules defining a bot review; the patterns change and must be learned from examples.
Solution 3
- (a) Each row is an example (one specific historical order).
- (b) Each column is a feature of the order: amount, number of items, and whether the customer was new.
- (c)
labelsholds the known correct answer for each example (1 = fraud, 0 = legitimate); it is what lets the algorithm learn the relationship between features and outcome. - (d) Training is done by
model.fit(orders, labels); prediction bymodel.predict([[580, 15, 1]]).
Conclusion
In this lesson we have defined Machine Learning both intuitively (learning from examples in order to generalize) and formally (Mitchell's T-E-P definition), and we have seen its essential difference from traditional programming: the rules are no longer written by hand, they emerge from the data during training. We have also pinned down the basic vocabulary — training data, model, training, prediction — and met MercaFresh, the online supermarket whose challenges (demand, churn, segmentation, anomalies) we will solve throughout the course. In the next lesson we will take a step back to understand where all of this comes from: we will walk through the history and evolution of Machine Learning, from Turing to today's era of generative AI.
Machine Learning Course
Module 1: Introduction to Machine Learning
- What is Machine Learning?
- History and evolution of Machine Learning
- Types of Machine Learning
- Applications of Machine Learning
- The Machine Learning project workflow
Module 2: Foundations of Statistics and Probability
- Basic statistics concepts
- Probability distributions
- Correlation and covariance
- Statistical inference
- Bayes' theorem
Module 3: Data Preprocessing
- Data cleaning
- Handling missing data
- Data transformation
- Encoding categorical variables
- Normalization and standardization
- Feature engineering
Module 4: Supervised Machine Learning Algorithms
- Linear regression
- Logistic regression
- Decision trees
- Support Vector Machines (SVM)
- K-Nearest Neighbors (K-NN)
- Naive Bayes
- Neural networks
Module 5: Unsupervised Machine Learning Algorithms
- Clustering: K-means
- Hierarchical clustering
- Principal Component Analysis (PCA)
- DBSCAN clustering
- Data visualization with t-SNE and UMAP
Module 6: Model Evaluation and Validation
- Data splitting: training, validation and test
- Evaluation metrics
- Cross-validation
- ROC curve and AUC
- Overfitting and underfitting
Module 7: Advanced Techniques and Optimization
- Regularization: Ridge, Lasso and Elastic Net
- Ensemble Learning
- Gradient Boosting
- Deep neural networks (Deep Learning)
- Hyperparameter optimization
Module 8: Model Implementation and Deployment
- Popular frameworks and libraries
- Deploying models to production
- Model maintenance and monitoring
- Ethical and privacy considerations
Module 9: Hands-On Projects
- Project 1: Housing price prediction
- Project 2: Image classification
- Project 3: Sentiment analysis on social media
- Project 4: Fraud detection
- Project 5: Customer segmentation
