Three modules spent preparing MercaFresh's data, and at last the payoff arrives: training a model. We start with linear regression, the oldest and simplest supervised algorithm — and precisely for that reason the best place to understand, with no black boxes, what it means for a machine to learn. In this lesson you will build the geometric intuition (a straight line running through a cloud of points), see which error it minimizes and how it minimizes it (least squares and gradient descent), learn to read the coefficients as business statements, and train your first real model: predicting the monthly spend of MercaFresh's customers.
Contents
- The geometric intuition: the best-fitting line
- The equation: simple and multiple regression
- Least squares: what "best fit" means
- Gradient descent: how the minimum is found
- Implementation with scikit-learn: monthly spend at MercaFresh
- Interpreting the coefficients
- Model assumptions in practice
- Limitations and when to be suspicious
The geometric intuition: the best-fitting line
Picture a scatter plot of MercaFresh's customers: on the X axis, their orders per month; on the Y axis, their monthly spend in euros. The cloud of points rises to the right — more orders, more spend — but it doesn't form a perfect line: there are customers with many small orders and customers with a few enormous ones.
Linear regression answers a very specific question: of all the possible lines that could run through that cloud, which one passes "closest" to all the points at once? That line is the model. Once you have it, predicting is trivial: for a new customer with 6 orders/month, you go straight up from x=6 to the line and read off the estimated spend.
Notice how this matches Mitchell's definition from module 1 exactly:
- Task T: predict a customer's monthly spend.
- Experience E: the historical customers with their known spend.
- Measure P: how far off the predictions are (we formalize it two sections from now).
"Learning" is, quite literally, moving the line until measure P is as good as possible on experience E.
The equation: simple and multiple regression
Simple regression (one variable)
The line you know from school:
$$\hat{y} = w_0 + w_1 x$$
| Symbol | Name | Reading at MercaFresh |
|---|---|---|
| $\hat{y}$ | Prediction | Estimated monthly spend (the hat means "estimated", not actual) |
| $x$ | Feature | The customer's orders per month |
| $w_1$ | Slope (coefficient) | Extra euros of spend for each additional monthly order |
| $w_0$ | Intercept (bias) | Estimated baseline spend when x = 0 |
The parameters the model learns are just two numbers: $w_0$ and $w_1$. Nothing more. All of the model's "knowledge" fits inside them.
Multiple regression (several variables)
In practice you never predict with a single feature. A customer's spend depends on their orders, their recency, their plan... The equation generalizes by adding one term per feature:
$$\hat{y} = w_0 + w_1 x_1 + w_2 x_2 + \dots + w_n x_n$$
Geometrically it is no longer a line but a hyperplane: with 2 features it's a plane in 3D; with the features in the MercaFresh dataset, a hyperplane in a space we can't draw but that works by exactly the same logic. Each $w_i$ is still "how much the prediction changes when $x_i$ goes up by one unit, holding everything else constant" — that final qualifier will be key when we interpret.
Least squares: what "best fit" means
"The line that passes closest to all the points" sounds nice but is ambiguous. We have to define close with a formula. The classic definition:
- For each customer $i$, compute the residual: $e_i = y_i - \hat{y}_i$ (actual spend minus predicted spend). It is the vertical distance from the point to the line.
- Square each residual: $e_i^2$. That way positive and negative errors don't cancel out, and large errors weigh disproportionately more (being off by €20 counts 4 times as much as being off by €10).
- Average: that is the mean squared error (MSE), the cost function.
$$MSE(w) = \frac{1}{n} \sum_{i=1}^{n} (y_i - \hat{y}_i)^2$$
The least squares line is the one that makes this number as small as possible. Now "learning" has an exact operational definition: find the values of $w_0, \dots, w_n$ that minimize the MSE on the training data. In module 6 (06-02) we'll cover MSE and its relatives in detail as evaluation metrics; here its role as a cost function is all we need — the objective the algorithm chases.
One detail worth knowing: for linear regression there is a closed-form formula (the "normal equation") that yields the optimal weights in one stroke, and that is what scikit-learn uses under the hood. But the general method is worth understanding, because almost no other model in the course has a closed-form solution.
Gradient descent: how the minimum is found
Gradient descent is the universal iterative method for minimizing cost functions, and it will resurface in logistic regression (04-02) and in neural networks (04-07). The intuition:
Think of the MSE as a landscape: each point on the terrain is a combination of weights $(w_0, w_1)$, and the altitude at that point is the error that combination produces. For linear regression, that landscape is a bowl-shaped valley with a single bottom. You are on the hillside, in fog, and you want to reach the bottom:
- Start anywhere: random weights.
- Feel the slope under your feet: the gradient points in the direction where the error climbs fastest.
- Take a step in the opposite direction, with a size proportional to the learning rate.
- Repeat until the ground is flat: you have reached the minimum.
flowchart TD
A["Random initial weights<br/>w0, w1"] --> B["Compute predictions<br/>and the current MSE"]
B --> C["Compute the gradient:<br/>which direction the error grows in"]
C --> D["Update weights:<br/>step in the opposite direction<br/>w = w - rate * gradient"]
D --> E{"Has the error<br/>almost stopped dropping?"}
E -- "No" --> B
E -- "Yes" --> F["Final weights:<br/>trained model"]
The learning rate is the central trade-off: steps that are too small take forever; steps that are too large leap from one hillside to the opposite one without ever descending. It's the first hyperparameter you meet in the course — a number you set before training, not something the model learns (its systematic optimization arrives in 07-05).
Implementation with scikit-learn: monthly spend at MercaFresh
Business goal: MercaFresh wants to estimate each customer's future monthly spend to size campaigns and stock. It's a regression: the label is a continuous number. We use features derived from the RFM table we built in 03-06:
import numpy as np
import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_squared_error
rng = np.random.default_rng(42)
n = 500
# Simulate the MercaFresh customer table (features from 03-06)
customers = pd.DataFrame({
"orders_per_month": rng.gamma(3, 1.2, n).round(1),
"avg_order_spend": rng.gamma(9, 5, n).round(2),
"recency_days": rng.integers(1, 120, n),
"months_tenure": rng.integers(3, 60, n),
})
# Actual monthly spend: orders x basket, penalized by inactivity, plus noise
customers["monthly_spend"] = (
customers["orders_per_month"] * customers["avg_order_spend"]
- 0.8 * customers["recency_days"]
+ rng.normal(0, 15, n)
).clip(lower=0).round(2)
X = customers[["orders_per_month", "avg_order_spend",
"recency_days", "months_tenure"]]
y = customers["monthly_spend"]
# Train/test split: train on 80% and check on the 20%
# the model never saw (the full why comes in 06-01)
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42)
model = LinearRegression()
model.fit(X_train, y_train) # this is where the "learning" happens
predictions = model.predict(X_test)
mse = mean_squared_error(y_test, predictions)
print(f"MSE on test: {mse:.1f}")
print(f"Typical error: ~{np.sqrt(mse):.1f} EUR")A step-by-step walkthrough for anyone training for the first time:
train_test_split: sets aside 20% of customers the model won't see during training, in order to measure the error on new data — the same fit-on-train discipline we applied to scaling in 03-05. Module 6 develops this idea in depth.model.fit(X_train, y_train): the line where everything happens. scikit-learn computes the weights that minimize the MSE on the training set. It's the samefitinterface you used withSimpleImputerorStandardScaler: adjusting parameters from data.model.predict(X_test): applies the equation $\hat{y} = w_0 + w_1 x_1 + \dots$ to each test customer.mean_squared_error: the MSE from section 3, now acting as the judge. Its square root brings the error back to euros, the original unit: "we're typically off by about €15".
A practical note that picks up 03-05: LinearRegression works without scaling the features (the exact solution doesn't depend on scale), but if you trained with gradient descent or added regularization (07-01), scaling would be essential. Scaling by default, as our module 3 preprocessor does, never hurts.
Interpreting the coefficients
The great virtue of linear regression is that the trained model can be read:
coefs = pd.Series(model.coef_, index=X.columns).round(3)
print(coefs)
print(f"Intercept: {model.intercept_:.2f}")Typical output (your numbers will vary slightly):
| Feature | Coefficient | Business reading |
|---|---|---|
orders_per_month |
+44.1 | Each additional monthly order adds ~€44 of estimated spend, all else being equal |
avg_order_spend |
+3.6 | Each extra euro of average basket adds ~€3.6 per month |
recency_days |
−0.81 | Each day without buying subtracts ~€0.81 of expected monthly spend |
months_tenure |
~0.0 | Tenure adds almost nothing once the others are known |
Three professional caveats:
- "All else being equal" is not a decorative tag: the coefficient of
recency_dayscompares customers with the same orders and basket but different recency. If the features are correlated with each other (and the RFM ones are, as we saw in the heatmap in 02-03), the individual coefficients become unstable and must be read with caution. - Units matter: a coefficient of 44 on "orders" and one of 3.6 on "euros of basket" are not directly comparable — they measure things in different units. To compare importances, train on standardized features (03-05).
- Correlation is not causation (02-03): the model says high recency goes along with low spend, not that forcing a purchase will rejuvenate the customer.
Model assumptions in practice
Linear regression assumes things about the data. You don't need the full statistical formality; you do need to know how to check the two that hurt most in practice:
- Linearity: the real relationship between features and target should be approximately a sum of linear effects. Check: a scatter plot of each feature against the target. If you see a curve (diminishing returns of spend with frequency, for instance), the line will fall short — the log transformations from 03-03 or the interactions from 03-06 can linearize the relationship.
- Residuals with no structure: the model's errors should look like noise — centered on zero, patternless, with stable variance. The check is a plot of residuals against predictions:
import matplotlib.pyplot as plt
residuals = y_test - predictions
plt.scatter(predictions, residuals, alpha=0.5)
plt.axhline(0, color="red", linestyle="--")
plt.xlabel("Predicted spend (EUR)")
plt.ylabel("Residual (actual - predicted)")
plt.show()How to read it: a horizontal, shapeless cloud around zero is a good sign. A U shape says "there's nonlinearity I'm not capturing"; a funnel (residuals growing with the prediction) says "I predict far worse for big customers" — the histogram and boxplot from 02-01, applied to the residuals, complete the diagnosis.
Limitations and when to be suspicious
| Limitation | Why it happens | Mitigation |
|---|---|---|
| Doesn't capture nonlinear relationships | The model is a hyperplane | Transform features (03-03), interactions (03-06), or nonlinear models (04-03 onward) |
| Sensitive to outliers | Squaring the residual makes an extreme point pull on the line with enormous force | The outlier cleaning from 03-01 before training |
| Unstable coefficients with correlated features | Several features "compete" to explain the same thing | The correlation filter from 03-06; Ridge/Lasso regularization (07-01) extends linear regression by penalizing large weights and stabilizes exactly this problem |
| Extrapolates cheerfully | The line stays straight beyond the training range | Distrust predictions for customers very different from those seen |
The first limitation is the most important conceptually: there are problems where no straight line works. The most notorious one: predicting a category (will this customer churn, yes or no?). There linear regression fails by design — and that failure is exactly where the next lesson begins.
Common Mistakes and Tips
- Using linear regression to classify (churn yes/no coded as 0/1 and off you go). It produces absurd predictions like "probability 1.4". The right tool arrives in 04-02.
- Interpreting coefficients of correlated features as isolated truths. With
orders_per_monthandtotal_spendboth in the model, their coefficients split the effect arbitrarily. Review the correlation heatmap (02-03) before reading coefficients. - Forgetting to look at the residuals. An acceptable MSE can hide a model that fails systematically on one segment (for instance, all the premium customers). The residual plot takes three lines and reveals it.
- Fitting the model with uncleaned outliers. A single corporate customer spending €10,000/month shifts the line for everyone. Revisit the IQR/z-score treatment from 03-01.
- Tip: always start with linear regression even if you plan to use complex models. It's the baseline: cheap, interpretable, and if a sophisticated model doesn't clearly beat it, it isn't paying for its complexity.
Exercises
Exercise 1. With the model trained in the lesson, compute by hand (without predict) the estimated monthly spend of a customer with 4 orders/month, an average basket of €30, recency of 10 days and 24 months of tenure. Use model.intercept_ and model.coef_. Then check with model.predict.
Exercise 2. Add a new feature to the dataset, interaction = orders_per_month * avg_order_spend (the hand-picked interaction, as in 03-06), retrain and compare the MSE on test. Why does this interaction help so much on this particular dataset?
Exercise 3. Artificially introduce an outlier: a customer with monthly_spend = 50000 in the training set. Retrain and compare the coefficients with the originals. Which coefficient changes most, and why?
Solutions
Exercise 1
customer = np.array([4, 30, 10, 24])
manual_prediction = model.intercept_ + np.dot(model.coef_, customer)
print(f"By hand: {manual_prediction:.2f} EUR")
customer_df = pd.DataFrame([[4, 30, 10, 24]], columns=X.columns)
print(f"With predict: {model.predict(customer_df)[0]:.2f} EUR")Both values match exactly: predict does nothing beyond the hyperplane equation. Internalizing this demystifies the model — all its "intelligence" is five numbers.
Exercise 2
X2 = X.copy()
X2["interaction"] = X2["orders_per_month"] * X2["avg_order_spend"]
X2_train, X2_test, y_train, y_test = train_test_split(
X2, y, test_size=0.2, random_state=42)
model2 = LinearRegression().fit(X2_train, y_train)
mse2 = mean_squared_error(y_test, model2.predict(X2_test))
print(f"MSE without interaction: {mse:.1f} | with interaction: {mse2:.1f}")The MSE drops dramatically because the actual spend was generated as orders × basket: a multiplicative relationship that no hyperplane of standalone features can express, but which becomes linear the moment the interaction exists as a column. Moral: feature engineering (03-06) can turn a nonlinear problem into a linear one.
Exercise 3
X_out, y_out = X_train.copy(), y_train.copy()
y_out.iloc[0] = 50000
model3 = LinearRegression().fit(X_out, y_out)
print(pd.DataFrame({"original": model.coef_, "with_outlier": model3.coef_},
index=X.columns).round(2))The coefficients of the features where that customer has high values shoot up: the squared €50,000 residual dominates the cost function, and the model finds it worthwhile to misfit the other 399 customers just to get slightly closer to that one. It's outlier sensitivity in action — and the definitive argument for the cleaning in 03-01.
Conclusion
You have now trained your first real model: you know that linear regression seeks the hyperplane that minimizes the MSE, that gradient descent is the general method for finding that minimum (and will come back again and again throughout the course), that coefficients read as business statements — with the caveats about correlated features —, and that residuals are the telltale of broken assumptions. You've also seen its limits: nonlinearity, outliers, and above all its inability to answer yes-or-no questions.
And that last limitation is the doorway to the next lesson: MercaFresh doesn't just want to know how much a customer will spend, but whether they are going to churn. For that we need to bend the line into an S-shaped curve that speaks the language of probabilities: logistic regression.
Machine Learning Course
Module 1: Introduction to Machine Learning
- What is Machine Learning?
- History and evolution of Machine Learning
- Types of Machine Learning
- Applications of Machine Learning
- The Machine Learning project workflow
Module 2: Foundations of Statistics and Probability
- Basic statistics concepts
- Probability distributions
- Correlation and covariance
- Statistical inference
- Bayes' theorem
Module 3: Data Preprocessing
- Data cleaning
- Handling missing data
- Data transformation
- Encoding categorical variables
- Normalization and standardization
- Feature engineering
Module 4: Supervised Machine Learning Algorithms
- Linear regression
- Logistic regression
- Decision trees
- Support Vector Machines (SVM)
- K-Nearest Neighbors (K-NN)
- Naive Bayes
- Neural networks
Module 5: Unsupervised Machine Learning Algorithms
- Clustering: K-means
- Hierarchical clustering
- Principal Component Analysis (PCA)
- DBSCAN clustering
- Data visualization with t-SNE and UMAP
Module 6: Model Evaluation and Validation
- Data splitting: training, validation and test
- Evaluation metrics
- Cross-validation
- ROC curve and AUC
- Overfitting and underfitting
Module 7: Advanced Techniques and Optimization
- Regularization: Ridge, Lasso and Elastic Net
- Ensemble Learning
- Gradient Boosting
- Deep neural networks (Deep Learning)
- Hyperparameter optimization
Module 8: Model Implementation and Deployment
- Popular frameworks and libraries
- Deploying models to production
- Model maintenance and monitoring
- Ethical and privacy considerations
Module 9: Hands-On Projects
- Project 1: Housing price prediction
- Project 2: Image classification
- Project 3: Sentiment analysis on social media
- Project 4: Fraud detection
- Project 5: Customer segmentation
