Linear regression closed the previous lesson by running into a wall: it doesn't know how to answer yes-or-no questions. And MercaFresh's headline question is exactly of that kind — will this customer churn within the next 90 days? In this lesson we transform the straight line into an S-shaped curve that produces probabilities, understand why that transformation is necessary and what its coefficients mean in terms of odds, and train the course's first classifier on the churn dataset we built in module 3, preprocessor included. Despite its name, logistic regression is a classification algorithm — probably the most widely used in production anywhere in the world.
Contents
- Why linear regression can't classify
- The sigmoid function: from numbers to probabilities
- Probability and the decision threshold
- Log-odds and coefficients: interpretation via odds ratios
- The cost function: log-loss
- Implementation: MercaFresh churn with the module 3 preprocessor
- Multiclass classification: one-vs-rest and softmax
Why linear regression can't classify
First temptation: code churn as 0/1 and fit a linear regression. It sounds reasonable — the prediction would be "something like the probability of churn". It fails for three reasons:
- Out-of-range predictions: a hyperplane is unbounded. For a very inactive customer it will predict 1.4, and for a very loyal one, −0.3. A 140% probability? No interpretation can rescue that.
- Absurd sensitivity to extreme cases: remember exercise 3 in 04-01 — far-away points pull on the line. A blatantly obvious churn case (recency of 300 days) shifts the line and worsens the classification of the borderline cases, which are precisely the ones that matter.
- Squared error punishes the wrong thing: penalizing the model with MSE for predicting 1.4 on a customer who did in fact churn (label 1) is punishing it for being too right.
The solution is not to abandon the linear combination $z = w_0 + w_1 x_1 + \dots + w_n x_n$ — which remains the heart of the model — but to pass it through a function that squashes it into the interval (0, 1).
The sigmoid function: from numbers to probabilities
That function is the sigmoid (or logistic function, hence the algorithm's name):
$$\sigma(z) = \frac{1}{1 + e^{-z}}$$
Its properties are exactly the ones we need:
| Input $z$ | Output $\sigma(z)$ | Reading |
|---|---|---|
| $z \to -\infty$ | → 0 | Certainty of "no churn" |
| $z = -2$ | 0.12 | Probably stays |
| $z = 0$ | 0.50 | Maximum uncertainty |
| $z = +2$ | 0.88 | Probably churns |
| $z \to +\infty$ | → 1 | Certainty of "churn" |
import numpy as np
import matplotlib.pyplot as plt
z = np.linspace(-6, 6, 200)
sigma = 1 / (1 + np.exp(-z))
plt.plot(z, sigma)
plt.axhline(0.5, color="gray", linestyle="--")
plt.axvline(0, color="gray", linestyle="--")
plt.xlabel("z = w0 + w1*x1 + ... (linear score)")
plt.ylabel("P(churn = 1)")
plt.title("The sigmoid function")
plt.show()The curve is S-shaped: flat at the extremes (where the model is already sure, extra evidence barely changes anything) and steep in the middle (where every tenth of $z$ moves the probability a lot). The full model is:
$$P(\text{churn} = 1 \mid \mathbf{x}) = \sigma(w_0 + w_1 x_1 + \dots + w_n x_n)$$
Look at the notation: it's a conditional probability, the same creature we handled in Bayes' theorem (02-05). Logistic regression estimates $P(\text{class} \mid \text{data})$ directly — what we called the posterior there — without going through priors and likelihoods: it learns the conditional probability in one go, by fitting weights. In 04-06 we'll see the opposite approach (Naive Bayes), which does build the posterior piece by piece as in 02-05.
Probability and the decision threshold
The model returns probabilities; the business needs decisions. The bridge is the threshold: by default, if $P(\text{churn}) \geq 0.5$ churn is predicted, otherwise staying. Geometrically, $\sigma(z) = 0.5$ happens when $z = 0$: the equation $w_0 + w_1 x_1 + \dots = 0$ defines a linear decision boundary — a hyperplane, as in 04-01, but now separating classes instead of fitting values.
The important part: the 0.5 threshold is a convention, not a law. For MercaFresh, letting a valuable customer slip away (false negative) costs far more than sending an unnecessary coupon (false positive) — the same asymmetric-cost dilemma as the fraud detector in 02-05. Lowering the threshold to 0.3, the retention campaign captures more at-risk customers in exchange for more wasted coupons:
probs = model.predict_proba(X_test)[:, 1] # probability of churn
standard_decision = probs >= 0.5
cautious_decision = probs >= 0.3 # when in doubt, retainHow to choose the optimal threshold and how to measure that trade-off rigorously is the subject of the ROC curve (06-04) and the classification metrics (06-02); here it's enough to know that the probability is the rich output and the decision is a layer on top, tunable to the business cost.
Log-odds and coefficients: interpretation via odds ratios
In linear regression, "raising $x_1$ by one unit adds $w_1$ euros". And here? The sigmoid complicates a direct reading in probabilities, but there is an elegant reformulation. Solving for $z$:
$$\ln\left(\frac{P}{1-P}\right) = w_0 + w_1 x_1 + \dots + w_n x_n$$
The ratio $\frac{P}{1-P}$ is the odds: with P = 0.75, the odds are 3 — "3 to 1 in favor of churn". The model is linear in the logarithm of the odds (log-odds). That's where the practical interpretation comes from:
Raising $x_i$ by one unit multiplies the odds by $e^{w_i}$ (the odds ratio), holding everything else constant.
| Coefficient $w_i$ | Odds ratio $e^{w_i}$ | MercaFresh reading (example) |
|---|---|---|
| +0.9 (scaled recency) | 2.46 | Each extra unit of recency multiplies the odds of churn by ~2.5 |
| 0.0 | 1.00 | The feature adds nothing |
| −0.7 (scaled frequency) | 0.50 | Each extra unit of frequency halves the odds of churn |
Two caveats inherited from 04-01: scaled features (our preprocessor uses RobustScaler) make "one unit" a robust IQR-style quantity, not a natural unit; and with correlated features the individual coefficients share out the effect unstably — the reading is indicative, not a causal verdict (02-03).
The cost function: log-loss
We can't train by minimizing the MSE: combined with the sigmoid it produces an error landscape full of false valleys where gradient descent gets stuck. The natural cost function for probabilities is log-loss (or binary cross-entropy). Its logic, case by case:
- If the true label is 1 (churn), the cost is $-\ln(p)$: predicting $p = 0.9$ costs little (0.105); predicting $p = 0.01$ costs a fortune (4.6).
- If the true label is 0, the cost is $-\ln(1-p)$: the mirror image.
$$\text{LogLoss} = -\frac{1}{n}\sum_{i=1}^{n} \left[ y_i \ln(p_i) + (1 - y_i)\ln(1 - p_i) \right]$$
The key property: the cost of a confident mistake tends to infinity. Saying "97% sure they'll stay" about a customer who churns is extremely expensive; saying "60%" about the same customer, much less so. Log-loss doesn't just reward being right: it rewards honest probabilities. With it, the error landscape becomes a bowl with a single minimum again, and the gradient descent from 04-01 finds it without a closed-form solution (there is no normal equation here: training is always iterative).
Implementation: MercaFresh churn with the module 3 preprocessor
Time to cash in the module 3 investment: the ColumnTransformer from 03-06 is chained with the classifier in a single Pipeline. One object makes the whole journey raw data → prediction:
import numpy as np
import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline
from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import (OneHotEncoder, OrdinalEncoder,
PowerTransformer, RobustScaler)
from sklearn.linear_model import LogisticRegression
# --- The preprocessor built in 03-06 (abridged) ---
numeric_branch = Pipeline([
("impute", SimpleImputer(strategy="median", add_indicator=True)),
("unskew", PowerTransformer(method="yeo-johnson")),
("scale", RobustScaler()),
])
preprocessor = ColumnTransformer([
("num", numeric_branch,
["age", "satisfaction", "recency_days", "orders_per_month",
"avg_order_spend", "inactivity_ratio", "trend"]),
("cat_nominal", OneHotEncoder(sparse_output=False, handle_unknown="ignore"),
["city"]),
("cat_ordinal", OrdinalEncoder(categories=[["basic", "standard", "premium"]]),
["plan"]),
])
# --- Simulated churn dataset with the 03-06 structure ---
rng = np.random.default_rng(7)
n = 800
df = pd.DataFrame({
"age": rng.integers(18, 75, n),
"satisfaction": np.where(rng.random(n) < 0.1, np.nan, rng.integers(1, 11, n)),
"recency_days": rng.integers(1, 180, n),
"orders_per_month": rng.gamma(3, 1.2, n).round(1),
"avg_order_spend": rng.gamma(9, 5, n).round(2),
"city": rng.choice(["Valencia", "Madrid", "Barcelona"], n),
"plan": rng.choice(["basic", "standard", "premium"], n, p=[0.5, 0.3, 0.2]),
})
df["inactivity_ratio"] = (df["recency_days"] / 365).round(3)
df["trend"] = rng.uniform(0.1, 2.5, n).round(2)
# Churn depends mostly on high recency and a declining trend
logits = 0.03 * df["recency_days"] - 1.5 * df["trend"] - 0.4 * df["orders_per_month"] + 1.0
df["churn"] = (rng.random(n) < 1 / (1 + np.exp(-logits))).astype(int)
X = df.drop(columns="churn")
y = df["churn"]
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, stratify=y, random_state=42)
# --- Full pipeline: preprocess + classify ---
clf = Pipeline([
("prep", preprocessor),
("model", LogisticRegression(max_iter=1000)),
])
clf.fit(X_train, y_train) # fits the preprocessor AND the model, on train only
accuracy = clf.score(X_test, y_test) # % of correct answers, as a first reference
print(f"Accuracy on test: {accuracy:.2%}")
# Probabilities for the retention campaign
at_risk = clf.predict_proba(X_test)[:, 1]
print(X_test.assign(p_churn=at_risk.round(2))
.sort_values("p_churn", ascending=False)
.head(3)[["recency_days", "trend", "p_churn"]])Points that deserve a slow read:
stratify=y: keeps the churn proportion equal in train and test — important when one class is a minority (more in 06-01).- The
Pipelinesolves data leakage: when you callclf.fit(X_train, ...), the preprocessor learns medians and quartiles from train only; when predicting on test, it applies those same parameters. Exactly the discipline of 03-02 and 03-05, now automated. predict_proba: the real business output. The list sorted byp_churnis the retention team's call list — it connects with thevalue_at_riskfeature from the 03-06 exercise.scorereturns the accuracy (fraction of correct answers): useful as a first reference, but a misleading metric with imbalanced classes; its critique and its alternatives arrive in 06-02.max_iter=1000: training is iterative (gradient descent and variants); with badly scaled features or too few iterations it doesn't converge — another reason why scaling (03-05) matters here, unlike in 04-01.
As with linear regression, there is a regularized version that penalizes large weights — in fact LogisticRegression ships with it on by default (the C parameter); understanding it in depth is the subject of 07-01.
Multiclass classification: one-vs-rest and softmax
What if there are more than two classes? MercaFresh might want to predict a customer's favorite category (fresh / pantry / cleaning). Two strategies:
- One-vs-rest (OvR): train one binary classifier per class ("fresh versus the rest", "pantry versus the rest"...) and pick the class whose model gives the highest probability. Simple and parallelizable.
- Softmax (multinomial logistic regression): generalizes the sigmoid to K classes at once — computes one score $z_k$ per class and turns them into K probabilities that sum to 1. It's what
LogisticRegressionuses by default for multiclass problems, and it will reappear as the output layer of neural networks (04-07).
# Nothing changes in the code: sklearn detects the classes automatically
# model = LogisticRegression(max_iter=1000) # softmax if y has 3+ classes
# model.predict_proba(X) -> one probability column per classFor the rest of the course the binary case — the churn case — is all we need.
Common Mistakes and Tips
- Reading
predictwhen the business needspredict_proba. The hard label throws away information: a customer with p = 0.51 and another with p = 0.99 are both "churn", but they don't deserve the same phone call. Work with probabilities and decide the threshold at the end. - Interpreting coefficients as probabilities. A coefficient of 0.9 does not mean "+90% probability": it means odds multiplied by e^0.9 ≈ 2.5. The effect on probability depends on the starting point (the S is flat at the extremes).
- Forgetting the scaling. Unlike
LinearRegression, training here is iterative and regularized by default: without scaling, it converges poorly and the penalty punishes large-scale features arbitrarily. OurPipelinemakes it impossible to forget — always use it. - Trusting a high accuracy with imbalanced classes. With 90% loyal customers, a model that always says "stays" is right 90% of the time... and useless. The complete solution, in 06-02.
- Tip: logistic regression is the mandatory baseline for any classification, just as linear regression was for regression. Interpretable, fast, with well-calibrated probabilities. The models in the coming lessons have to earn their keep against it.
Exercises
Exercise 1. Without running any code: a churn model has $w_0 = -1$ and a single coefficient $w_1 = 0.5$ on scaled_recency. Compute $P(\text{churn})$ for customers with scaled recency 0, 2 and 6. From what scaled recency onward does the model predict churn with the 0.5 threshold?
Exercise 2. With the lesson's Pipeline trained, extract the model's coefficients (clf.named_steps["model"].coef_) together with the names of the transformed features (clf.named_steps["prep"].get_feature_names_out()). Sort by absolute value and translate the two most influential features into odds ratios with one business sentence each.
Exercise 3. The retention team can only call 10% of customers. Write the code that selects, from the test set, the 10% with the highest churn probability, and explain why this beats using predict with a 0.5 threshold.
Solutions
Exercise 1
- Recency 0: $z = -1$, $\sigma(-1) = 1/(1+e^{1}) \approx 0.27$.
- Recency 2: $z = 0$, $\sigma(0) = 0.50$ — right on the boundary.
- Recency 6: $z = 2$, $\sigma(2) \approx 0.88$.
The boundary sits where $z = 0$: $-1 + 0.5 \cdot x = 0 \Rightarrow x = 2$. Above a scaled recency of 2, the model predicts churn. Notice that the decision boundary is a point (in 1D), a line (in 2D), a hyperplane in general: logistic regression is still a linear classifier.
Exercise 2
names = clf.named_steps["prep"].get_feature_names_out()
coefs = clf.named_steps["model"].coef_[0]
table = (pd.DataFrame({"feature": names, "coef": coefs})
.assign(odds_ratio=lambda t: np.exp(t["coef"]).round(2))
.reindex(pd.Series(coefs).abs().sort_values(ascending=False).index))
print(table.head())With the simulated data, the top spots go to num__recency_days (positive coefficient: each extra robust unit of recency multiplies the odds of churn by its odds ratio — the customer who goes quiet is the one who leaves) and num__trend (negative coefficient, odds ratio < 1: a rising activity trend divides the odds of churning). Both sentences match the business intuition that motivated those features in 03-06 — a good sign: the model has learned what we expected it to learn.
Exercise 3
probs = clf.predict_proba(X_test)[:, 1]
cutoff = np.quantile(probs, 0.90) # 90th percentile (02-01)
call_list = X_test[probs >= cutoff]
print(f"Effective threshold: {cutoff:.2f} | customers to call: {len(call_list)}")With a 0.5 threshold the list would have an arbitrary size — maybe 30% of customers (impossible to call), maybe 2% (capacity wasted). Sorting by probability and cutting by capacity uses exactly the resource available and concentrates it on the highest-risk cases: the threshold is dictated by the business, not by mathematical convention.
Conclusion
You have turned the line from 04-01 into a classifier: the sigmoid squashes the linear score into a probability, log-loss punishes confident mistakes and rewards probabilistic honesty, coefficients read as odds ratios, and the decision threshold is a business lever, not a sacred constant. You've also debuted the definitive professional pattern: Pipeline(preprocessor, model), where the module 3 work and the classifier travel together with no risk of leakage.
But logistic regression has the same soul as linear regression: its decision boundary is a hyperplane. If churning customers don't separate from loyal ones with one straight cut — "churn if recency is high and the plan is basic, but also if spend is high and the trend has collapsed" — we need a model that thinks in rules rather than weights. That is exactly a decision tree, and it's the next lesson.
Machine Learning Course
Module 1: Introduction to Machine Learning
- What is Machine Learning?
- History and evolution of Machine Learning
- Types of Machine Learning
- Applications of Machine Learning
- The Machine Learning project workflow
Module 2: Foundations of Statistics and Probability
- Basic statistics concepts
- Probability distributions
- Correlation and covariance
- Statistical inference
- Bayes' theorem
Module 3: Data Preprocessing
- Data cleaning
- Handling missing data
- Data transformation
- Encoding categorical variables
- Normalization and standardization
- Feature engineering
Module 4: Supervised Machine Learning Algorithms
- Linear regression
- Logistic regression
- Decision trees
- Support Vector Machines (SVM)
- K-Nearest Neighbors (K-NN)
- Naive Bayes
- Neural networks
Module 5: Unsupervised Machine Learning Algorithms
- Clustering: K-means
- Hierarchical clustering
- Principal Component Analysis (PCA)
- DBSCAN clustering
- Data visualization with t-SNE and UMAP
Module 6: Model Evaluation and Validation
- Data splitting: training, validation and test
- Evaluation metrics
- Cross-validation
- ROC curve and AUC
- Overfitting and underfitting
Module 7: Advanced Techniques and Optimization
- Regularization: Ridge, Lasso and Elastic Net
- Ensemble Learning
- Gradient Boosting
- Deep neural networks (Deep Learning)
- Hyperparameter optimization
Module 8: Model Implementation and Deployment
- Popular frameworks and libraries
- Deploying models to production
- Model maintenance and monitoring
- Ethical and privacy considerations
Module 9: Hands-On Projects
- Project 1: Housing price prediction
- Project 2: Image classification
- Project 3: Sentiment analysis on social media
- Project 4: Fraud detection
- Project 5: Customer segmentation
