Fourth project and a new enemy: extreme class imbalance. You are going to build a fraudulent card transaction detector where only 1% of the cases are fraud — the scenario where accuracy deceives, thresholds are decided in euros and the metrics from module 6 earn their keep. The problem will feel familiar twice over: it is the spam detector from Bayes' theorem in 02-05 with money on the line, and the supervised cousin of the order anomalies you hunted with DBSCAN in 05-04 — except here we do have historical fraud labels, so we can train classifiers. We will use a synthetic dataset generated right in the code (real fraud datasets are confidential by nature; the classic public one on Kaggle comes anonymized through PCA), which also gives us full control over the experiment.
Contents
- Problem definition, costs and target metric
- Generating the synthetic dataset
- EDA with imbalanced classes
- Why accuracy lies at 99%
- Strategies against imbalance
- Modeling and comparison with CV
- The precision-recall curve and AUC-PR
- The optimal threshold, in euros
- Final evaluation on the test set
- Conclusions and the link to production
Problem definition, costs and target metric
- Problem: binary classification — is this transaction fraudulent (1) or legitimate (0)?
- The asymmetry that governs everything: the two errors do not cost the same. An undetected fraud (false negative) costs the defrauded amount plus handling, say €150 on average; blocking a legitimate transaction (false positive) costs the manual review plus customer friction, say €5. This cost matrix, as you learned in 06-04 "ROC curve and AUC", will decide the threshold.
- Metrics: precision, recall and F1 for the fraud class (06-02 "Evaluation metrics"), the precision-recall curve with its AUC-PR as the model-comparison metric, and the total cost in euros as the final business metric. Accuracy is banned, and you are about to see why.
| Predicted: legitimate | Predicted: fraud | |
|---|---|---|
| Actual: legitimate | €0 | €5 (review + friction) |
| Actual: fraud | €150 (loss) | €0 (fraud prevented) |
Generating the synthetic dataset
make_classification generates a classification problem with controlled structure; with weights=[0.99, 0.01] we impose the imbalance. Then we dress the abstract features with realistic names and scales so we can reason like analysts:
import numpy as np
import pandas as pd
from sklearn.datasets import make_classification
X_raw, y = make_classification(
n_samples=50_000, n_features=8, n_informative=5, n_redundant=1,
weights=[0.99, 0.01], # 1% fraud
class_sep=1.0, flip_y=0.005, # some label noise: realism
random_state=42,
)
columns = [
"amount_eur", # transaction amount
"hour_of_day", # time of day (fraud prefers the small hours)
"dist_home_km", # distance from the merchant to the cardholder's home
"merchant_freq_month", # times the holder buys at that merchant per month
"txns_last_hour", # card transactions in the last hour
"card_age_months", # months since the card was issued
"ratio_avg_amount", # amount / holder's historical average spend
"online_payment", # channel signal (higher => more online)
]
df = pd.DataFrame(X_raw, columns=columns)
df["fraud"] = y
print(df["fraud"].value_counts()) # ~49500 / ~500
print(f"Fraud rate: {y.mean():.3%}")What we are simulating and what we are not: the generated columns are abstract Gaussian combinations that we name so we can reason with them — the structure of the problem (few informative features, 1% positives, some label noise) is faithful to reality; the values are not literal hours or euros. In a real case, features like dist_home_km or ratio_avg_amount would come out of the feature engineering from 03-06 applied to the cardholder's history — exactly the way you built MercaFresh's RFM.
We split with stratification, essential with 1% positives (06-01):
from sklearn.model_selection import train_test_split
X = df.drop(columns="fraud")
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.25, stratify=y, random_state=42
)
print("Frauds in train:", y_train.sum(), "| in test:", y_test.sum())EDA with imbalanced classes
With imbalance, global histograms are useless: the frauds vanish under the legitimate mass. The right tool is comparing distributions per class:
import matplotlib.pyplot as plt
train = X_train.copy(); train["fraud"] = y_train
fig, axes = plt.subplots(2, 4, figsize=(15, 7))
for ax, col in zip(axes.ravel(), columns):
for cls, color in [(0, "tab:blue"), (1, "tab:red")]:
ax.hist(train.loc[train.fraud == cls, col], bins=40,
density=True, alpha=0.5, color=color)
ax.set_title(col)
plt.tight_layout(); plt.show()
print(train.groupby("fraud").mean().round(2).T)The density=True is the key: it normalizes each histogram so you can overlay 370 frauds on 37,000 legitimate transactions. Look for the features where the two bells separate (the informative ones) and those that overlap completely (the ones that will barely contribute). This very plot, comparing "today" against "the historical record", will be your drift detector in production — we come back to it at the end.
Why accuracy lies at 99%
The experiment that proves it, with DummyClassifier playing the role of metrics con artist:
from sklearn.dummy import DummyClassifier
from sklearn.metrics import accuracy_score, recall_score, precision_score
lazy = DummyClassifier(strategy="most_frequent").fit(X_train, y_train)
pred = lazy.predict(X_test)
print("Accuracy:", round(accuracy_score(y_test, pred), 4)) # ~0.99
print("Fraud recall:", recall_score(y_test, pred)) # 0.0
print("Fraud precision:", precision_score(y_test, pred, zero_division=0)) # 0.0A model that says "everything is legitimate" is right 99% of the time and does not catch a single fraud: total cost = every fraud paid in full. It is the canonical example from 06-02: when classes are imbalanced, accuracy mostly measures the share of the majority class. From here on, precision/recall/F1 for the fraud class, and euros.
Strategies against imbalance
Three families, from simplest to most invasive:
class_weight="balanced": the model weights every fraud example ~99 times more in its loss function. It does not touch the data, it is one line, and it is almost always the first thing to try.- Moving the threshold: train normally, but do not cut at 0.5 — cut wherever the cost matrix says (06-04). It is the most honest strategy because it separates the model (probabilities) from the decision (business threshold).
- Resampling: changing the training set itself. Oversample the minority (duplicate frauds, or generate interpolated synthetics with SMOTE, available in the
imbalanced-learnlibrary) or undersample the majority. In code we illustrate it with sklearn'sresample:
from sklearn.utils import resample
train_legit = train[train.fraud == 0]
train_fraud = train[train.fraud == 1]
fraud_over = resample(train_fraud, replace=True,
n_samples=len(train_legit) // 10, # 10:1 ratio, not 1:1
random_state=42)
train_bal = pd.concat([train_legit, fraud_over])
print(train_bal["fraud"].value_counts())Golden rule: resampling is applied to the training set only, never to the test set (we want to evaluate against the 1% reality) and, if combined with CV, inside each fold — resampling before splitting duplicates the same frauds across train and validation: the leak from 06-01, imbalanced edition.
Modeling and comparison with CV
Three models from the course with class_weight where it exists, compared with stratified CV (06-03) on AUC-PR (average_precision), recall and precision:
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import RandomForestClassifier, HistGradientBoostingClassifier
from sklearn.model_selection import cross_validate, StratifiedKFold
models = {
"Logistic": Pipeline([
("scale", StandardScaler()),
("m", LogisticRegression(class_weight="balanced", max_iter=1000)),
]),
"Random Forest": RandomForestClassifier(
n_estimators=200, class_weight="balanced", n_jobs=-1, random_state=42),
"HistGradientBoosting": HistGradientBoostingClassifier(random_state=42),
}
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
scoring = {"ap": "average_precision", "recall": "recall", "precision": "precision"}
rows = []
for name, m in models.items():
r = cross_validate(m, X_train, y_train, cv=cv, scoring=scoring)
rows.append({"Model": name,
"AUC-PR": r["test_ap"].mean(),
"[email protected]": r["test_recall"].mean(),
"[email protected]": r["test_precision"].mean()})
print(pd.DataFrame(rows).round(3).to_string(index=False))Ballpark results:
| Model | AUC-PR (CV) | Recall @0.5 | Precision @0.5 | Reading |
|---|---|---|---|---|
| HistGradientBoosting | ≈ 0.80 | ≈ 0.55 | ≈ 0.85 | Best ranking; its 0.5 threshold is conservative (no class_weight) |
| Random Forest | ≈ 0.75 | ≈ 0.65 | ≈ 0.70 | Solid; the class weight lifts its recall |
| Logistic | ≈ 0.65 | ≈ 0.85 | ≈ 0.10 | The 99:1 weight tips it into shouting "fraud" with little precision |
Notice that the @0.5 columns are almost anecdotal: they depend on a threshold we have not decided yet. AUC-PR, by contrast, evaluates the quality of the probability ranking across all thresholds at once — that is why it is the comparison metric. And a note on ROC-AUC: with 1% positives it usually comes out sky-high (0.95+) for everyone, because the abundant true negatives inflate it; the precision-recall curve, which ignores true negatives, is far more demanding and realistic here. It is the final nuance that was left as a note in 06-04.
The precision-recall curve and AUC-PR
With HistGradientBoosting chosen, we draw its curve using CV probabilities (without touching the test set):
from sklearn.model_selection import cross_val_predict
from sklearn.metrics import precision_recall_curve, average_precision_score
best = HistGradientBoostingClassifier(random_state=42)
proba_cv = cross_val_predict(best, X_train, y_train, cv=cv,
method="predict_proba")[:, 1]
prec, rec, thresholds = precision_recall_curve(y_train, proba_cv)
print("AUC-PR:", round(average_precision_score(y_train, proba_cv), 3))
plt.plot(rec, prec)
plt.axhline(y_train.mean(), ls="--", color="gray",
label=f"Chance = {y_train.mean():.1%}")
plt.xlabel("Recall (frauds caught)")
plt.ylabel("Precision (hits among alerts)")
plt.title("Precision-recall curve (CV)"); plt.legend(); plt.show()The gray line is humbling and necessary: a random classifier has precision = 1% on this problem. Everything above that line is the model's merit. The curve is a menu of trade-offs: each point is a possible threshold — more recall (catching more frauds) paid for with precision (more false alarms). Which one to pick? The one that minimizes euros.
The optimal threshold, in euros
We apply the cost matrix to every candidate threshold, as in 06-04 but with currency:
COST_FN = 150 # undetected fraud
COST_FP = 5 # legitimate transaction blocked/reviewed
results = []
for t in np.linspace(0.01, 0.99, 99):
pred = (proba_cv >= t).astype(int)
fn = ((y_train == 1) & (pred == 0)).sum()
fp = ((y_train == 0) & (pred == 1)).sum()
results.append({"threshold": t, "cost": fn * COST_FN + fp * COST_FP})
costs = pd.DataFrame(results)
optimal = costs.loc[costs.cost.idxmin()]
print(f"Optimal threshold: {optimal.threshold:.2f} | CV cost: {optimal.cost:,.0f} EUR")
plt.plot(costs.threshold, costs.cost)
plt.axvline(optimal.threshold, color="red", ls="--")
plt.xlabel("Threshold"); plt.ylabel("Total cost (EUR)")
plt.title("Expected cost by threshold"); plt.show()With a false negative 30 times more expensive than a false positive, the optimal threshold falls well below 0.5 (typically between 0.05 and 0.15): it pays to review plenty of legitimate transactions in order to catch more frauds. Compare the optimal threshold's cost against two references: the "default" 0.5 threshold and the "block nothing" policy (= every fraud paid). The difference, in euros, is the argument this project defends itself with in front of the business — far more persuasive than any F1.
Final evaluation on the test set
We train on the full training set, apply the chosen threshold, and open the test set exactly once:
from sklearn.metrics import classification_report, confusion_matrix
best.fit(X_train, y_train)
proba_test = best.predict_proba(X_test)[:, 1]
pred_test = (proba_test >= optimal.threshold).astype(int)
print(confusion_matrix(y_test, pred_test))
print(classification_report(y_test, pred_test,
target_names=["legitimate", "fraud"], digits=3))
fn = ((y_test == 1) & (pred_test == 0)).sum()
fp = ((y_test == 0) & (pred_test == 1)).sum()
cost_model = fn * COST_FN + fp * COST_FP
cost_no_model = y_test.sum() * COST_FN
print(f"Cost with model: {cost_model:,.0f} EUR | without model: {cost_no_model:,.0f} EUR")
print(f"Savings: {cost_no_model - cost_model:,.0f} EUR")The final report must tell three numbers: fraud recall (what fraction of the fraud we catch), precision (what fraction of the alerts are real — that is the review team's workload) and the savings in euros versus doing nothing. If they match the CV estimates reasonably well, the protocol was clean.
Common Mistakes and Tips
- Splitting without stratifying. With 1% positives, a random split can leave the test set almost fraud-free and the metrics turn into noise.
stratify=y, always. - Resampling before splitting. Duplicating frauds and then dealing out train/test puts copies of the same fraud on both sides: inflated, illusory metrics. Resample only within the training set, and inside each fold if there is CV.
- Showing off ROC-AUC. A 0.95 ROC-AUC with 1% positives can coexist with 8% precision. Report AUC-PR and the precision-recall curve.
- Setting the threshold with the test set. The threshold is a hyperparameter of the decision: it is chosen with CV on the training set. Picking it by looking at the test set is overfitting the final decision.
- Treating the cost matrix as eternal. The €150/€5 are business estimates that change (average amounts, cost of the review team). Document the figures and recompute the threshold when they change.
Challenges to go further
- IsolationForest, the unsupervised cousin. Train
sklearn.ensemble.IsolationForestignoring the labels and check how many real frauds show up among its anomalies, connecting with the approach from 05-04 "DBSCAN clustering". Hint: usecontamination=0.01; it is the tool for the cold start, when there are no historical labels yet. - Sensitivity to the cost matrix. Recompute the optimal threshold for several scenarios (FN = €50, €300; FP = €1, €20) and plot how it moves. Hint: if the optimal threshold barely changes across plausible scenarios, your decision is robust; if it dances around, the business must refine its estimates before deploying.
- Temporal drift. Simulate fraud evolving: generate a second dataset with a different
random_stateand slightly moreclass_sepor an offset added to two features, and evaluate the old model on it. Hint: measure how much recall and AUC-PR drop and which per-class histograms give the change away — you are reproducing the cycle from 08-03 in miniature.
Conclusions and the link to production
This project has taught you how to work when the class that matters is a needle in a haystack: stratify everything, distrust accuracy (and even ROC-AUC), compare models by AUC-PR, and above all separate the model from the decision — the algorithm supplies the probabilities, the euros supply the threshold, exactly the method from 06-04 taken to its logical conclusion. You have also seen the arsenal against imbalance (class weights, threshold, resampling) and its traps. And there remains the final warning, which connects with 08-02 "Deploying models to production" and 08-03 "Model maintenance and monitoring": fraud is the most drift-prone problem in existence, because on the other side there are adversaries adapting to your detector — the model that saves thousands of euros today degrades within months, and histogram monitoring plus periodic retraining are not optional, they are part of the product. With this you close the block of supervised projects; in the last project we go back home, to MercaFresh, to close the circle with unsupervised learning: segmenting customers from start to finish.
Machine Learning Course
Module 1: Introduction to Machine Learning
- What is Machine Learning?
- History and evolution of Machine Learning
- Types of Machine Learning
- Applications of Machine Learning
- The Machine Learning project workflow
Module 2: Foundations of Statistics and Probability
- Basic statistics concepts
- Probability distributions
- Correlation and covariance
- Statistical inference
- Bayes' theorem
Module 3: Data Preprocessing
- Data cleaning
- Handling missing data
- Data transformation
- Encoding categorical variables
- Normalization and standardization
- Feature engineering
Module 4: Supervised Machine Learning Algorithms
- Linear regression
- Logistic regression
- Decision trees
- Support Vector Machines (SVM)
- K-Nearest Neighbors (K-NN)
- Naive Bayes
- Neural networks
Module 5: Unsupervised Machine Learning Algorithms
- Clustering: K-means
- Hierarchical clustering
- Principal Component Analysis (PCA)
- DBSCAN clustering
- Data visualization with t-SNE and UMAP
Module 6: Model Evaluation and Validation
- Data splitting: training, validation and test
- Evaluation metrics
- Cross-validation
- ROC curve and AUC
- Overfitting and underfitting
Module 7: Advanced Techniques and Optimization
- Regularization: Ridge, Lasso and Elastic Net
- Ensemble Learning
- Gradient Boosting
- Deep neural networks (Deep Learning)
- Hyperparameter optimization
Module 8: Model Implementation and Deployment
- Popular frameworks and libraries
- Deploying models to production
- Model maintenance and monitoring
- Ethical and privacy considerations
Module 9: Hands-On Projects
- Project 1: Housing price prediction
- Project 2: Image classification
- Project 3: Sentiment analysis on social media
- Project 4: Fraud detection
- Project 5: Customer segmentation
