Fourth project and a new enemy: extreme class imbalance. You are going to build a fraudulent card transaction detector where only 1% of the cases are fraud — the scenario where accuracy deceives, thresholds are decided in euros and the metrics from module 6 earn their keep. The problem will feel familiar twice over: it is the spam detector from Bayes' theorem in 02-05 with money on the line, and the supervised cousin of the order anomalies you hunted with DBSCAN in 05-04 — except here we do have historical fraud labels, so we can train classifiers. We will use a synthetic dataset generated right in the code (real fraud datasets are confidential by nature; the classic public one on Kaggle comes anonymized through PCA), which also gives us full control over the experiment.

Contents

  1. Problem definition, costs and target metric
  2. Generating the synthetic dataset
  3. EDA with imbalanced classes
  4. Why accuracy lies at 99%
  5. Strategies against imbalance
  6. Modeling and comparison with CV
  7. The precision-recall curve and AUC-PR
  8. The optimal threshold, in euros
  9. Final evaluation on the test set
  10. Conclusions and the link to production

Problem definition, costs and target metric

  • Problem: binary classification — is this transaction fraudulent (1) or legitimate (0)?
  • The asymmetry that governs everything: the two errors do not cost the same. An undetected fraud (false negative) costs the defrauded amount plus handling, say €150 on average; blocking a legitimate transaction (false positive) costs the manual review plus customer friction, say €5. This cost matrix, as you learned in 06-04 "ROC curve and AUC", will decide the threshold.
  • Metrics: precision, recall and F1 for the fraud class (06-02 "Evaluation metrics"), the precision-recall curve with its AUC-PR as the model-comparison metric, and the total cost in euros as the final business metric. Accuracy is banned, and you are about to see why.
Predicted: legitimate Predicted: fraud
Actual: legitimate €0 €5 (review + friction)
Actual: fraud €150 (loss) €0 (fraud prevented)

Generating the synthetic dataset

make_classification generates a classification problem with controlled structure; with weights=[0.99, 0.01] we impose the imbalance. Then we dress the abstract features with realistic names and scales so we can reason like analysts:

import numpy as np
import pandas as pd
from sklearn.datasets import make_classification

X_raw, y = make_classification(
    n_samples=50_000, n_features=8, n_informative=5, n_redundant=1,
    weights=[0.99, 0.01],          # 1% fraud
    class_sep=1.0, flip_y=0.005,   # some label noise: realism
    random_state=42,
)

columns = [
    "amount_eur",           # transaction amount
    "hour_of_day",          # time of day (fraud prefers the small hours)
    "dist_home_km",         # distance from the merchant to the cardholder's home
    "merchant_freq_month",  # times the holder buys at that merchant per month
    "txns_last_hour",       # card transactions in the last hour
    "card_age_months",      # months since the card was issued
    "ratio_avg_amount",     # amount / holder's historical average spend
    "online_payment",       # channel signal (higher => more online)
]
df = pd.DataFrame(X_raw, columns=columns)
df["fraud"] = y

print(df["fraud"].value_counts())        # ~49500 / ~500
print(f"Fraud rate: {y.mean():.3%}")

What we are simulating and what we are not: the generated columns are abstract Gaussian combinations that we name so we can reason with them — the structure of the problem (few informative features, 1% positives, some label noise) is faithful to reality; the values are not literal hours or euros. In a real case, features like dist_home_km or ratio_avg_amount would come out of the feature engineering from 03-06 applied to the cardholder's history — exactly the way you built MercaFresh's RFM.

We split with stratification, essential with 1% positives (06-01):

from sklearn.model_selection import train_test_split

X = df.drop(columns="fraud")
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.25, stratify=y, random_state=42
)
print("Frauds in train:", y_train.sum(), "| in test:", y_test.sum())

EDA with imbalanced classes

With imbalance, global histograms are useless: the frauds vanish under the legitimate mass. The right tool is comparing distributions per class:

import matplotlib.pyplot as plt

train = X_train.copy(); train["fraud"] = y_train

fig, axes = plt.subplots(2, 4, figsize=(15, 7))
for ax, col in zip(axes.ravel(), columns):
    for cls, color in [(0, "tab:blue"), (1, "tab:red")]:
        ax.hist(train.loc[train.fraud == cls, col], bins=40,
                density=True, alpha=0.5, color=color)
    ax.set_title(col)
plt.tight_layout(); plt.show()

print(train.groupby("fraud").mean().round(2).T)

The density=True is the key: it normalizes each histogram so you can overlay 370 frauds on 37,000 legitimate transactions. Look for the features where the two bells separate (the informative ones) and those that overlap completely (the ones that will barely contribute). This very plot, comparing "today" against "the historical record", will be your drift detector in production — we come back to it at the end.

Why accuracy lies at 99%

The experiment that proves it, with DummyClassifier playing the role of metrics con artist:

from sklearn.dummy import DummyClassifier
from sklearn.metrics import accuracy_score, recall_score, precision_score

lazy = DummyClassifier(strategy="most_frequent").fit(X_train, y_train)
pred = lazy.predict(X_test)

print("Accuracy:", round(accuracy_score(y_test, pred), 4))  # ~0.99
print("Fraud recall:", recall_score(y_test, pred))          # 0.0
print("Fraud precision:", precision_score(y_test, pred, zero_division=0))  # 0.0

A model that says "everything is legitimate" is right 99% of the time and does not catch a single fraud: total cost = every fraud paid in full. It is the canonical example from 06-02: when classes are imbalanced, accuracy mostly measures the share of the majority class. From here on, precision/recall/F1 for the fraud class, and euros.

Strategies against imbalance

Three families, from simplest to most invasive:

  1. class_weight="balanced": the model weights every fraud example ~99 times more in its loss function. It does not touch the data, it is one line, and it is almost always the first thing to try.
  2. Moving the threshold: train normally, but do not cut at 0.5 — cut wherever the cost matrix says (06-04). It is the most honest strategy because it separates the model (probabilities) from the decision (business threshold).
  3. Resampling: changing the training set itself. Oversample the minority (duplicate frauds, or generate interpolated synthetics with SMOTE, available in the imbalanced-learn library) or undersample the majority. In code we illustrate it with sklearn's resample:
from sklearn.utils import resample

train_legit = train[train.fraud == 0]
train_fraud = train[train.fraud == 1]

fraud_over = resample(train_fraud, replace=True,
                      n_samples=len(train_legit) // 10,  # 10:1 ratio, not 1:1
                      random_state=42)
train_bal = pd.concat([train_legit, fraud_over])
print(train_bal["fraud"].value_counts())

Golden rule: resampling is applied to the training set only, never to the test set (we want to evaluate against the 1% reality) and, if combined with CV, inside each fold — resampling before splitting duplicates the same frauds across train and validation: the leak from 06-01, imbalanced edition.

Modeling and comparison with CV

Three models from the course with class_weight where it exists, compared with stratified CV (06-03) on AUC-PR (average_precision), recall and precision:

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import RandomForestClassifier, HistGradientBoostingClassifier
from sklearn.model_selection import cross_validate, StratifiedKFold

models = {
    "Logistic": Pipeline([
        ("scale", StandardScaler()),
        ("m", LogisticRegression(class_weight="balanced", max_iter=1000)),
    ]),
    "Random Forest": RandomForestClassifier(
        n_estimators=200, class_weight="balanced", n_jobs=-1, random_state=42),
    "HistGradientBoosting": HistGradientBoostingClassifier(random_state=42),
}

cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
scoring = {"ap": "average_precision", "recall": "recall", "precision": "precision"}

rows = []
for name, m in models.items():
    r = cross_validate(m, X_train, y_train, cv=cv, scoring=scoring)
    rows.append({"Model": name,
                 "AUC-PR": r["test_ap"].mean(),
                 "[email protected]": r["test_recall"].mean(),
                 "[email protected]": r["test_precision"].mean()})
print(pd.DataFrame(rows).round(3).to_string(index=False))

Ballpark results:

Model AUC-PR (CV) Recall @0.5 Precision @0.5 Reading
HistGradientBoosting ≈ 0.80 ≈ 0.55 ≈ 0.85 Best ranking; its 0.5 threshold is conservative (no class_weight)
Random Forest ≈ 0.75 ≈ 0.65 ≈ 0.70 Solid; the class weight lifts its recall
Logistic ≈ 0.65 ≈ 0.85 ≈ 0.10 The 99:1 weight tips it into shouting "fraud" with little precision

Notice that the @0.5 columns are almost anecdotal: they depend on a threshold we have not decided yet. AUC-PR, by contrast, evaluates the quality of the probability ranking across all thresholds at once — that is why it is the comparison metric. And a note on ROC-AUC: with 1% positives it usually comes out sky-high (0.95+) for everyone, because the abundant true negatives inflate it; the precision-recall curve, which ignores true negatives, is far more demanding and realistic here. It is the final nuance that was left as a note in 06-04.

The precision-recall curve and AUC-PR

With HistGradientBoosting chosen, we draw its curve using CV probabilities (without touching the test set):

from sklearn.model_selection import cross_val_predict
from sklearn.metrics import precision_recall_curve, average_precision_score

best = HistGradientBoostingClassifier(random_state=42)
proba_cv = cross_val_predict(best, X_train, y_train, cv=cv,
                             method="predict_proba")[:, 1]

prec, rec, thresholds = precision_recall_curve(y_train, proba_cv)
print("AUC-PR:", round(average_precision_score(y_train, proba_cv), 3))

plt.plot(rec, prec)
plt.axhline(y_train.mean(), ls="--", color="gray",
            label=f"Chance = {y_train.mean():.1%}")
plt.xlabel("Recall (frauds caught)")
plt.ylabel("Precision (hits among alerts)")
plt.title("Precision-recall curve (CV)"); plt.legend(); plt.show()

The gray line is humbling and necessary: a random classifier has precision = 1% on this problem. Everything above that line is the model's merit. The curve is a menu of trade-offs: each point is a possible threshold — more recall (catching more frauds) paid for with precision (more false alarms). Which one to pick? The one that minimizes euros.

The optimal threshold, in euros

We apply the cost matrix to every candidate threshold, as in 06-04 but with currency:

COST_FN = 150  # undetected fraud
COST_FP = 5    # legitimate transaction blocked/reviewed

results = []
for t in np.linspace(0.01, 0.99, 99):
    pred = (proba_cv >= t).astype(int)
    fn = ((y_train == 1) & (pred == 0)).sum()
    fp = ((y_train == 0) & (pred == 1)).sum()
    results.append({"threshold": t, "cost": fn * COST_FN + fp * COST_FP})

costs = pd.DataFrame(results)
optimal = costs.loc[costs.cost.idxmin()]
print(f"Optimal threshold: {optimal.threshold:.2f} | CV cost: {optimal.cost:,.0f} EUR")

plt.plot(costs.threshold, costs.cost)
plt.axvline(optimal.threshold, color="red", ls="--")
plt.xlabel("Threshold"); plt.ylabel("Total cost (EUR)")
plt.title("Expected cost by threshold"); plt.show()

With a false negative 30 times more expensive than a false positive, the optimal threshold falls well below 0.5 (typically between 0.05 and 0.15): it pays to review plenty of legitimate transactions in order to catch more frauds. Compare the optimal threshold's cost against two references: the "default" 0.5 threshold and the "block nothing" policy (= every fraud paid). The difference, in euros, is the argument this project defends itself with in front of the business — far more persuasive than any F1.

Final evaluation on the test set

We train on the full training set, apply the chosen threshold, and open the test set exactly once:

from sklearn.metrics import classification_report, confusion_matrix

best.fit(X_train, y_train)
proba_test = best.predict_proba(X_test)[:, 1]
pred_test = (proba_test >= optimal.threshold).astype(int)

print(confusion_matrix(y_test, pred_test))
print(classification_report(y_test, pred_test,
                            target_names=["legitimate", "fraud"], digits=3))

fn = ((y_test == 1) & (pred_test == 0)).sum()
fp = ((y_test == 0) & (pred_test == 1)).sum()
cost_model = fn * COST_FN + fp * COST_FP
cost_no_model = y_test.sum() * COST_FN
print(f"Cost with model: {cost_model:,.0f} EUR | without model: {cost_no_model:,.0f} EUR")
print(f"Savings: {cost_no_model - cost_model:,.0f} EUR")

The final report must tell three numbers: fraud recall (what fraction of the fraud we catch), precision (what fraction of the alerts are real — that is the review team's workload) and the savings in euros versus doing nothing. If they match the CV estimates reasonably well, the protocol was clean.

Common Mistakes and Tips

  • Splitting without stratifying. With 1% positives, a random split can leave the test set almost fraud-free and the metrics turn into noise. stratify=y, always.
  • Resampling before splitting. Duplicating frauds and then dealing out train/test puts copies of the same fraud on both sides: inflated, illusory metrics. Resample only within the training set, and inside each fold if there is CV.
  • Showing off ROC-AUC. A 0.95 ROC-AUC with 1% positives can coexist with 8% precision. Report AUC-PR and the precision-recall curve.
  • Setting the threshold with the test set. The threshold is a hyperparameter of the decision: it is chosen with CV on the training set. Picking it by looking at the test set is overfitting the final decision.
  • Treating the cost matrix as eternal. The €150/€5 are business estimates that change (average amounts, cost of the review team). Document the figures and recompute the threshold when they change.

Challenges to go further

  1. IsolationForest, the unsupervised cousin. Train sklearn.ensemble.IsolationForest ignoring the labels and check how many real frauds show up among its anomalies, connecting with the approach from 05-04 "DBSCAN clustering". Hint: use contamination=0.01; it is the tool for the cold start, when there are no historical labels yet.
  2. Sensitivity to the cost matrix. Recompute the optimal threshold for several scenarios (FN = €50, €300; FP = €1, €20) and plot how it moves. Hint: if the optimal threshold barely changes across plausible scenarios, your decision is robust; if it dances around, the business must refine its estimates before deploying.
  3. Temporal drift. Simulate fraud evolving: generate a second dataset with a different random_state and slightly more class_sep or an offset added to two features, and evaluate the old model on it. Hint: measure how much recall and AUC-PR drop and which per-class histograms give the change away — you are reproducing the cycle from 08-03 in miniature.

Conclusions and the link to production

This project has taught you how to work when the class that matters is a needle in a haystack: stratify everything, distrust accuracy (and even ROC-AUC), compare models by AUC-PR, and above all separate the model from the decision — the algorithm supplies the probabilities, the euros supply the threshold, exactly the method from 06-04 taken to its logical conclusion. You have also seen the arsenal against imbalance (class weights, threshold, resampling) and its traps. And there remains the final warning, which connects with 08-02 "Deploying models to production" and 08-03 "Model maintenance and monitoring": fraud is the most drift-prone problem in existence, because on the other side there are adversaries adapting to your detector — the model that saves thousands of euros today degrades within months, and histogram monitoring plus periodic retraining are not optional, they are part of the product. With this you close the block of supervised projects; in the last project we go back home, to MercaFresh, to close the circle with unsupervised learning: segmenting customers from start to finish.

Machine Learning Course

Module 1: Introduction to Machine Learning

Module 2: Foundations of Statistics and Probability

Module 3: Data Preprocessing

Module 4: Supervised Machine Learning Algorithms

Module 5: Unsupervised Machine Learning Algorithms

Module 6: Model Evaluation and Validation

Module 7: Advanced Techniques and Optimization

Module 8: Model Implementation and Deployment

Module 9: Hands-On Projects

Module 10: Additional Resources

© Copyright 2026. All rights reserved