In 05-02 we distilled the principle of autoencoders into one sentence: "the unusual reconstructs badly". And we promised to apply it to fraud. This third project delivers on that: you will build a fraudulent transaction detector for TecnoMarket by training an autoencoder only on normal transactions and using the reconstruction error as an anomaly score. Beyond the model, this project teaches you two disciplines that are indispensable in real problems: evaluating with the right metrics when the classes are brutally imbalanced, and translating the confusion matrix into business cost.

Contents

  1. Project brief and why this is not supervised classification
  2. Phase 1: the transaction dataset
  3. Phase 2: splitting and scaling without leakage
  4. Phase 3: the autoencoder, trained only on the normal
  5. Phase 4: reconstruction error as an anomaly score
  6. Phase 5: choosing the threshold
  7. Phase 6: metrics under imbalance and the confusion matrix as business cost
  8. Phase 7: delivery — review queue and API

Project brief and why this is not supervised classification

Business context. TecnoMarket processes tens of thousands of orders a day; a tiny fraction is fraud (stolen cards, compromised accounts). The anti-fraud team can't keep up reviewing by hand.

Why not train a supervised fraud/not-fraud classifier like the ones in 04-03? Two reasons:

  • There are almost no fraud examples (here, 0.5%), and the ones that exist are biased toward the frauds we already know how to detect.
  • Fraud mutates: a new pattern doesn't resemble the labeled ones. A supervised classifier only recognizes the fraud of the past.

The unsupervised approach from 05-02 sidesteps both: the autoencoder learns to compress and reconstruct only normal transactions. Anything that doesn't fit that mold — old fraud or new — will reconstruct badly and set off the alarm.

Phase 1: the transaction dataset

We generate a fictional dataset with numpy, but with realistic structure: five variables per transaction, coherent normal patterns and 0.5% of frauds with different patterns (the same controlled-synthetic-data philosophy we used with the sales series in 04-04):

import numpy as np

rng = np.random.default_rng(42)   # seed: 06-04
N_NORMAL, N_FRAUD = 20_000, 100   # 0.5% fraud

# --- Normal transactions ---
amount   = rng.lognormal(mean=3.6, sigma=0.7, size=N_NORMAL)       # ~€20-150
hour     = np.clip(rng.normal(15, 4, N_NORMAL) % 24, 0, 24)        # afternoon peak
n_items  = rng.poisson(2, N_NORMAL) + 1                            # 1-5 items
dist_km  = rng.exponential(30, N_NORMAL)                           # nearby shipping
age_days = rng.exponential(400, N_NORMAL)                          # long-standing accounts

X_normal = np.column_stack([amount, hour, n_items, dist_km, age_days])

# --- Frauds: different patterns, not a single recipe ---
half = N_FRAUD // 2
# Pattern A: high amount, small hours, freshly created account
fraud_a = np.column_stack([
    rng.lognormal(6.2, 0.4, half),             # ~€300-900
    rng.uniform(1, 5, half),                   # 1-5 in the small hours
    rng.poisson(1, half) + 1,
    rng.exponential(30, half),
    rng.uniform(0, 10, half),                  # days-old account
])
# Pattern B: many expensive items shipped to a distant address
fraud_b = np.column_stack([
    rng.lognormal(5.8, 0.3, N_FRAUD - half),
    np.clip(rng.normal(15, 4, N_FRAUD - half) % 24, 0, 24),
    rng.poisson(8, N_FRAUD - half) + 3,        # 8-15 items
    rng.exponential(600, N_FRAUD - half),      # shipping very far away
    rng.uniform(0, 60, N_FRAUD - half),
])
X_fraud = np.vstack([fraud_a, fraud_b])

COLUMNS = ["amount", "hour", "n_items", "dist_km", "age_days"]

Explore before modeling (phase 1 of 07-01): per-variable histograms overlaying normal and fraud. You'll see that no single variable separates fraud on its own — it's the combination (high amount + small hours + new account) that is anomalous. Exactly what an autoencoder can capture and a column-by-column manual rule cannot.

Phase 2: splitting and scaling without leakage

Two golden rules, both already familiar:

  1. The autoencoder is trained ONLY on normal transactions. Fraud never enters training: it is reserved in full for evaluation.
  2. The scaler is fitted only on the train set (the anti-leakage rule from 04-04): if the scaling saw the test set — or the frauds — we would be leaking information from the future.
from sklearn.preprocessing import StandardScaler

rng.shuffle(X_normal)
n_train = 14_000
X_train      = X_normal[:n_train]            # normals only: training
X_val_normal = X_normal[n_train:17_000]      # normals only: threshold and stopping
X_test = np.vstack([X_normal[17_000:], X_fraud])   # normals + ALL the fraud
y_test = np.concatenate([np.zeros(3_000), np.ones(N_FRAUD)])

scaler = StandardScaler().fit(X_train)       # ONLY on train
X_train_s      = scaler.transform(X_train)
X_val_s        = scaler.transform(X_val_normal)
X_test_s       = scaler.transform(X_test)

Note the deliberate asymmetry of the split: train is pure normal; test concentrates the 100 frauds alongside 3,000 normals. This is the standard structure of an anomaly detection problem.

Phase 3: the autoencoder, trained only on the normal

The architecture is the dense autoencoder from 05-02, adapted from images to tabular data: 5 → bottleneck → 5. The 2-3 unit bottleneck forces compression; only the frequent (normal) patterns survive the squeeze:

import tensorflow as tf
from tensorflow.keras import layers, models

tf.random.set_seed(42)

autoencoder = models.Sequential([
    layers.Input(shape=(5,)),
    layers.Dense(16, activation="relu"),
    layers.Dense(3, activation="relu", name="bottleneck"),   # compression 5 -> 3
    layers.Dense(16, activation="relu"),
    layers.Dense(5, activation="linear"),                    # reconstructs all 5 vars
])

autoencoder.compile(optimizer="adam", loss="mse")

history = autoencoder.fit(
    X_train_s, X_train_s,              # input = target: learn to copy
    validation_data=(X_val_s, X_val_s),
    epochs=100, batch_size=64,
    callbacks=[tf.keras.callbacks.EarlyStopping(patience=10,
                                                restore_best_weights=True)],
    verbose=0,
)
print(f"Final val MSE: {history.history['val_loss'][-1]:.4f}")  # typical ~0.05-0.15

Key details:

  • linear output, MSE loss: the scaled data lives on the whole real line, not in [0,1] like the pixels of 05-02.
  • Validation here is of normals: it measures whether the model generalizes to unseen normals, and it drives EarlyStopping (06-01). Fraud still never appears.
  • It trains in seconds: the value of this project lies not in the model's size but in the protocol.

Phase 4: reconstruction error as an anomaly score

Now we apply the 05-02 principle: run each transaction through the autoencoder and measure how closely the reconstruction resembles the original.

def anomaly_score(model, X):
    reconstruction = model.predict(X, verbose=0)
    return np.mean((X - reconstruction) ** 2, axis=1)   # MSE per transaction

score_val  = anomaly_score(autoencoder, X_val_s)
score_test = anomaly_score(autoencoder, X_test_s)

print(f"Mean score, normals (test): {score_test[y_test == 0].mean():.3f}")  # ~0.1
print(f"Mean score, frauds  (test): {score_test[y_test == 1].mean():.3f}")  # ~3-15

Plot both histograms (normal vs. fraud) on a log scale: you'll see two separated distributions with an overlap — some atypical normals score high and the odd discreet fraud scores low. That overlap is why choosing a threshold is a trade-off, not a magic boundary.

Phase 5: choosing the threshold

The threshold turns the continuous score into a binary decision. A practical approach: set it as a percentile of the validation scores (normals only) — for example, the 99th percentile means "we accept flagging ~1% of legitimate transactions":

threshold = np.percentile(score_val, 99)
y_pred = (score_test > threshold).astype(int)

But a blindly chosen percentile may not be the best trade-off. Explore several and measure their consequences:

from sklearn.metrics import precision_score, recall_score, f1_score

print(f"{'percentile':>10} | {'threshold':>9} | {'precision':>9} | {'recall':>6} | {'F1':>5}")
for p in [90, 95, 97, 99, 99.5]:
    t = np.percentile(score_val, p)
    pred = (score_test > t).astype(int)
    print(f"{p:>10} | {t:>9.2f} | {precision_score(y_test, pred):>9.2f} "
          f"| {recall_score(y_test, pred):>6.2f} | {f1_score(y_test, pred):>5.2f}")

Typical result with this dataset (your figures will vary somewhat):

Percentile Precision Recall Business reading
90 ~0.23 ~0.98 Catches almost all the fraud, but 3 out of 4 alerts are false
95 ~0.38 ~0.96 Less noise, still catches nearly everything
99 ~0.72 ~0.88 Alerts mostly genuine; ~12% of the fraud slips through
99.5 ~0.83 ~0.80 Very reliable, but 1 in 5 frauds gets past

Raising the threshold buys precision (fewer customers bothered) by selling recall (more fraud let through). There is no correct threshold in the abstract: there is a correct threshold for a given set of costs.

Phase 6: metrics under imbalance and the confusion matrix as business cost

Why accuracy lies here. With 0.5% fraud, a "model" that always answers "normal" is right 99.5% of the time. Spectacular and useless: recall 0. Under heavy imbalance, accuracy mostly measures the size of the majority class. That's why this project is evaluated with precision, recall and F1:

Metric Question it answers At TecnoMarket
Precision Of what I flag, how much is real fraud? How many legitimate customers I bother
Recall Of the real fraud, how much do I catch? How much money I stop losing
F1 Harmonic mean of the two The trade-off in a single number

The confusion matrix, in euros. Assume the 99th-percentile threshold and its typical matrix on the test set (3,000 normals + 100 frauds):

Predicted normal Predicted fraud
Actual normal 2,966 (TN) 34 (FP)
Actual fraud 12 (FN) 88 (TP)

Translation into cost, with explicit assumptions (average undetected fraud ≈ €400 lost; false alert ≈ €5 of review cost + customer friction):

  • FN (missed fraud): 12 × 400 = €4,800 of direct loss.
  • FP (annoyed customer): 34 × 5 = €170 + reputational risk.

Under these assumptions, an FN costs ~80 times more than an FP: it pays to lower the threshold (95th-97th percentile) and accept more false alerts. If your review process were expensive or automatically blocked orders, the balance would shift. The threshold is a business decision informed by the model, not a technical hyperparameter.

Phase 7: delivery — review queue and API

The final decision isn't binary but three-way, reusing the human review queue from 03-04:

REVIEW_THRESHOLD = np.percentile(score_val, 95)   # suspicious: to review
BLOCK_THRESHOLD  = np.percentile(score_val, 99.9) # flagrant: hold the order

def score_transaction(x_raw):
    x = scaler.transform(x_raw.reshape(1, -1))          # SAME scaler as in training
    score = float(anomaly_score(autoencoder, x)[0])
    if score > BLOCK_THRESHOLD:
        return {"score": score, "action": "hold_and_verify"}
    if score > REVIEW_THRESHOLD:
        return {"score": score, "action": "review_queue"}
    return {"score": score, "action": "approve"}

For deployment we follow 06-05 to the letter: scaler and model are saved and versioned as one unit (for example, the scaler serialized alongside the .keras file, or built in as a Normalization layer inside the inference model), validated against a frozen set before every update, and served behind a FastAPI endpoint (POST /score-transaction) identical in structure to the review classifier's. An important operational point: what counts as "normal" shifts with the seasons (remember the Black Friday in the 04-04 series) — schedule periodic retraining on recent normal data so the sales campaign doesn't turn into a storm of false alarms.

Common Mistakes and Tips

  • Letting frauds slip into training: if the autoencoder learns to reconstruct fraud, fraud stops scoring high. The purity of the train set is the premise of the whole method.
  • Fitting the scaler on all the data: classic leakage (04-04). The scaler is fitted on the normal train set and applied, frozen, to everything else.
  • Boasting about accuracy: 99% accuracy means nothing here. Always report precision, recall and F1, plus the confusion matrix.
  • Choosing the threshold on the test set: the test set (frauds included) is the frozen set. The threshold is set with validation (percentiles of normals) or with business costs; the test set only measures the outcome.
  • A single, immovable threshold: the score distribution drifts over time. Monitor the daily alert percentage: if it spikes with no campaign to explain it, it's time to retrain or recalibrate.

Exercises

  1. Shrink the bottleneck to 2 units and widen it to 4. Compare the separation between normal and fraud scores (for example, using each group's median) in the three cases. What happens with a bottleneck that is too wide?
  2. Add a third fraud pattern that doesn't exist in the current data (for example, a normal amount but hundreds of items) and evaluate the detector without retraining it. Does it catch it? What does this tell you compared with a supervised classifier?
  3. Define a total cost function cost(threshold) = 400 * FN + 5 * FP, evaluate it for percentiles from 90 to 99.9 and find the threshold that minimizes it. Compare with the one you eyeballed in phase 5.

Solutions

  1. With a bottleneck of 2, the reconstruction of normals worsens a little (the baseline score rises) but the relative separation usually holds or improves; with 4-5 units the autoencoder approaches the identity function and reconstructs the frauds reasonably well too: the separation collapses. Lesson: the narrow bottleneck is not a defect, it is the mechanism.
  2. Generate, for example, np.column_stack([rng.lognormal(3.6,0.7,50), normal_hours, rng.poisson(200,50), normal_dists, normal_ages]), scale with the same scaler and score it. The detector typically nails it with sky-high scores: n_items=200 is miles away from everything it learned. A supervised classifier trained only on patterns A and B would have no guarantee whatsoever of detecting it — this is the structural advantage of the unsupervised approach.
  3. Sweep the percentiles, compute the confusion matrix with confusion_matrix(y_test, pred).ravel() (order tn, fp, fn, tp) and evaluate 400*fn + 5*fp. With the given costs, the minimum usually falls between the 95th and 97th percentiles — more aggressive than the "intuitive" 99th, confirming the phase 6 analysis. If you repeat with an FP cost of €50, the optimum shifts upward: the optimal threshold is a function of the costs.

Conclusion

Third project delivered: a fraud detector that needs no fraud examples to train. You have applied the autoencoder from 05-02 with the data discipline of 04-04 (leak-free scaling, honest splits), turned the reconstruction error into an anomaly score, chosen a threshold as a precision-recall trade-off and translated the confusion matrix into euros — the step that turns a technical exercise into a business tool with its review queue (03-04) and its API (06-05). In the next lesson we return to images and settle module 5's most anticipated debt: the complete training loop of a GAN with GradientTape, for TecnoMarket's creatives-generation prototype.

© Copyright 2026. All rights reserved