In 09-02 you closed each practice with a business decision and with an open question: when does a neural network pay off against a well-built classical model? This lesson answers it hands on, in PyTorch on CPU, with four guided projects that consolidate module 5: the returns MLP with a manual search over architecture and learning rate (05-02, 05-03), the incident-photo CNN with a third class and data augmentation (05-04), a GRU on the NovaClean weekly demand (05-04) and a sentiment classifier with embeddings trained from scratch (05-05). Each project is compared with the classical reference from module 4 under the same conditions, and in several of them the network does not win: understanding why is the goal.
How to work through the lesson: each project has numbered steps; try to write the code for each step before looking at ours, run it (they all take seconds on CPU; the longest, a couple of minutes) and compare with the outputs, which vary slightly with the seed and the PyTorch version. You need novamarket_ml.py (module 4), reviews_nm.py (the POSITIVES, NEGATIVES and PRODUCTS templates from 05-05), torch, scikit-learn, pandas and numpy. Guideline time: 45-60 minutes per project.
Contents
- Project 1: returns MLP with an architecture and learning-rate search
- Project 2: incident-photo CNN with a third class and data augmentation
- Project 3: GRU for weekly demand against the baselines and the linear model
- Project 4: review sentiment with embeddings from scratch against the bag of words
- Common Mistakes and Tips
- Conclusion
- Project 1: returns MLP with an architecture and learning-rate search
Reminder (05-03, section 9). The 21 → 16 → 8 → 1 MLP with dropout 0.2, Adam (0.001), batches of 64 and early stopping with patience 15 on a validation set of 450 orders obtained AUC 0.834 on the 750-order test set (the same one as in 04-05, where logistic regression gives 0.844). The loop is zero_grad → backward → step, with BCEWithLogitsLoss and the sigmoid only for reading off probabilities.
Statement. Marta wants to know whether a better-tuned MLP beats logistic regression. Steps: (1) prepare the data exactly as in 05-03 (train_test_split with random_state=42, then 20 % of the training set for validation; the preprocessing is fitted on training data only); (2) write build_net(layers, dropout), which builds an MLP from any list of sizes; (3) write train with early stopping, returning the best epoch; (4) try at least six configurations varying layers ((16, 8), (32, 16), (64, 32, 16), (8,)), rate (0.01, 0.001, 0.0001) and dropout, and record in a table parameters, best epoch, validation loss and AUC, test AUC and seconds; (5) choose by validation, repeat the chosen one with 5 seeds and compare it with logistic regression trained on the same 1,799 orders.
Solution
import numpy as np, pandas as pd, torch, torch.nn as nn, time
from torch.utils.data import TensorDataset, DataLoader
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import Pipeline
from sklearn.metrics import roc_auc_score
from novamarket_ml import generate_orders_ml, dirty_orders, prepare_orders, build_preprocessing
X, y = prepare_orders(dirty_orders(generate_orders_ml(3000, 42), 42))
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.25, random_state=42, stratify=y)
Xtr2, Xval, ytr2, yval = train_test_split(Xtr, ytr, test_size=0.2, random_state=42, stratify=ytr)
prep = build_preprocessing().fit(Xtr2) # training data only
def to_tensor(Xd, yd):
return torch.tensor(prep.transform(Xd).astype(np.float32)), torch.tensor(yd.values.astype(np.float32)).view(-1, 1)
Xtr_t, ytr_t = to_tensor(Xtr2, ytr2); Xval_t, yval_t = to_tensor(Xval, yval); Xte_t, yte_t = to_tensor(Xte, yte)
def build_net(layers=(16, 8), dropout=0.2, n_in=21):
parts, n = [], n_in
for h in layers:
parts += [nn.Linear(n, h), nn.ReLU(), nn.Dropout(dropout)]; n = h
return nn.Sequential(*parts, nn.Linear(n, 1)) # no sigmoid: BCEWithLogitsLoss adds it
def evaluate(net, X_t, y_t):
net.eval()
with torch.no_grad():
logits = net(X_t)
return nn.BCEWithLogitsLoss()(logits, y_t).item(), roc_auc_score(y_t.numpy().ravel(), torch.sigmoid(logits).numpy().ravel())
def train(net, rate=0.001, epochs=300, batch=64, patience=15, seed=0):
torch.manual_seed(seed)
loader = DataLoader(TensorDataset(Xtr_t, ytr_t), batch_size=batch, shuffle=True)
opt, loss_fn = torch.optim.Adam(net.parameters(), lr=rate), nn.BCEWithLogitsLoss()
best_val, best_state, no_improvement, best_ep = float("inf"), None, 0, 0
for ep in range(1, epochs + 1):
net.train()
for xb, yb in loader:
opt.zero_grad(); loss_fn(net(xb), yb).backward(); opt.step()
L_val, _ = evaluate(net, Xval_t, yval_t)
if L_val < best_val - 1e-4:
best_val, no_improvement, best_ep = L_val, 0, ep
best_state = {k: v.clone() for k, v in net.state_dict().items()}
else:
no_improvement += 1
if no_improvement >= patience: break
net.load_state_dict(best_state)
return best_ep, ep
configs = [((16, 8), 0.001, 0.2), ((16, 8), 0.01, 0.2), ((16, 8), 0.0001, 0.2), ((32, 16), 0.001, 0.2),
((64, 32, 16), 0.001, 0.3), ((8,), 0.001, 0.0), ((16, 8), 0.001, 0.0)]
rows = []
for layers, rate, do in configs:
t = time.time(); torch.manual_seed(42); net = build_net(layers, do)
best_ep, last = train(net, rate=rate)
(Lv, auc_v), (_, auc_t) = evaluate(net, Xval_t, yval_t), evaluate(net, Xte_t, yte_t)
rows.append([str(layers), rate, do, sum(p.numel() for p in net.parameters()), best_ep, last, round(Lv, 4), round(auc_v, 3), round(auc_t, 3), round(time.time() - t, 1)])
table = pd.DataFrame(rows, columns=["layers", "rate", "dropout", "params", "best_ep", "last_ep", "val_loss", "AUC_val", "AUC_test", "sec"])
print(table.to_string(index=False))
best = table.sort_values("AUC_val", ascending=False).iloc[0]
aucs = []
for s in range(5):
torch.manual_seed(s); net = build_net(eval(best["layers"]), best["dropout"]); train(net, rate=best["rate"], seed=s)
aucs.append(evaluate(net, Xte_t, yte_t)[1])
print("Best by validation:", best["layers"], best["rate"], "| AUC test 5 seeds:", np.round(aucs, 3), round(np.mean(aucs), 3))
log = Pipeline([("prep", build_preprocessing()), ("model", LogisticRegression(max_iter=1000))]).fit(Xtr2, ytr2)
print("Logistic (same 1799): AUC val", round(roc_auc_score(yval, log.predict_proba(Xval)[:, 1]), 3), "| AUC test", round(roc_auc_score(yte, log.predict_proba(Xte)[:, 1]), 3)) layers rate dropout params best_ep last_ep val_loss AUC_val AUC_test sec
(16, 8) 0.0010 0.2 497 75 90 0.3241 0.835 0.834 3.1
(16, 8) 0.0100 0.2 497 11 26 0.3316 0.825 0.830 0.7
(16, 8) 0.0001 0.2 497 299 300 0.3423 0.821 0.831 8.1
(32, 16) 0.0010 0.2 1249 26 41 0.3317 0.829 0.839 1.2
(64, 32, 16) 0.0010 0.3 4033 17 32 0.3337 0.831 0.838 1.1
(8,) 0.0010 0.0 185 97 112 0.3232 0.843 0.836 2.5
(16, 8) 0.0010 0.0 497 49 64 0.3219 0.837 0.832 1.7
Best by validation: (8,) 0.001 | AUC test 5 seeds: [0.837 0.841 0.837 0.84 0.842] 0.84
Logistic (same 1799): AUC val 0.832 | AUC test 0.842Reading the table:
- The first row reproduces 05-03 (best epoch 75, test AUC 0.834). With rate 0.01 the network converges in 11 epochs but worse (0.825 on validation): steps too large; with 0.0001 it uses up all 300 epochs without having finished learning (0.821): the learning rate is the most sensitive hyperparameter, as 05-03 said.
- The larger networks (1,249 and 4,033 parameters) stop earlier (epoch 17-26) because they start overfitting straight away with 1,799 examples; their test AUC is slightly higher (0.838-0.839), but not their validation AUC, and choosing by test would be cheating.
- The best by validation is the smallest: a single layer of 8 neurons without dropout (185 parameters), 0.843 on validation and 0.840 ± 0.002 on test over 5 seeds. Logistic regression on the same data gives 0.842. All the differences are within ±0.005, the noise of a 750-order test set.
- Decision: the one from 05-03, reinforced. On small tabular data with an almost linear truth, a well-tuned MLP ties with logistic regression; no configuration beats it in a way that justifies the extra complexity. Marta keeps the logistic regression and files this table as evidence.
Feedback
- Typical mistake: choosing the configuration by looking at
AUC_test(here you would have picked (32, 16) with 0.839, and that figure would no longer be an honest estimate). Another one: not settingtorch.manual_seedbefore creating the network and attributing to the architecture what is initialisation luck; that is why the chosen one is repeated with 5 seeds. - Early stopping plays the role of the number of epochs: do not search for it by hand. And the validation loss is a better stopping criterion than the AUC (smoother).
- Variants: add
weight_decay=1e-4in Adam (L2 regularisation) and compare with dropout; trynn.BatchNorm1dbetween layers; measure the business cost at threshold 0.2 (09-02) of the best MLP against logistic regression; useoptunaorRandomizedSearchCVwithskorchif you want to automate the search (beyond the scope of the course).
- Project 2: incident-photo CNN with a third class and data augmentation
Reminder (05-04, sections 3-4). The synthetic 16×16 photos from generate_photos (bright box on a dark background; 1 = damaged with a diagonal crack or a dent) are classified by a CNN with two Conv2d → ReLU → MaxPool2d blocks and two dense layers (9,538 parameters, 98.7 % accuracy), while an MLP with more parameters stays at 57 %. CrossEntropyLoss combines softmax and cross-entropy for multiple classes.
Statement. Diego wants a third category: "wet box" (a damp stain: a large area of medium grey with diffuse edges, different from the small dark dent). Steps: (1) write generate_photos3(n=900, seed=42), which extends the generator reproducibly with classes 0 intact, 1 damaged, 2 wet; (2) print one photo of each class as a matrix of digits and check that they can be told apart by eye; (3) adapt the CNN to 3 outputs and train it for 30 epochs (675 training photos, 225 test); (4) compute the confusion matrix and list what it confuses; (5) write augment(xb, rng), which applies random flips and 90° rotations to each image in the batch (labels do not change), and compare accuracy with and without augmentation using 150 training photos (60 epochs) and using all 675, with 4 seeds; (6) compare with an MLP.
Solution
from sklearn.metrics import confusion_matrix
CLASSES = ["intact", "damaged", "wet"]
def generate_photos3(n=900, seed=42, size=16):
"""16x16 greyscale: bright box on a dark background. 0 = intact, 1 = damaged (crack/dent), 2 = wet (diffuse stain)."""
rng = np.random.default_rng(seed)
X, y = np.zeros((n, 1, size, size), dtype=np.float32), np.zeros(n, dtype=np.int64)
for i in range(n):
img = rng.normal(0.1, 0.05, (size, size))
x0, y0 = rng.integers(1, 5, size=2); x1, y1 = rng.integers(size - 5, size - 1, size=2)
img[y0:y1, x0:x1] = rng.normal(0.7, 0.05, (y1 - y0, x1 - x0))
cls = rng.integers(0, 3); y[i] = cls
if cls == 1: # same as in 05-04
if rng.random() < 0.5:
r0, c0 = rng.integers(y0, y1 - 4), rng.integers(x0, x1 - 4); length = rng.integers(4, min(y1 - r0, x1 - c0) + 1)
for k in range(length): img[r0 + k, c0 + k] = 0.05
else:
r0, c0 = rng.integers(y0, y1 - 3), rng.integers(x0, x1 - 3); side = rng.integers(2, 4)
img[r0:r0 + side, c0:c0 + side] = rng.normal(0.15, 0.05, (side, side))
elif cls == 2: # stain: 5-8 pixels, medium grey, soft edges
height, width = rng.integers(5, 9, size=2)
r0 = rng.integers(y0, max(y0 + 1, y1 - height)); c0 = rng.integers(x0, max(x0 + 1, x1 - width))
r1, c1 = min(r0 + height, y1), min(c0 + width, x1)
stain = rng.normal(0.45, 0.06, (r1 - r0, c1 - c0))
rr, cc = np.linspace(-1, 1, r1 - r0)[:, None], np.linspace(-1, 1, c1 - c0)[None, :]
weight = np.clip(1.2 - (rr ** 2 + cc ** 2), 0, 1) # strongest at the centre, fades out at the edge
img[r0:r1, c0:c1] = weight * stain + (1 - weight) * img[r0:r1, c0:c1]
X[i, 0] = np.clip(img, 0, 1)
return X, y
X, y = generate_photos3(); print(X.shape, np.bincount(y))
np.set_printoptions(linewidth=120)
print("wet:"); print((X[np.where(y == 2)[0][0], 0] * 9).astype(int))
Xtr, ytr, Xte, yte = X[:675], y[:675], X[675:], y[675:]
Xtr_t, ytr_t, Xte_t, yte_t = map(torch.tensor, (Xtr, ytr, Xte, yte))
def build_cnn(n_classes=3):
return nn.Sequential(nn.Conv2d(1, 8, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2), # 1x16x16 -> 8x8x8
nn.Conv2d(8, 16, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2), # -> 16x4x4
nn.Flatten(), nn.Linear(256, 32), nn.ReLU(), nn.Linear(32, n_classes))
def augment(xb, rng):
"""Random horizontal/vertical flip and 0/90/180/270° rotation, image by image."""
out = xb.clone()
for i in range(len(out)):
if rng.random() < 0.5: out[i] = torch.flip(out[i], dims=[2])
if rng.random() < 0.5: out[i] = torch.flip(out[i], dims=[1])
k = int(rng.integers(0, 4))
if k: out[i] = torch.rot90(out[i], k, dims=[1, 2])
return out
def accuracy(net, Xt, yt):
net.eval()
with torch.no_grad(): return (net(Xt).argmax(1) == yt).float().mean().item()
def train(net, Xt, yt, epochs=30, rate=0.003, batch=32, aug=False, seed=0, verbose=False):
torch.manual_seed(seed); rng = np.random.default_rng(seed)
opt, loss_fn = torch.optim.Adam(net.parameters(), lr=rate), nn.CrossEntropyLoss()
for ep in range(1, epochs + 1):
net.train(); perm = torch.randperm(len(Xt))
for i in range(0, len(perm), batch):
idx = perm[i:i + batch]; xb = augment(Xt[idx], rng) if aug else Xt[idx]
opt.zero_grad(); loss_fn(net(xb), yt[idx]).backward(); opt.step()
if verbose and ep in (1, 5, 10, 20, 30):
print(f" epoch {ep:2d} train accuracy {accuracy(net, Xt, yt):.3f} test {accuracy(net, Xte_t, yte_t):.3f}")
t = time.time(); torch.manual_seed(0); cnn = build_cnn(); print("parameters:", sum(p.numel() for p in cnn.parameters()))
train(cnn, Xtr_t, ytr_t, verbose=True); print(f"{time.time() - t:.1f} s")
with torch.no_grad(): pred = cnn(Xte_t).argmax(1).numpy()
print(confusion_matrix(yte, pred)); print("Test accuracy:", round((pred == yte).mean(), 3))
print("Errors:", [(CLASSES[a], "->", CLASSES[b]) for a, b in zip(yte[pred != yte], pred[pred != yte])])
for n_tr, ep in ((150, 60), (675, 30)):
for aug in (False, True):
accs = []
for s in range(4):
torch.manual_seed(s); net = build_cnn(); train(net, Xtr_t[:n_tr], ytr_t[:n_tr], epochs=ep, aug=aug, seed=s); accs.append(accuracy(net, Xte_t, yte_t))
print(f"n_train={n_tr:3d} epochs={ep} aug={aug!s:5s}: {np.round(accs, 3)} mean {np.mean(accs):.3f}")
torch.manual_seed(0); mlp = nn.Sequential(nn.Flatten(), nn.Linear(256, 64), nn.ReLU(), nn.Linear(64, 3)); train(mlp, Xtr_t, ytr_t)
print("MLP:", sum(p.numel() for p in mlp.parameters()), "parameters, test accuracy", round(accuracy(mlp, Xte_t, yte_t), 3))(900, 1, 16, 16) [298 293 309]
wet:
[[1 0 1 1 1 0 1 1 1 1 1 0 0 0 1 1]
[0 0 1 0 6 6 6 5 6 5 6 5 6 0 0 1]
[0 0 1 0 5 5 6 5 2 4 6 6 6 0 1 1]
[0 1 1 1 6 6 5 4 4 3 4 7 7 0 0 0]
[0 0 1 0 6 5 6 4 4 4 6 6 5 0 0 1]
[1 0 1 0 6 7 6 6 5 6 5 6 6 0 1 1]
...
parameters: 9571
epoch 1 train accuracy 0.493 test 0.458
epoch 5 train accuracy 0.630 test 0.600
epoch 10 train accuracy 0.647 test 0.627
epoch 20 train accuracy 0.963 test 0.911
epoch 30 train accuracy 0.990 test 0.978
1.1 s
[[70 0 1]
[ 1 74 3]
[ 0 0 76]]
Test accuracy: 0.978
Errors: [('damaged', '->', 'wet'), ('damaged', '->', 'intact'), ('damaged', '->', 'wet'), ('damaged', '->', 'wet'), ('intact', '->', 'wet')]
n_train=150 epochs=60 aug=False: [0.924 0.862 0.907 0.867] mean 0.890
n_train=150 epochs=60 aug=True : [0.951 0.92 0.916 0.911] mean 0.924
n_train=675 epochs=30 aug=False: [0.978 0.978 0.951 0.982] mean 0.972
n_train=675 epochs=30 aug=True : [0.978 0.973 0.942 0.964] mean 0.964
MLP: 16643 parameters, test accuracy 0.48Reading:
- The third class is harder than the two from 05-04: the stain (values 3-4 on a box of 6) is subtle, and the CNN goes from 9,538 to 9,571 parameters (33 more in the output layer) and takes longer to "get going" (epoch 10: 63 %; epoch 20: 91 %; epoch 30: 97.8 %). Almost all the errors are damaged → wet: a 3×3 dent and a small stain look alike. The MLP with 16,643 parameters stays at 48 %, close to chance with three classes (33 %): without convolution there is no invariance to position.
- Data augmentation: with only 150 photos, flips and rotations raise the mean accuracy from 89.0 % to 92.4 % and reduce the variance across seeds (0.86-0.92 → 0.91-0.95): the network sees each photo in 8 orientations and cannot memorise positions. With 675 photos it does not help (97.2 % versus 96.4 %): the natural variety already covers what augmentation invents. It is the rule from 05-03: augmentation is a regularisation, valuable when data is scarce.
- Note that the transformations must preserve the label: a rotated wet box is still wet; on the other hand, in a problem of reading labels or "this side up" arrows, a rotation would change the meaning.
Feedback
- Typical mistake: augmenting the test set too (you always evaluate on original images) or augmenting once at the start instead of in every epoch (you lose the variety). Another one: using
torch.flipon the wrong dimensions (here each image is(1, 16, 16): dims 1 and 2 are height and width). - With more epochs the CNN without augmentation would end up memorising; with early stopping on a validation set (project 1) you would avoid having to choose 30 by hand.
- Variants: add a fourth class "illegible label" (a small, very bright rectangle with noise) and see whether the confusion matrix degrades; visualise the 8 filters of the first layer (
cnn[0].weight) and check that one of them detects diagonal edges; train withn_tr=75and see how far augmentation gets you.
- Project 3: GRU for weekly demand against the baselines and the linear model
Reminder (05-04, section 5). A recurrent network processes a sequence step by step while keeping a hidden state; GRUs/LSTMs use gates to remember long dependencies. nn.GRU(input_size=1, hidden_size=h, batch_first=True) receives tensors of shape (batch, steps, 1) and returns the states of every step and the last one. In 09-02, the calendar linear model has MAE 16.8 on weeks 79-104 and the naive baseline around 30-39 depending on how it is defined.
Statement. Steps: (1) build windows of 12 weeks → next week (92 windows; the last 26, weeks 79-104, are the test set); (2) represent each window relatively (subtracting its last week, so that the network predicts the increment and does not depend on the level) and scale with a single deviation computed on training data only; (3) define GRUNet (GRU + linear layer on the last state) and train_gru with Adam and MSELoss on the full batch; (4) choose hidden size and number of epochs using the last 13 training windows as validation (weeks 66-78); (5) retrain with the 66 windows and evaluate on the test set with 5 seeds; (6) compare MAE with the naive baseline (last week), the 4-week mean, a linear regression on the window (12 coefficients) and the calendar linear model from 04-05, breaking down the error of week 100 (Black Friday) and week 101.
Solution
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_absolute_error
from novamarket_ml import generate_weekly_demand
d = generate_weekly_demand(104, 42); series = d["units"].values.astype(np.float32)
WINDOW, N_TEST, N_VAL = 12, 26, 13
def windows(series, window):
return np.array([series[t - window:t] for t in range(window, len(series))]), series[window:]
Xw, yw = windows(series, WINDOW) # 92 windows -> target weeks 13..104
target_week = np.arange(WINDOW + 1, len(series) + 1)
is_test = target_week > len(series) - N_TEST # weeks 79-104
last = Xw[:, -1] # last week of each window
Xr, yr = Xw - last[:, None], yw - last # relative: the network predicts the increment
scale = Xr[~is_test].std() # a single scale, from training data only
tensor = lambda a: torch.tensor(a / scale, dtype=torch.float32)
Xtr, ytr, Xte = tensor(Xr[~is_test]).unsqueeze(-1), tensor(yr[~is_test]).view(-1, 1), tensor(Xr[is_test]).unsqueeze(-1)
print("windows:", Xw.shape, "| training:", tuple(Xtr.shape), "| test:", tuple(Xte.shape))
class GRUNet(nn.Module):
def __init__(self, hidden=8):
super().__init__()
self.gru = nn.GRU(input_size=1, hidden_size=hidden, batch_first=True)
self.output = nn.Linear(hidden, 1)
def forward(self, x):
states, last_state = self.gru(x) # states: (batch, 12, hidden); last_state: (1, batch, hidden)
return self.output(last_state[0])
def train_gru(X_t, y_t, hidden=8, epochs=50, rate=0.01, seed=0):
torch.manual_seed(seed); net = GRUNet(hidden)
opt, loss_fn = torch.optim.Adam(net.parameters(), lr=rate), nn.MSELoss()
for _ in range(epochs): # full batch: 66 windows fit easily
net.train(); opt.zero_grad(); loss_fn(net(X_t), y_t).backward(); opt.step()
return net.eval()
def predict(net, X_t, last_week):
with torch.no_grad(): return net(X_t).numpy().ravel() * scale + last_week
Xa, ya, Xv = Xtr[:-N_VAL], ytr[:-N_VAL], Xtr[-N_VAL:] # validation: weeks 66-78
y_val, last_val = yw[~is_test][-N_VAL:], last[~is_test][-N_VAL:]
results = {}
for hidden in (8, 16, 32):
for epochs in (50, 100, 200, 400):
maes = [mean_absolute_error(y_val, predict(train_gru(Xa, ya, hidden, epochs, seed=s), Xv, last_val)) for s in range(3)]
results[(hidden, epochs)] = np.mean(maes)
print({k: round(float(v), 1) for k, v in results.items()}); print("Naive on validation:", round(mean_absolute_error(y_val, last_val), 1))
best = min(results, key=results.get); print("Best configuration:", best)
maes = [mean_absolute_error(yw[is_test], predict(train_gru(Xtr, ytr, *best, seed=s), Xte, last[is_test])) for s in range(5)]
print(f"GRU test, 5 seeds: {np.round(maes, 1)} mean {np.mean(maes):.1f}")
net = train_gru(Xtr, ytr, *best, seed=0); print("parameters:", sum(p.numel() for p in net.parameters()))
d["sin"], d["cos"] = np.sin(2 * np.pi * d["week"] / 52), np.cos(2 * np.pi * d["week"] / 52)
d["black_friday"] = ((d["week"] - 1) % 52 == 47).astype(int)
feats, tr, te = ["week", "sin", "cos", "black_friday"], d["week"] <= 78, d["week"] > 78
compare = {"GRU (seed 0)": predict(net, Xte, last[is_test]),
"Naive: last week": Xw[is_test][:, -1], "4-week mean": Xw[is_test][:, -4:].mean(1),
"Linear on the window (12 coef.)": LinearRegression().fit(Xw[~is_test], yw[~is_test]).predict(Xw[is_test]),
"Calendar linear (04-05)": LinearRegression().fit(d.loc[tr, feats], d.loc[tr, "units"]).predict(d.loc[te, feats])}
for name, p in compare.items():
err = pd.Series(np.abs(yw[is_test] - p), index=target_week[is_test])
print(f"{name:36s} MAE {err.mean():5.1f} | wk 100 (BF) {err[100]:5.1f} | wk 101 {err[101]:5.1f} | rest {err.drop([100, 101]).mean():4.1f} | bias {np.mean(yw[is_test] - p):+.1f}")windows: (92, 12) | training: (66, 12, 1) | test: (26, 12, 1)
{(8, 50): 15.9, (8, 100): 16.7, (8, 200): 25.5, (8, 400): 32.2, (16, 50): 16.3, (16, 100): 20.6, (16, 200): 19.9, (16, 400): 28.7, (32, 50): 19.3, (32, 100): 19.3, (32, 200): 19.5, (32, 400): 19.2}
Naive on validation: 18.7
Best configuration: (8, 50)
GRU test, 5 seeds: [33.9 29.3 32.2 34.3 30.1] mean 32.0
parameters: 273
GRU (seed 0) MAE 33.9 | wk 100 (BF) 212.9 | wk 101 118.7 | rest 22.9 | bias -12.7
Naive: last week MAE 34.3 | wk 100 (BF) 215.0 | wk 101 238.0 | rest 18.3 | bias -1.5
4-week mean MAE 31.4 | wk 100 (BF) 241.5 | wk 101 59.5 | rest 21.5 | bias -4.1
Linear on the window (12 coef.) MAE 29.8 | wk 100 (BF) 231.8 | wk 101 92.2 | rest 18.8 | bias -0.4
Calendar linear (04-05) MAE 16.8 | wk 100 (BF) 27.5 | wk 101 8.6 | rest 16.7 | bias -7.0Reading:
- On validation the GRU (15.9) seems to beat the naive baseline (18.7), and the long configurations get worse (400 epochs: 32.2), a sign of overfitting with 53 windows. But on the test set the GRU gives 32.0 ± 2 (5 seeds), the same as the naive baseline (34.3), the 4-week mean (31.4) and the linear model on the window (29.8), and a long way from the calendar linear model (16.8). Outside weeks 100 and 101 the GRU (22.9) is even worse than the naive baseline (18.3).
- Why it does not win: (1) it has 66 examples for 273 parameters; (2) its 12-week window is shorter than the seasonal period (52), so it cannot see the seasonality that the linear model receives directly through
sinandcos; (3) nobody tells it that week 100 is Black Friday (error 213, like every method that only looks at the past; and week 101 pays for having the spike inside the window), whereas the linear model knows through theblack_fridayvariable; (4) the noise in the series is independent from one week to the next, so there is no short-term dynamics to learn: the best it can do with the window is approximate persistence. Without the relative representation it was even worse: with absolute normalisation the network does not extrapolate the trend and systematically underestimates (bias +25 units). - With 4 or 8 years of synthetic data (
generate_weekly_demand(208),(416)) the GRU stays at 32-40 against 18-21 for the linear model: the problem is not only the size, it is that the information that matters (calendar) is not in the window. A recurrent network competes when there are rich temporal dependencies that the calendar does not capture (runs of promotions, stock-outs, campaign effects that drag on) and windows that cover the relevant cycle, or when it is also given the calendar variables as exogenous inputs. - Decision: the calendar linear model. And a lesson in method: the small validation (13 weeks) was misleading; with short series, comparison against baselines and against the reference model on the same test set is essential before believing an improvement.
Feedback
- Typical mistake: normalising with the mean and deviation of the whole series (leakage from the future) or shuffling the windows before separating the test set (test windows overlapping with training). Another one: passing
nn.GRUtensors of shape(batch, 12)without the feature dimension; you needunsqueeze(-1). MSELosson scaled data is the usual choice; the MAE is computed after undoing the scale and adding back the last week.- Variants: give the GRU a second channel with
black_fridayfor each week of the window and a third input with the target week's value (an exogenous variable known in advance); trynn.LSTM; use 52-week windows withgenerate_weekly_demand(416)(then it can see the yearly cycle) and check whether it gets closer to the linear model.
- Project 4: review sentiment with embeddings from scratch against the bag of words
Reminder (05-05, sections 3-4). The bag of words (CountVectorizer) + logistic regression obtains 0.767 accuracy in cross-validation on 56 reviews; embeddings represent each word as a learned dense vector; nn.EmbeddingBag(mode="mean") averages the vectors of the words in a sentence in a single pass and a linear layer decides. A pre-trained model (05-05, section 5) brings those vectors already learned from billions of sentences.
Statement. Steps: (1) write generate_reviews_extended(n=400, seed=42), which uses the templates from 05-05 and also returns the template identifier; (2) evaluate the bag of words + logistic regression with StratifiedKFold and with GroupKFold by template, and explain the difference; (3) write a tokeniser, a vocabulary and to_indices (concatenated indices + offsets, the EmbeddingBag format); (4) define EmbNet (EmbeddingBag of dimension 8 + linear) and train it for 100 epochs with Adam and weight_decay=1e-3; evaluate with the same two validations; (5) compare predictions on new sentences and look at which words have ended up most positive and most negative in the learned space; (6) discuss when a pre-trained model would pay off (without running it).
Solution
import re
from collections import Counter
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.model_selection import GroupKFold, StratifiedKFold, cross_val_score
from reviews_nm import POSITIVES, NEGATIVES, PRODUCTS
def generate_reviews_extended(n=400, seed=42):
rng = np.random.default_rng(seed); rows = []
for i in range(n):
sentiment = i % 2; pool = POSITIVES if sentiment else NEGATIVES
k = int(rng.integers(len(pool))); p = rng.choice(PRODUCTS)
rows.append({"review_id": f"R{1000 + i}", "text": pool[k].format(p=p, P=p[0].upper() + p[1:]),
"sentiment": sentiment, "template": f"{'P' if sentiment else 'N'}{k:02d}"})
return pd.DataFrame(rows).sample(frac=1, random_state=seed).reset_index(drop=True)
rev = generate_reviews_extended(); print(rev.shape, "distinct templates:", rev["template"].nunique())
bow = Pipeline([("bow", CountVectorizer()), ("log", LogisticRegression(max_iter=1000))])
skf, gkf = StratifiedKFold(5, shuffle=True, random_state=42), GroupKFold(5)
print("BoW, stratified CV:", cross_val_score(bow, rev["text"], rev["sentiment"], cv=skf).round(3))
sc = cross_val_score(bow, rev["text"], rev["sentiment"], cv=gkf, groups=rev["template"]); print("BoW, CV by template:", sc.round(3), round(sc.mean(), 3))
def tokenize(t): return re.findall(r"\w+", t.lower())
def build_vocab(texts):
vocab = {"<unk>": 0}
for w in Counter(w for t in texts for w in tokenize(t)): vocab[w] = len(vocab)
return vocab
def to_indices(texts, vocab):
idx, offsets = [], [0]
for t in texts:
toks = [vocab.get(w, 0) for w in tokenize(t)]; idx += toks; offsets.append(offsets[-1] + len(toks))
return torch.tensor(idx), torch.tensor(offsets[:-1])
class EmbNet(nn.Module):
def __init__(self, n_vocab, dim=8):
super().__init__(); self.emb = nn.EmbeddingBag(n_vocab, dim, mode="mean"); self.output = nn.Linear(dim, 1)
def forward(self, idx, off): return self.output(self.emb(idx, off))
def train_emb(texts, y, vocab, dim=8, epochs=100, rate=0.02, wd=1e-3, seed=0):
torch.manual_seed(seed); net = EmbNet(len(vocab), dim)
opt, loss_fn = torch.optim.Adam(net.parameters(), lr=rate, weight_decay=wd), nn.BCEWithLogitsLoss()
idx, off = to_indices(texts, vocab); y_t = torch.tensor(y, dtype=torch.float32).view(-1, 1)
for _ in range(epochs):
net.train(); opt.zero_grad(); loss_fn(net(idx, off), y_t).backward(); opt.step()
return net.eval()
def accuracy_emb(net, texts, y, vocab):
idx, off = to_indices(texts, vocab)
with torch.no_grad(): return float(((net(idx, off) > 0).numpy().ravel().astype(int) == np.asarray(y)).mean())
for name, cv, groups in (("stratified", skf, None), ("by template", gkf, rev["template"])):
accs = []
for tr, te in cv.split(rev["text"], rev["sentiment"], groups=groups):
vocab = build_vocab(rev["text"].iloc[tr]) # vocabulary from training data only
net = train_emb(rev["text"].iloc[tr].tolist(), rev["sentiment"].iloc[tr].values, vocab)
accs.append(accuracy_emb(net, rev["text"].iloc[te].tolist(), rev["sentiment"].iloc[te].values, vocab))
print(f"EmbeddingBag, CV {name}:", np.round(accs, 3), round(np.mean(accs), 3))
vocab = build_vocab(rev["text"]); net = train_emb(rev["text"].tolist(), rev["sentiment"].values, vocab)
print("vocabulary:", len(vocab), "| parameters:", sum(p.numel() for p in net.parameters()))
new_reviews = ["The coffee maker is great, very happy with the purchase", "The TV arrived broken and support does not answer",
"The robot vacuum is not quiet at all, a disaster", "Not bad at all, the air fryer works well and arrived fast",
"I expected it to be poor but the headphones are superb", "The laptop battery lasts hardly any time at all, disappointing"]
bow.fit(rev["text"], rev["sentiment"]); idx, off = to_indices(new_reviews, vocab)
with torch.no_grad(): p_emb = torch.sigmoid(net(idx, off)).numpy().ravel()
for t, pb, pe in zip(new_reviews, bow.predict_proba(new_reviews)[:, 1], p_emb): print(f"BoW {pb:.2f} Emb {pe:.2f} {t}")
E, inv = net.emb.weight.detach(), {i: w for w, i in vocab.items()}
score = (E @ net.output.weight[0].detach()).numpy(); order = np.argsort(score) # projection of each word onto the output direction
print("Most negative:", [inv[i] for i in order[:8]]); print("Most positive:", [inv[i] for i in order[-8:][::-1]])(400, 4) distinct templates: 56 BoW, stratified CV: [1. 1. 1. 1. 1.] BoW, CV by template: [0.85 0.638 0.675 0.8 0.725] 0.737 EmbeddingBag, CV stratified: [1. 1. 1. 1. 1.] 1.0 EmbeddingBag, CV by template: [0.888 0.588 0.75 0.738 0.7 ] 0.733 vocabulary: 267 | parameters: 2145 BoW 0.95 Emb 0.99 The coffee maker is great, very happy with the purchase BoW 0.05 Emb 0.02 The TV arrived broken and support does not answer BoW 0.15 Emb 0.21 The robot vacuum is not quiet at all, a disaster BoW 0.31 Emb 0.21 Not bad at all, the air fryer works well and arrived fast BoW 0.54 Emb 0.69 I expected it to be poor but the headphones are superb BoW 0.19 Emb 0.28 The laptop battery lasts hardly any time at all, disappointing Most negative: ['not', 'bad', 'does', 'came', 'noisy', 'no', 'late', 'expensive'] Most positive: ['that', 'excellent', 'works', 'well', 'price', 'good', 'great', 'perfectly']
Because the corpus is in English, the exact vocabulary size, probabilities and word lists may differ slightly from those shown here; the conclusions do not change.
Reading:
- Stratified validation gives 100 % for both models, and it is a lie: with 400 sentences generated from 56 templates, each template appears about 7 times (with different products) and the model sees in the test set sentences almost identical to the training ones. It is the leak from 04-03 in its text version, and that is why you have to validate by groups (
GroupKFoldby template): then the bag of words gives 0.737 (0.64-0.85 depending on the fold) and the embeddings 0.733. "Having n = 400" has added nothing over the 56 sentences of 05-05 (0.767), because the diversity of wordings is still 56. - The embeddings from scratch tie with the bag of words: 2,145 parameters learned from 56 formulations cannot capture more than the frequency of "not", "bad", "great", which the bag already counts. The most negative and most positive words in the learned space are almost the same as the logistic regression coefficients in 05-05; and both fail in the same way on sentences with negation or contrast ("Not bad at all", "I expected it to be poor but…"): averaging vectors or counting words ignores order.
- When a pre-trained model pays off: when the problem requires understanding new sentences that do not resemble the training ones (synonyms, irony, negations, spelling mistakes), and you do not have thousands of labelled examples of real diversity. A pre-trained encoder (for example, a multilingual BERT or a sentence embeddings model, 05-05) brings vectors in which "disappointing", "a rip-off" and "I don't recommend it" are already close together, and so are "great", "a hit" and "worth every euro"; with those fixed vectors and a logistic regression on top, 56 templates would be enough to exceed 0.9 on genuinely new sentences. The price: dependence on a model of hundreds of MB, more latency, and the obligation to check for inherited biases (02-04). With NovaMarket's real reviews (thousands, diverse), the first option today would be a fine-tuned pre-trained model; project 09-04 leaves it as a scripted alternative.
Feedback
- Typical mistake: building the vocabulary (or fitting
CountVectorizer) on all the sentences before cross-validation: it is a minor but real leak (test words sneak into the vocabulary). Another one: forgetting<unk>and havingto_indicesblow up on a new word. - If the
EmbeddingBagdoes not learn, checkoffsets(they must be the start positions of each sentence, without the last one) and the learning rate (0.02 here, high because the embedding gradients are sparse). - Variants: use
mode="sum"and compare; add bigrams to the tokeniser ("not working" as one unit) and see whether the sentences with negation improve (the binary bag with unigrams already rises to 0.775 by template); write 10 new reviews by hand, very different from the templates, and use them as the final test for both models.
Common Mistakes and Tips
- Choosing by test. In all four projects there is a validation set (or grouped cross-validation) for choosing architecture, epochs and rate; the test set is looked at once. If you choose by test, your final figure is optimistic and you do not know it.
- One seed is not a result. Small networks on small data vary by ±0.01 of AUC, ±3 points of accuracy and ±3 units of MAE across seeds. Repeat 3-5 times and report mean and range.
- Leaks specific to each data type: scaling series with statistics from the whole series; test windows overlapping with training; vocabulary fitted with the test set; templates repeated across folds; data augmentation applied to the test set.
- Comparing the network with a weak version of the classic. The calendar linear model, logistic regression with its pipeline and a well-validated bag of words are serious opponents; beating them with small data is the exception, not the rule (05-03).
- Forgetting
eval()and evaluating with dropout active, orzero_grad()and accumulating gradients; and withEmbeddingBag, mixing up indices and offsets. - Tip: every project ends with a table and a decision sentence, as in 09-02: "MLP ties (0.840 versus 0.842): logistic regression"; "CNN 97.8 % with three classes; augmentation useful with few photos"; "GRU 32 versus linear 16.8: linear"; "embeddings 0.733 versus BoW 0.737: BoW, and pre-trained if there is budget".
Conclusion
Four projects and a nuanced answer to the question from 05-03. The returns MLP, after seven configurations and five seeds, ties with logistic regression (0.840 versus 0.842): on small tabular data it does not pay off. The photo CNN absorbs a third class with 97.8 % accuracy (the errors are damaged → wet) against 48 % for a larger MLP, and augmentation by flips and rotations is worth it with 150 photos (89 → 92 %) and not with 675: on images the network is the tool, and augmentation its regularisation. The GRU on demand (MAE 32, like the baselines) loses to the calendar linear model (16.8) because its window sees neither the yearly cycle nor Black Friday and there is no short-term dynamics to learn: recurrent networks need real temporal dependencies and data. And the embeddings from scratch tie with the bag of words (0.733 versus 0.737) once validation by template unmasks the fictitious 100 %; the leap would come from a pre-trained model, not from more parameters. In all of them, the same discipline as in 09-01 and 09-02: honest validation, baselines and a written decision.
The last step of the module remains: in 09-04, Capstone Project, you will stop solving isolated exercises and apply the full method from 08-01 to a NovaMarket case that has not yet been closed end to end: the incident diagnosis of case 9, combining the Bayesian network from 06-03 with a classifier trained on synthetic incidents and the business rules from 06-04, from the definition document to the model card and the presentation to Diego; and with case 4 (reviews and response priority) as a scripted alternative.
Fundamentals of Artificial Intelligence (AI)
Module 1: Introduction to Artificial Intelligence
Module 2: Basic Principles of AI
- Fundamental Concepts: Agents, Environments and Rationality
- Types of Artificial Intelligence
- Data as the Raw Material of AI
- Ethics and Considerations in AI
Module 3: Algorithms in AI
- Introduction to Algorithms
- Search Algorithms
- Adversarial Search: Games and Minimax
- Optimization Algorithms
Module 4: Machine Learning
- Basic Concepts of Machine Learning
- Types of Machine Learning
- Data Preparation and Feature Engineering
- Machine Learning Algorithms
- Model Evaluation and Validation
- Overfitting, Regularization and Hyperparameter Tuning
Module 5: Neural Networks and Deep Learning
- Introduction to Neural Networks
- Neural Network Architecture
- How a Network Learns: Gradient Descent and Backpropagation
- Deep Learning and Its Applications
- Transformers, Large Language Models and Generative AI
Module 6: Logic and Expert Systems
- Logic in AI
- Expert Systems
- Reasoning under Uncertainty: Probability and Bayesian Networks
- Applications of Expert Systems
Module 7: Tools and Programming Languages in AI
- Programming Languages for AI
- Scientific Python: NumPy, pandas and Matplotlib
- Popular Tools and Libraries
- Development Environments
Module 8: Projects and Case Studies
Module 9: Exercises and Practice
- Algorithm Exercises
- Machine Learning Practice
- Neural Network Projects
- Capstone Project: from Idea to Prototype
