The heuristic recommender from the previous lesson is already in production, and it works. But AlpinaShop has two questions that SQL cannot answer.
The first one comes from Marta, who runs the website: out of every hundred carts that get filled, which ones are going to end in a purchase? If you knew in advance, you could reserve the welcome discount for the person who is about to abandon instead of giving it away to someone who was going to buy anyway. That is not a pattern you read off a co-occurrence table: it depends on a dozen signals that interact with each other.
The second one comes from the catalogue: 60 GB of unlabelled product images. Nobody knows, looking at the alpinashop-catalogo bucket, which ones are backpacks, which ones ice axes and which ones jackets, beyond whatever the product page says — which does not always match the photo the supplier uploaded.
Both questions need a trained model. And none of the three people on the team is a machine learning engineer.
AutoML exists for exactly this: training reasonably good models on your own data, without writing modelling code, letting the platform take the technical decisions. This lesson is about what it does underneath, how far it goes, where it falls short, and — the part most often neglected — how to read the results without fooling yourself.
Contents
- What AutoML is and what it really does underneath
- Who it is for and what its limits are
- Problem types and minimum data requirements
- Case 1: predicting a cart's purchase
- The temporal split: why a random one would be cheating
- Training budget and cost
- Reading the results without fooling yourself
- The decision threshold is a business decision
- Feature importance and explainability
- Case 2: classifying the catalogue photos
- Labelling well: the cost nobody budgets for
- Deploying: endpoint or batch prediction
- The three traps: leakage, imbalance and overfitting
- When AutoML is not the answer
- Bias, automated decision-making and the legal framework
- What AutoML is and what it really does underneath
AutoML is not magic, nor is it a specific algorithm. It is the automation of the tasks an ML engineer would do by hand, executed systematically and with more patience than any person has.
Four things it does for you:
Feature engineering. It detects the type of each column, normalises the numeric ones, encodes the categorical ones, extracts components from dates (day of the week, month, whether it is a public holiday), imputes missing values and discards useless columns — the ones with a single value or the ones that are a unique identifier per row.
Architecture and algorithm search. It tries different model families — boosted trees, neural networks, linear models — and, within each one, different configurations. For image and text it also searches for the network architecture through neural architecture search.
Hyperparameter tuning. It optimises learning rates, depths, regularisation and the rest, with the same kind of Bayesian search you saw in 05-01, but without you having to declare it.
Ensembling. The final model is rarely a single one: it is usually a weighted combination of several, because combining diverse models almost always beats the best individual one.
flowchart TD
A[Your labelled data] --> B[Automatic analysis and cleaning]
B --> C[Feature engineering]
C --> D[Architecture and algorithm search]
D --> E[Hyperparameter tuning]
E --> F[Evaluation on the test set]
F -->|budget left| D
F -->|budget exhausted| G[Ensembling of the best model]
G --> H[Model in Model Registry]
What it does not do, and it is worth being clear about from the start:
- It does not decide which problem to solve or what success means.
- It does not obtain data, nor fix data that is wrong.
- It does not detect that you have slipped in a column that leaks the future.
- It does not know what is fair, nor what the consequences of being wrong are.
All of that is still yours. And it is, as it happens, where most of the value and most of the risk sit.
- Who it is for and what its limits are
AutoML has a very well-defined audience: whoever knows the domain and the data but not model engineering. Lucía is the exact case. She knows what a session is, what an abandoned cart is and what every column of visitas means. She does not know how to choose between XGBoost and a neural network, and she does not need to.
| Advantages | Real limits |
|---|---|
| A decent result in hours, not weeks | Usually falls short of a good custom model |
| Requires no knowledge of modelling | Little control over preprocessing |
| Evaluation and explainability included | Higher training cost per result |
| Integrated with Model Registry and endpoints | The model is a fairly closed box |
| A good baseline to compare against | No custom architectures or loss functions |
And one function that is almost never mentioned but is among the most valuable: AutoML is an excellent detector of badly framed problems. If AutoML, with all its effort, cannot get anything better than chance, it is very likely that the signal is not in the data, and no custom model is going to invent it. And if AutoML gets 99.8 % on the first attempt, you almost certainly have data leakage. In both cases it has saved you weeks.
- Problem types and minimum data requirements
| Data type | Supported problems | Technical minimum | Reasonable minimum in practice |
|---|---|---|---|
| Tabular | Binary and multiclass classification, regression, forecasting | 1,000 rows | 10,000+ rows, and ≥100 examples of the minority class |
| Image | Classification (single or multi-label), object detection | 10 images per label | 100–500 per label, with real variety |
| Text | Classification, entity extraction, sentiment analysis | 20 examples per label | 100–1,000 per label |
| Video | Classification, action recognition, tracking | Dozens of clips | Hundreds, and well trimmed |
The difference between the last two columns matters a great deal. The technical minimum is what the platform accepts; the reasonable minimum is what produces a usable model. With 10 images per label AutoML trains and gives you a number, but that model is of no use whatsoever in production.
And a rule that holds for the whole module: variety matters more than quantity. Five hundred photos of backpacks all taken against the same white background with the same lighting teach the model to recognise that background, not the backpack. Two hundred varied photos — different backgrounds, angles, lighting — generalise far better.
A compulsory check before training anything:
SELECT
COUNT(*) AS filas,
COUNTIF(compro = 1) AS compraron,
ROUND(100 * COUNTIF(compro = 1) / COUNT(*), 2) AS pct_positivos,
COUNT(DISTINCT sesion_id) AS sesiones_unicas,
MIN(fecha) AS desde,
MAX(fecha) AS hasta,
COUNTIF(importe_carrito IS NULL) AS sin_importe
FROM `alpinashop-datos.alpinashop_analitica.ml_sesiones`;If pct_positivos is 0.4 %, if sesiones_unicas is lower than filas (there are duplicates) or if sin_importe is half the set, you have work to do before training. Training first and discovering it afterwards costs machine hours.
- Case 1: predicting a cart's purchase
The goal, written as a business decision and not as a technical problem: estimate the probability that a session with a cart ends in a purchase, in order to decide who gets shown a closing incentive.
The starting table is ml_sesiones, the one you built in 05-01 by crossing visitas with v_pedidos_analitica. We are going to enrich it with customer and cart context, taking care that everything is information available at the instant of the decision:
CREATE OR REPLACE TABLE `alpinashop-datos.alpinashop_analitica.automl_carritos` AS
SELECT
v.sesion_id,
v.fecha,
-- Session context
v.dispositivo,
v.canal,
v.paginas_vistas,
v.minutos_sesion,
v.productos_vistos,
-- Cart context
v.unidades_carrito,
ROUND(v.importe_carrito, 2) AS importe_carrito,
ROUND(v.importe_carrito / NULLIF(v.unidades_carrito,0), 2) AS precio_medio_articulo,
c.categoria_principal,
-- Temporal context
EXTRACT(DAYOFWEEK FROM v.fecha) AS dia_semana,
EXTRACT(HOUR FROM v.hora_inicio) AS hora_dia,
-- Customer context, computed BEFORE this session
IFNULL(h.pedidos_previos, 0) AS pedidos_previos,
IFNULL(ROUND(h.ticket_medio_previo, 2), 0) AS ticket_medio_previo,
IFNULL(h.dias_desde_ultimo_pedido, 999) AS dias_desde_ultimo,
-- Label
v.compro
FROM `alpinashop-datos.alpinashop_analitica.ml_sesiones` v
LEFT JOIN `alpinashop-datos.alpinashop_analitica.carrito_categorias` c
ON c.sesion_id = v.sesion_id
LEFT JOIN `alpinashop-datos.alpinashop_analitica.hist_cliente_diario` h
ON h.cliente_hash = v.cliente_hash
AND h.fecha = DATE_SUB(v.fecha, INTERVAL 1 DAY)
WHERE v.unidades_carrito > 0;The detail that decides whether this model is any use is in the last JOIN: h.fecha = DATE_SUB(v.fecha, INTERVAL 1 DAY). The customer's history is taken from the day before the session, not from today. If it were joined with the current history, every training row would carry information about orders placed after the session being predicted. The model would learn that "customers with many orders buy", which is true and also useless: in production, at the moment of deciding, that future order does not exist.
This kind of table — a snapshot of the state of each entity on each date — is called a historical feature table and it is what prevents almost every temporal leak. It is exactly what a Feature Store stores for you (05-01, section 7).
Creating the dataset and launching training:
gcloud ai datasets create \
--project=alpinashop-datos --region=europe-west1 \
--display-name=carritos-conversion \
--metadata-schema-uri="gs://google-cloud-aiplatform/schema/dataset/metadata/tabular_1.0.0.yaml" \
--metadata="{\"inputConfig\":{\"bigquerySource\":{\"uri\":\"bq://alpinashop-datos.alpinashop_analitica.automl_carritos\"}}}"From the console, the flow asks for: target column (compro), objective type (binary classification), columns to exclude (sesion_id and fecha — the first is an identifier, the second defines the split and must not go in as a feature), split method, budget and optimisation metric.
In Python it comes out more explicit and, above all, reproducible:
from google.cloud import aiplatform
aiplatform.init(project="alpinashop-datos", location="europe-west1")
ds = aiplatform.TabularDataset("projects/.../datasets/1234567890")
job = aiplatform.AutoMLTabularTrainingJob(
display_name="carritos-conversion-v1",
optimization_prediction_type="classification",
optimization_objective="maximize-au-prc", # PR-AUC, not accuracy
)
model = job.run(
dataset=ds,
target_column="compro",
predefined_split_column_name="conjunto", # our own temporal split
budget_milli_node_hours=2000, # 2 node-hours
model_display_name="automl-carritos-v1",
disable_early_stopping=False,
)Two choices that deserve justification:
maximize-au-prcinstead of accuracy. With imbalanced classes, the area under the precision-recall curve reflects what you care about — finding the positives — far better than the percentage of correct answers. We come back to this in section 7.disable_early_stopping=False: if the model stops improving, AutoML halts and does not charge you the remaining budget. Turning it off only makes sense in very specific cases.
- The temporal split: why a random one would be cheating
By default AutoML splits into 80 % training, 10 % validation and 10 % test, at random. For AlpinaShop's problem that is wrong, and it is worth understanding properly because it is the most repeated conceptual mistake.
Imagine that on 14 February AlpinaShop launches a campaign and conversions shoot up. With a random split, some 14 February sessions end up in training and others in test. The model, while training, sees how that particular day behaved and learns to recognise it. When evaluated on the other sessions of the same day, it gets them spectacularly right.
That result will never reproduce in production, because in production the model predicts on days it has never seen. A random split measures "can it interpolate within the known?" when the real question is "can it extrapolate to the unknown?".
The rule: if the problem has a temporal dimension, the split must be temporal. You train with the past, validate with the immediate past and evaluate with the future.
CREATE OR REPLACE VIEW `alpinashop-datos.alpinashop_analitica.automl_carritos_split` AS
SELECT
*,
CASE
WHEN fecha < DATE '2025-11-01' THEN 'TRAIN'
WHEN fecha < DATE '2026-01-01' THEN 'VALIDATE'
ELSE 'TEST'
END AS conjunto
FROM `alpinashop-datos.alpinashop_analitica.automl_carritos`
WHERE fecha BETWEEN DATE '2024-01-01' AND DATE '2026-03-31';Vertex AI accepts this column with predefined_split_column_name="conjunto", and the values must be exactly TRAIN, VALIDATE and TEST.
There is a side effect worth anticipating: the metrics will drop. That is normal and it is good. If with a random split the AUC was 0.91 and with a temporal split it is 0.83, the 0.83 is the real number. The 0.91 was an illusion, and finding that out now is far cheaper than finding it out after deployment.
One further nuance: the test period must cover at least one complete business cycle. In a mountain gear shop, two winter months do not represent the summer. If you can, it is worth also evaluating over a seasonally different period.
- Training budget and cost
The budget is expressed in node-hours and limits how much AutoML explores. It is given in thousandths: budget_milli_node_hours=2000 is 2 node-hours.
| Situation | Indicative budget | Comment |
|---|---|---|
| First trial, is there signal? | 1 node-hour | Cheap and enough to rule it out |
| Working tabular model | 3–6 node-hours | The usual point of diminishing returns |
| Large, complex dataset | 10–20 node-hours | Only if the short trial looked promising |
| Image | Several hours | Higher hourly cost than tabular |
The price per node-hour varies by data type and by region and changes over time: always check the current official documentation. As an order of magnitude, a tabular training run of a few hours runs to tens of euros; an image one with many hours can reach several hundred.
Three pieces of advice that genuinely save money:
- Always start with 1 node-hour. If with that budget the AUC is 0.52 — that is, chance — spending twenty times more will not fix it: the problem is in the data.
- Leave early stopping enabled. You do not pay for what is not used.
- Before spending, compare against BigQuery ML. A logistic regression costs cents and on simple tabular problems comes surprisingly close. If AutoML improves on BigQuery ML by two AUC points, the legitimate question is whether those two points are worth the difference in cost and in opacity.
- Reading the results without fooling yourself
This is where the person who uses AutoML is separated from the person who understands it.
Suppose the cart model's result on the test set, with 10,000 sessions of which 1,200 ended in a purchase (12 %):
| Predicted purchase | Predicted no purchase | |
|---|---|---|
| Actually bought | 780 (TP) | 420 (FN) |
| Did not buy | 890 (FP) | 7,910 (TN) |
The metrics that come out of it:
| Metric | Formula | Value | What it means here |
|---|---|---|---|
| Accuracy | (TP+TN)/total | 86.9 % | Misleading: saying "nobody buys" would give 88 % |
| Precision | TP/(TP+FP) | 46.7 % | Of every 100 I give the incentive to, 47 were going to buy |
| Recall | TP/(TP+FN) | 65.0 % | I detect 65 out of every 100 real buyers |
| F1 | harmonic mean | 54.5 % | Balance between the previous two |
| ROC-AUC | — | 0.83 | Probability of ranking a positive above a negative correctly |
| PR-AUC | — | 0.51 | The honest metric with imbalanced classes |
Accuracy is the metric that misleads most, and by a wide margin. In this case, a model that always answered "no purchase" would have 88 % accuracy — more than ours — and would be completely useless. Every time you see someone boasting about accuracy on an imbalanced problem, ask what the percentage of the majority class is.
ROC-AUC and PR-AUC are not interchangeable. ROC-AUC incorporates the true negatives, which here are 7,910 and extremely abundant, and that inflates it: 0.83 sounds good. PR-AUC only looks at what happens with the positives, and its 0.51 is a far more faithful description of the real difficulty. With imbalanced classes, always look at PR-AUC.
And the reference you must never forget: how much would a silly rule give? If ranking sessions by cart amount already identified 55 % of the buyers, the model contributes ten points, not sixty-five.
- The decision threshold is a business decision
A classifier does not return "yes" or "no": it returns a probability between 0 and 1. Turning it into a decision requires a threshold, and that threshold is not chosen by the model. You choose it, and it is an economic decision.
For AlpinaShop, the incentive is a 10 % discount. The numbers per session:
- True positive: I give the discount to someone who was going to buy → I lose 10 % of the margin unnecessarily. On an average cart of €95, about –€4.
- False positive: I give the discount to someone who was not going to buy → if the discount convinces a fraction of them, I gain; if not, it costs nothing because there is no sale. Expected value slightly positive.
- False negative: I do not give the discount to someone who was going to abandon → I lose a sale I could have recovered, about –€20 of expected margin.
- True negative: I do not give a discount to someone who was not going to buy → €0.
With those numbers, the expensive error is the false negative, and that pushes the threshold downwards: it pays to be generous handing out incentives. The trade-off table:
| Threshold | Sessions with incentive | Precision | Recall | Business reading |
|---|---|---|---|---|
| 0.20 | 4,100 | 26 % | 89 % | Almost everyone gets a discount; margin is given away |
| 0.35 | 2,400 | 38 % | 76 % | A reasonable compromise |
| 0.50 | 1,670 | 47 % | 65 % | The default, not necessarily the best |
| 0.70 | 620 | 68 % | 35 % | Very selective; two thirds are lost |
The right way to choose is to compute the expected profit of each threshold with the business values, not to look at which one gives the best F1:
WITH escenarios AS (
SELECT
umbral,
COUNTIF(prob >= umbral AND compro = 1) AS vp,
COUNTIF(prob >= umbral AND compro = 0) AS fp,
COUNTIF(prob < umbral AND compro = 1) AS fn
FROM `alpinashop-datos.alpinashop_analitica.predicciones_test`,
UNNEST([0.2, 0.3, 0.35, 0.4, 0.5, 0.6, 0.7]) AS umbral
GROUP BY umbral
)
SELECT
umbral, vp, fp, fn,
ROUND(vp * -4.0 + fp * 0.5 + fn * -20.0, 0) AS beneficio_estimado_eur
FROM escenarios
ORDER BY beneficio_estimado_eur DESC;This query turns a technical discussion into a table that management understands. And it makes explicit something that usually stays implicit: the coefficients -4, 0.5 and -20 are business hypotheses, not truths. Writing them down forces you to discuss them, which is exactly what has to happen.
One final warning about the threshold: it can be different per segment. For new customers, where the acquisition value is higher, a lower threshold may be worth it. That is no longer a model decision, it is product design.
- Feature importance and explainability
AutoML returns two levels of explanation, and they serve different purposes.
Global importance: which features weigh most in the model as a whole. For the cart model, a plausible result:
| Feature | Importance | Interpretation |
|---|---|---|
pedidos_previos |
0.24 | Whoever has bought before buys again |
minutos_sesion |
0.19 | Long sessions indicate intent |
importe_carrito |
0.15 | Expensive carts are abandoned more |
dias_desde_ultimo |
0.12 | Recency as a signal of engagement |
canal |
0.10 | Brand search converts more than display |
dispositivo |
0.07 | Mobile converts less than desktop |
hora_dia |
0.05 | Marginal |
dia_semana |
0.03 | Almost irrelevant |
How to read this table, and how not to. Importance is association, not causation. The fact that minutos_sesion weighs a lot does not mean that artificially lengthening the session increases purchases: it means that whoever is going to buy tends to spend more time. Confusing the two produces absurd product decisions.
What you should do with this table is a sanity check. If a feature turned up with an importance of 0.80 and all the others at zero, that is an alarm signal: that column probably leaks the outcome. And if all the importances were similar and low, that is a sign that no feature carries real signal.
Local attributions: why the model predicted what it predicted for one specific row. They are requested in the prediction call with explain instead of predict, and they return each feature's contribution to that individual prediction. They are indispensable when a decision has to be justified to a person, and they will return in 05-07 as part of the responsible ML checklist.
- Case 2: classifying the catalogue photos
The second problem is different in nature: unstructured input, 60 GB of images in alpinashop-catalogo under productos/<sku>/original|web|thumb/.
The goal: classify each photo into AlpinaShop's own taxonomy — backpack, ice axe, helmet, jacket, boot, crampon, rope, sleeping bag — in order to detect product pages with wrongly assigned images and to enrich the catalogue metadata.
Why the Cloud Vision API is not enough. The pre-trained API (lesson 05-05) recognises "backpack" perfectly, but it does not tell a 60-litre trekking backpack from a 25-litre summit pack, and it will call a technical ice axe a "hatchet". AutoML Vision exists precisely for when the taxonomy is your own and the general vocabulary does not cover it.
The input data is declared in an index CSV in Cloud Storage:
TRAIN,gs://alpinashop-catalogo/productos/MOC-4471/web/01.jpg,backpack
TRAIN,gs://alpinashop-catalogo/productos/PIO-1120/web/01.jpg,ice-axe
VALIDATE,gs://alpinashop-catalogo/productos/CAS-8802/web/02.jpg,helmet
TEST,gs://alpinashop-catalogo/productos/BOT-3391/web/01.jpg,bootThree columns: set, image URI and label. If the first column is omitted, Vertex AI splits for you.
Here the split can be random — there is no temporal dimension — but with one critical condition: all the photos of the same SKU must fall in the same set. A product usually has five or six almost identical photos; if some go to training and others to test, the model memorises that particular product and the evaluation lies. Grouping by entity is to the image problem what the temporal split is to the tabular one.
job = aiplatform.AutoMLImageTrainingJob(
display_name="catalogo-categorias-v1",
prediction_type="classification",
multi_label=False,
model_type="CLOUD", # served on an endpoint; EDGE to export
base_model=None,
)
image_model = job.run(
dataset=image_ds,
budget_milli_node_hours=8000, # 8 node-hours
training_filter_split=None,
model_display_name="automl-catalogo-v1",
disable_early_stopping=False,
)model_type="CLOUD" produces a model optimised for serving on Vertex AI. "EDGE" generates an exportable variant (TensorFlow Lite, container) to run outside, smaller and somewhat less accurate. AlpinaShop does not need edge: its images are already in the cloud.
- Labelling well: the cost nobody budgets for
This section is the most boring in the lesson and the one that sinks the most projects.
To train AutoML Vision you need labelled images, and AlpinaShop's 60 GB are not. The real numbers of labelling:
- 8 categories × 300 images = 2,400 images to label.
- At around 6 seconds per image with a good interface, that is about 4 hours of human work.
- Plus the time to agree the taxonomy, which always takes longer than expected.
And the problem is not the time, it is the consistency. The questions that come up half an hour in:
- A photo showing a backpack with an ice axe strapped to it: what is it?
- A climbing helmet and a ski helmet: are they the same label?
- A close-up of the stitching on a jacket: is that "jacket" or is it discarded?
- A pack of rope + carabiners: which label does it get?
If those decisions are not written down beforehand and are not consistent, the model learns the inconsistency. And a model trained with contradictory labels cannot exceed the quality of its labels: it is a hard ceiling.
The minimum protocol, which costs an afternoon and saves the project:
- Write the labelling guide with the definition of each category and at least one borderline case resolved per category.
- An
otroslabel for whatever does not fit. Without it, people force labels and contaminate the set. - Double-labelling of a sample: two people label the same 200 images. If they agree on less than 90 %, the taxonomy is ambiguous and has to be fixed before going on.
- Start with the
web/photos, which are normalised, before theoriginal/ones.
Vertex AI offers labelling workflows in the console and managed labelling services with human labellers. Be careful with the latter if the images contain anything sensitive: sending customer photos to an external labelling service has privacy implications that must be assessed beforehand. AlpinaShop's catalogue photos do not have them; the ones customers upload with their reviews do.
- Deploying: endpoint or batch prediction
The same dilemma as in 05-01, and the answer is once again the same for the same reason.
For the image model, the task is to classify 60 GB of photos once, and after that only the new ones that come in. That is a textbook case of batch prediction:
gcloud ai batch-prediction-jobs create \
--project=alpinashop-datos --region=europe-west1 \
--display-name=clasifica-catalogo-inicial \
--model=MODEL_ID \
--input-paths-uri=gs://alpinashop-datalake/entradas/catalogo_imagenes.jsonl \
--input-format=jsonl \
--output-uri-prefix=gs://alpinashop-datalake/salidas/catalogo-clases/ \
--output-format=jsonlAnd the results are loaded into the governed dataset to cross them with the catalogue:
CREATE OR REPLACE TABLE `alpinashop-datos.alpinashop_analitica.imagenes_clasificadas` AS
SELECT
REGEXP_EXTRACT(instance.content, r'productos/([^/]+)/') AS sku,
instance.content AS uri_imagen,
prediction.displayNames[OFFSET(0)] AS categoria_predicha,
ROUND(prediction.confidences[OFFSET(0)], 3) AS confianza
FROM `alpinashop-datos.alpinashop_analitica.raw_predicciones_imagen`;-- Product pages with a possibly wrong image
SELECT i.sku, p.categoria AS categoria_ficha,
i.categoria_predicha, i.confianza, i.uri_imagen
FROM `alpinashop-datos.alpinashop_analitica.imagenes_clasificadas` i
JOIN `alpinashop-datos.alpinashop_analitica.productos` p USING (sku)
WHERE i.categoria_predicha != p.categoria
AND i.confianza > 0.85
ORDER BY i.confianza DESC;This query is the real product of all the work. It is not a deployed model: it is a list of product pages a person can review, ordered by confidence. A model that produces a prioritised work list contributes more value with less risk than one that acts on its own.
For new photos, classification is triggered from the imagenes-subidas Pub/Sub topic, with the implementation that arrives in 06-03 with Cloud Functions.
- The three traps: leakage, imbalance and overfitting
Data leakage. A feature contains information that in production will not be available at prediction time. Unmistakable symptom: metrics that are too good. At AlpinaShop, the candidates are metodo_pago (it only exists if there was a purchase), direccion_envio_confirmada, codigo_descuento_aplicado and any customer history computed as of today instead of as of the session date. The smell test is always the same question: did this piece of data exist and was it known at the exact instant of the decision? If the answer is "more or less", then it did not.
Class imbalance. When a class is rare, the model learns to ignore it. With 12 % conversion the problem is moderate; with 0.3 % fraud it is serious. Remedies: choose PR-AUC as the objective, weight the classes, and in extreme cases undersample the majority class — never the test set, which must keep the real proportion so the metrics mean something. And sometimes the right remedy is to reframe: instead of classifying, rank by risk and review the top N.
Overfitting. The model memorises instead of generalising. Symptom: excellent on training, mediocre on test. AutoML controls it fairly well with regularisation and early stopping, but it cannot protect you from two causes that depend on you: too little data for the complexity of the problem, and leakage between sets — the same SKU or the same customer in training and in test.
| Trap | Symptom | Usual cause | What to do |
|---|---|---|---|
| Leakage | Suspicious AUC > 0.97 | A column from the future | Audit every column with the instant question |
| Imbalance | Tiny recall | Rare class | PR-AUC, weights, reframe as ranking |
| Overfitting | Test much worse than train | Little data or leakage | Group by entity, more data, less complexity |
- When AutoML is not the answer
| Criterion | Pre-trained API | BigQuery ML | AutoML | Custom training |
|---|---|---|---|---|
| Own data needed | None | Yes, in BigQuery | Yes, labelled | Yes, labelled |
| Time to result | Minutes | Hours | Hours or days | Weeks |
| Knowledge required | Calling an API | SQL | Interface + judgement | ML and engineering |
| Control over the model | None | Low | Low | Total |
| Training cost | 0 | Very low | Medium-high | High |
| Cost per prediction | Per unit | Very low | Medium | Variable |
| Quality ceiling | The provider's | Medium | High | The highest |
| When to choose it | The problem is generic | Tabular and already in BigQuery | Own taxonomy, no ML team | Requirements nothing else covers |
AutoML is not the answer when:
- The problem is generic. "Does this image contain a person?" or "is this text positive?" are solved by the pre-trained APIs (05-04, 05-05) with no data, no training and for cents.
- You have no labelled data and you are not going to invest in labelling it. Without labels there is no AutoML.
- The data is tabular and already lives in BigQuery. Try BigQuery ML first: it is hours against days and cents against tens of euros.
- You need your own architecture or loss function. That is the case of the two-tower recommender in 05-03.
- A heuristic already solves 80 %. The DA-003 lesson from 05-01 still stands.
- You need to run the model outside the cloud with strict requirements. EDGE mode helps, but your own model gives you more control over size and latency.
- Bias, automated decision-making and the legal framework
The two models in this lesson have very different risk profiles, and seeing them side by side is the best way to understand the framework.
The image model is low risk. It classifies inanimate objects. The worst possible error is labelling an ice axe as a hatchet, and the remedy is for a person to correct it in a review list.
The cart model does process personal data and does affect people. And there you have to think before deploying:
Bias. The model learns from what happened. If historically customers coming in from mobile convert less — because the mobile site is worse — the model will learn not to offer them incentives, so they will convert even less, so the model will confirm itself. It is a feedback loop: the model does not describe reality, it manufactures it. The way to detect it is to measure performance by segment — device, channel, country, new or returning customer — and not only in aggregate. A global AUC of 0.83 can hide a 0.86 on desktop and a 0.61 on mobile.
Personal data and GDPR. Even if it is trained over v_pedidos_analitica with the customer pseudonymised, pseudonymisation is not anonymisation: it is still personal data for GDPR purposes and the legal basis, minimisation, purpose limitation and retention periods still apply. The rights of access, objection and erasure apply too, which means being able to withdraw a customer's data from the next cycle's training set.
Automated decision-making. Showing or hiding a commercial discount is not the same as denying credit. But the boundary of Article 22 of the GDPR — decisions based solely on automated processing with significant effects — is not always obvious, and systematic price differentiation between customers can come closer to it than it seems.
The EU AI Act. The AI Regulation classifies systems by risk level, and from that classification follow obligations on risk management, data quality, technical documentation, record-keeping, transparency and human oversight. A commercial recommendation system usually sits at a low level, but the classification depends on the specific use, and the use can drift over time without anyone reviewing it.
Express recommendation. Before deploying the cart model, a compliance professional or the DPO must determine the legal basis for the processing, whether the intended use constitutes automated decision-making under Article 22, what information has to be given to customers, the retention period for the training set, and the system's classification under the AI Act with the associated obligations. This lesson describes technical controls and does not constitute legal advice. All data is fictitious.
And one technical measure that is worth more than many declarations: evaluate by segment and publish that table alongside the global metric. If nobody looks at performance by group, the bias is not detected.
Common Mistakes and Tips
Trusting accuracy. With imbalanced classes it is the metric that misleads most. Look at PR-AUC, precision and recall, and always compare them against the percentage of the majority class.
Leaving the split random on a temporal problem. It produces inflated metrics that do not reproduce in production.
Letting photos of the same product fall into different sets. Pure leakage: the model memorises the product and the evaluation lies. Group by entity.
Putting the identifier in as a feature. sesion_id or sku as a column makes the model memorise. AutoML usually discards high-cardinality columns, but do not rely on it: exclude them yourself.
Accepting the 0.5 threshold without thinking. It is a default value, not a recommendation. The threshold is computed from the costs of each type of error.
Labelling with no written guide. Inconsistent labels put a ceiling on model quality that no training budget lifts.
Spending the big budget on the first attempt. One node-hour tells you whether there is signal. If there is not, twenty will not find it either.
Tip: keep the test predictions along with their real label. That is what lets you recompute thresholds, analyse by segment and compare against the next model without retraining.
Tip: put the result in a person's hands before putting it in a process's hands. The list of product pages with a suspicious image contributes value from day one and with no risk.
Exercises
Exercise 1
Lucía trains a model with AutoML to predict whether an order will be returned. She gets an accuracy of 94.2 % and wants to deploy it. You know that 6 % of AlpinaShop's orders are returned. What do you ask her before giving your approval, and what metrics do you ask for?
Exercise 2
Design the training set for a model that predicts, at the moment an order is placed, whether that order will arrive late (more than 5 days). List the features you would use, mark the ones that would be data leakage and explain how you would split train/validation/test.
Exercise 3
The image classification model gets 91 % global accuracy. Broken down by category, crampon has 34 % recall and backpack has 98 %. Diagnose the possible causes and propose a three-action plan.
Solutions
Solution 1
The key question, before any other: what accuracy would a model that always said "not returned" have? Answer: 94 %. Lucía's model contributes 0.2 points over doing nothing. In all likelihood it has learned to say "no" almost always.
Metrics to ask for:
- The complete confusion matrix over the test set. If the true positives are almost zero, the diagnosis is confirmed.
- Precision and recall of the "returned" class, which is the only one that matters. Global accuracy is irrelevant here.
- PR-AUC, not ROC-AUC, because of the 94/6 imbalance.
- Comparison against a baseline: what do you get with the rule "textile category + size at the end of the range"? Many returns of technical clothing are about sizing, and a simple rule can capture a good part of them.
Questions about the data:
- Which columns went in? If
motivo_devolucion,fecha_devolucionorestado_finalappear, there is leakage and the accuracy means nothing. - How was it split? If it was random and there is seasonality — and with returns there is, with the January peak — the metrics are inflated.
- How many returns are there in the test set? With 6 % of 2,000 rows that is 120 cases, a small number to estimate anything with precision.
And the business question that frames it all: what is going to be done with the prediction? If it is warning the customer to check the size before confirming, a model with moderate recall already contributes. If it is blocking orders, it is not deployed: a false positive means losing a legitimate sale and an angry customer.
Recommendation: do not deploy. Retrain optimising PR-AUC with auto_class_weights, audit the columns, split temporally and evaluate again against the baseline.
Solution 2
Objective: binary classification llego_tarde (delivery > 5 calendar days), decided at the instant the order is confirmed.
Valid features (all known at confirmation time):
| Group | Features |
|---|---|
| Order | Units, amount, number of lines, estimated total weight, volume |
| Product | Dominant category, whether any item is out of stock, whether it comes from an external supplier |
| Destination | Province, whether it is a rural area, whether it is the Balearics or the Canaries, grouped postcode |
| Shipping | Assigned carrier, service type, whether it is pick-up at a collection point |
| Temporal | Day of the week, hour, whether it is the eve of a public holiday, week of the year, whether it is a campaign |
| Historical | Carrier's average delay to that province over the last 30 days, external supplier's average delay |
Data leakage — features that CANNOT go in:
| Feature | Why it is leakage |
|---|---|
fecha_entrega |
It is the label in disguise |
dias_transito |
Same |
numero_incidencias |
Incidents happen during shipping |
estado_actual |
It only exists afterwards |
veces_reintentado_reparto |
Subsequent to shipping |
retraso_medio_transportista computed over the whole history |
It includes periods later than this order |
The last one deserves a pause: the feature is valid, it is the computation that leaks. It must be a rolling window over the 30 days before the order date, materialised in a daily historical table, exactly like the hist_cliente_diario in section 4.
CREATE OR REPLACE TABLE `alpinashop-datos.alpinashop_analitica.ml_entregas` AS
SELECT
p.pedido_id,
p.fecha_pedido,
p.unidades, p.n_lineas, ROUND(p.peso_kg,2) AS peso_kg,
p.provincia, p.transportista, p.tipo_servicio,
EXTRACT(DAYOFWEEK FROM p.fecha_pedido) AS dia_semana,
EXTRACT(WEEK FROM p.fecha_pedido) AS semana,
t.retraso_medio_30d, -- window BEFORE the order
IF(DATE_DIFF(p.fecha_entrega, p.fecha_pedido, DAY) > 5, 1, 0) AS llego_tarde,
CASE
WHEN p.fecha_pedido < DATE '2025-10-01' THEN 'TRAIN'
WHEN p.fecha_pedido < DATE '2025-12-15' THEN 'VALIDATE'
ELSE 'TEST'
END AS conjunto
FROM `alpinashop-datos.alpinashop_analitica.pedidos_entrega` p
LEFT JOIN `alpinashop-datos.alpinashop_analitica.hist_transportista_diario` t
ON t.transportista = p.transportista
AND t.provincia = p.provincia
AND t.fecha = DATE_SUB(p.fecha_pedido, INTERVAL 1 DAY)
WHERE p.fecha_entrega IS NOT NULL;Split: temporal, without exception. Logistics is seasonal in the worst sense: the Christmas campaign concentrates the delays. With a random split, the model would see December orders in training and would learn the pattern of that particular December.
An important caution about the test period: if TEST falls entirely within the Christmas campaign, the metrics will be pessimistic and not representative of the rest of the year. The correct approach is to evaluate over two periods — one campaign period and one normal one — and look at both. A single metric over an atypical period is a misleading metric.
Final note: fecha_entrega IS NOT NULL excludes orders still in transit. That is correct for training, but it introduces a bias if very delayed orders take so long that they fall outside the window. It is worth checking how many there are.
Solution 3
Diagnosis. The four causes, by probability:
1. Too few crampon examples (the most likely). It is a niche product: if there are 40 crampon photos against 600 backpack ones, the model has barely seen the class. Besides, the 91 % global accuracy is dominated by the large classes, so the failure hides.
2. Confusion with visually close classes. A crampon and an ice axe's metalwork share metal, points and straps. You have to look at the per-class confusion matrix to see where the errors go: if 50 % of crampons are classified as "ice axe", the problem is discrimination between similar classes, not lack of data.
3. Unrepresentative photos. If crampons are always photographed on a white background and from above, and in test they appear mounted on a boot, the model does not recognise the object in context.
4. Inconsistent labelling. Is a pack of crampon + case a "crampon"? And a photo of the boot with the crampon fitted? If those decisions were not taken in writing, there is noise in the label.
-- How many examples per class and where the errors go
SELECT etiqueta_real, COUNT(*) AS ejemplos
FROM `alpinashop-datos.alpinashop_analitica.imagenes_etiquetadas`
GROUP BY etiqueta_real ORDER BY ejemplos;
SELECT etiqueta_real, categoria_predicha, COUNT(*) AS casos
FROM `alpinashop-datos.alpinashop_analitica.eval_imagenes`
WHERE etiqueta_real = 'crampon'
GROUP BY 1,2 ORDER BY casos DESC;A three-action plan, in order:
Action 1 — Balance and expand crampon. Go from 40 to at least 200 images, and with deliberate variety: different models, backgrounds, angles, fitted and loose, with and without the case. If there are not 200 photos of your own, you can use every available variant from the catalogue (original, web, thumb count as different photos only if they genuinely differ; if they are the same image rescaled, they contribute nothing). It is the action with the highest return.
Action 2 — Clarify the taxonomy. If the confusion matrix shows that crampons and ice axes get mixed up, there are two ways out: merge them into a material_progresion_hielo label if the distinction adds no business value, or separate them better with an explicit guide and examples of every borderline case. Here, the right decision depends on what the classification is used for in the shop, not on the model.
Action 3 — Change the tracking metric and the use. Replace global accuracy with macro-average recall per class, which treats the rare category and the majority one alike, and always publish the per-class table. And in the meantime, use the model with a confidence threshold: automatically accept predictions above 0.90 and send the rest to human review. With that, the bad class stops being a production problem and becomes a short work queue, while more examples are collected — which, incidentally, come already labelled by the reviewer, feeding Action 1.
What you should NOT do: raise the training budget. The problem is the data, not the compute time, and spending more node-hours on 40 images does not create information that does not exist.
Conclusion
You have seen what AutoML does underneath — feature engineering, architecture search, hyperparameter tuning and ensembling — and, above all, what it does not do: it does not choose the problem, it does not obtain data, it does not detect that you have slipped in a column from the future and it does not know what the consequences of being wrong are. That is still human work, and it is where the value sits.
You know the supported problem types and the difference between the technical minimum and the reasonable minimum of data, with the rule that variety matters more than quantity.
You have built AlpinaShop's tabular case — predicting a cart's conversion — taking care of the only thing that really decides the result: that every feature was available at the instant of the decision, joining the customer's history with the previous day and not with today. And you have understood why a temporal split is compulsory when time is involved, and why the metrics drop when you apply it: because the previous ones were a lie.
You know how to read a confusion matrix and you no longer trust accuracy: with 12 % conversion, saying "nobody buys" gives 88 %. You tell ROC-AUC from PR-AUC and you know which one to look at when the classes are imbalanced. And you have turned the choice of threshold into what it really is: a table of expected profit with the costs of each type of error, up for discussion by management and not by the technical team. You know how to read feature importance as association and not as causation, and to use it as a sanity check.
You have applied AutoML Vision to the catalogue's 60 GB, understanding why the pre-trained API is not enough when the taxonomy is your own, why all the photos of one SKU must fall in the same set, and what labelling well really costs: not the hours, but the consistency. And the result you have put into production is not a model that decides on its own, but a prioritised list of suspicious product pages that a person reviews — more value and less risk.
You have the three traps identified — leakage, imbalance and overfitting — with their symptom and their remedy, and the table that decides between a pre-trained API, BigQuery ML, AutoML and custom training. And you are clear about the responsibility framework: measuring performance by segment to detect bias before the feedback loop entrenches it, treating pseudonymisation as what it is — personal data — and going through compliance before deploying any model that affects people, with the GDPR and the AI Act on the table.
But AutoML has a ceiling, and AlpinaShop is about to hit it. The recommender Marta wants — one that learns at the same time what customers are like and what products are like, and that knows how to place them in the same space to measure affinity — is not a tabular classification problem. There is no target column to predict: a representation has to be learned. That requires a specific architecture, a specific loss function and control over how the data is read.
In 05-03, TensorFlow on GCP, we go down to the code. You will see the minimum of TensorFlow and Keras needed to follow along without being an expert, you will build AlpinaShop's recommender with customer and product embeddings following the two-tower intuition, you will learn to read efficiently from Cloud Storage with tf.data and TFRecord, you will package the training for Vertex AI choosing between CPU, GPU and TPU with judgement, you will use checkpoints to survive Spot VMs, you will follow the training with managed TensorBoard and Experiments, and you will end up saving a SavedModel in the Model Registry ready to serve.
And you will still have to beat thirty lines of SQL.
Google Cloud Platform (GCP) Course
Module 1: Introduction to Google Cloud Platform
- What is Google Cloud Platform?
- Setting Up Your GCP Account
- A Tour of the GCP Console
- Projects, Resource Hierarchy and Billing
- Regions, Zones and the Shared Responsibility Model
- Cloud Shell and the gcloud CLI
Module 2: Core GCP Services
- Compute Engine: Virtual Machines on Google Cloud
- Cloud Storage: Object Storage
- Cloud SQL: Managed Relational Databases
- App Engine: Platform as a Service
- Google Kubernetes Engine (GKE)
- NoSQL Databases: Firestore, Bigtable and Spanner
- How to Choose the Right Compute Service
Module 3: Networking and Security
- VPC Networks
- Cloud Load Balancing
- Cloud CDN
- Identity and Access Management (IAM)
- Cloud Armor
- Secrets and Encryption: Secret Manager and Cloud KMS
- Cloud DNS, TLS Certificates and Publishing Services Securely
Module 4: Data and Analytics
- BigQuery: The Analytical Data Warehouse
- Cloud Dataflow: Batch and Streaming Data Processing
- Cloud Dataproc: Managed Spark and Hadoop
- Cloud Pub/Sub: Asynchronous Messaging
- Cloud Data Fusion: Code-Free Data Integration
- Orchestrating Pipelines with Cloud Composer and Workflows
- Data Governance and Dashboards with Dataplex and Looker Studio
Module 5: Machine Learning and AI
- Vertex AI: The Machine Learning Platform on GCP
- AutoML: Custom Models Without Writing Code
- TensorFlow on GCP: Training and Serving Models
- Natural Language API
- Vision API
- Generative AI on Vertex AI: Gemini Models and Embeddings
- MLOps: From Model to Product with Vertex AI Pipelines
Module 6: DevOps and Monitoring
- Cloud Build: Continuous Integration on GCP
- Cloud Source Repositories and Source Code Management
- Cloud Functions: Serverless Functions
- Cloud Monitoring (formerly Stackdriver): Metrics, Dashboards and Alerts
- Cloud Deployment Manager and Native Infrastructure as Code
- Cloud Logging and Cloud Trace: Logs, Traces and Diagnostics
- Terraform on GCP: Infrastructure as Code in Practice
Module 7: Advanced GCP Topics
- Hybrid and Multicloud with Anthos
- Serverless Computing with Cloud Run
- Advanced Networking: Shared VPC, Peering and Hybrid Connectivity
- Security Best Practices
- Cost Management and Optimization
- Reliability: SLOs, High Availability and Disaster Recovery
- Governance at Scale: Organization, Policies and Auditing
