The heuristic recommender from the previous lesson is already in production, and it works. But AlpinaShop has two questions that SQL cannot answer.

The first one comes from Marta, who runs the website: out of every hundred carts that get filled, which ones are going to end in a purchase? If you knew in advance, you could reserve the welcome discount for the person who is about to abandon instead of giving it away to someone who was going to buy anyway. That is not a pattern you read off a co-occurrence table: it depends on a dozen signals that interact with each other.

The second one comes from the catalogue: 60 GB of unlabelled product images. Nobody knows, looking at the alpinashop-catalogo bucket, which ones are backpacks, which ones ice axes and which ones jackets, beyond whatever the product page says — which does not always match the photo the supplier uploaded.

Both questions need a trained model. And none of the three people on the team is a machine learning engineer.

AutoML exists for exactly this: training reasonably good models on your own data, without writing modelling code, letting the platform take the technical decisions. This lesson is about what it does underneath, how far it goes, where it falls short, and — the part most often neglected — how to read the results without fooling yourself.

Contents

  1. What AutoML is and what it really does underneath
  2. Who it is for and what its limits are
  3. Problem types and minimum data requirements
  4. Case 1: predicting a cart's purchase
  5. The temporal split: why a random one would be cheating
  6. Training budget and cost
  7. Reading the results without fooling yourself
  8. The decision threshold is a business decision
  9. Feature importance and explainability
  10. Case 2: classifying the catalogue photos
  11. Labelling well: the cost nobody budgets for
  12. Deploying: endpoint or batch prediction
  13. The three traps: leakage, imbalance and overfitting
  14. When AutoML is not the answer
  15. Bias, automated decision-making and the legal framework

  1. What AutoML is and what it really does underneath

AutoML is not magic, nor is it a specific algorithm. It is the automation of the tasks an ML engineer would do by hand, executed systematically and with more patience than any person has.

Four things it does for you:

Feature engineering. It detects the type of each column, normalises the numeric ones, encodes the categorical ones, extracts components from dates (day of the week, month, whether it is a public holiday), imputes missing values and discards useless columns — the ones with a single value or the ones that are a unique identifier per row.

Architecture and algorithm search. It tries different model families — boosted trees, neural networks, linear models — and, within each one, different configurations. For image and text it also searches for the network architecture through neural architecture search.

Hyperparameter tuning. It optimises learning rates, depths, regularisation and the rest, with the same kind of Bayesian search you saw in 05-01, but without you having to declare it.

Ensembling. The final model is rarely a single one: it is usually a weighted combination of several, because combining diverse models almost always beats the best individual one.

flowchart TD
    A[Your labelled data] --> B[Automatic analysis and cleaning]
    B --> C[Feature engineering]
    C --> D[Architecture and algorithm search]
    D --> E[Hyperparameter tuning]
    E --> F[Evaluation on the test set]
    F -->|budget left| D
    F -->|budget exhausted| G[Ensembling of the best model]
    G --> H[Model in Model Registry]

What it does not do, and it is worth being clear about from the start:

  • It does not decide which problem to solve or what success means.
  • It does not obtain data, nor fix data that is wrong.
  • It does not detect that you have slipped in a column that leaks the future.
  • It does not know what is fair, nor what the consequences of being wrong are.

All of that is still yours. And it is, as it happens, where most of the value and most of the risk sit.

  1. Who it is for and what its limits are

AutoML has a very well-defined audience: whoever knows the domain and the data but not model engineering. Lucía is the exact case. She knows what a session is, what an abandoned cart is and what every column of visitas means. She does not know how to choose between XGBoost and a neural network, and she does not need to.

Advantages Real limits
A decent result in hours, not weeks Usually falls short of a good custom model
Requires no knowledge of modelling Little control over preprocessing
Evaluation and explainability included Higher training cost per result
Integrated with Model Registry and endpoints The model is a fairly closed box
A good baseline to compare against No custom architectures or loss functions

And one function that is almost never mentioned but is among the most valuable: AutoML is an excellent detector of badly framed problems. If AutoML, with all its effort, cannot get anything better than chance, it is very likely that the signal is not in the data, and no custom model is going to invent it. And if AutoML gets 99.8 % on the first attempt, you almost certainly have data leakage. In both cases it has saved you weeks.

  1. Problem types and minimum data requirements

Data type Supported problems Technical minimum Reasonable minimum in practice
Tabular Binary and multiclass classification, regression, forecasting 1,000 rows 10,000+ rows, and ≥100 examples of the minority class
Image Classification (single or multi-label), object detection 10 images per label 100–500 per label, with real variety
Text Classification, entity extraction, sentiment analysis 20 examples per label 100–1,000 per label
Video Classification, action recognition, tracking Dozens of clips Hundreds, and well trimmed

The difference between the last two columns matters a great deal. The technical minimum is what the platform accepts; the reasonable minimum is what produces a usable model. With 10 images per label AutoML trains and gives you a number, but that model is of no use whatsoever in production.

And a rule that holds for the whole module: variety matters more than quantity. Five hundred photos of backpacks all taken against the same white background with the same lighting teach the model to recognise that background, not the backpack. Two hundred varied photos — different backgrounds, angles, lighting — generalise far better.

A compulsory check before training anything:

SELECT
  COUNT(*)                                            AS filas,
  COUNTIF(compro = 1)                                 AS compraron,
  ROUND(100 * COUNTIF(compro = 1) / COUNT(*), 2)      AS pct_positivos,
  COUNT(DISTINCT sesion_id)                           AS sesiones_unicas,
  MIN(fecha)                                          AS desde,
  MAX(fecha)                                          AS hasta,
  COUNTIF(importe_carrito IS NULL)                    AS sin_importe
FROM `alpinashop-datos.alpinashop_analitica.ml_sesiones`;

If pct_positivos is 0.4 %, if sesiones_unicas is lower than filas (there are duplicates) or if sin_importe is half the set, you have work to do before training. Training first and discovering it afterwards costs machine hours.

  1. Case 1: predicting a cart's purchase

The goal, written as a business decision and not as a technical problem: estimate the probability that a session with a cart ends in a purchase, in order to decide who gets shown a closing incentive.

The starting table is ml_sesiones, the one you built in 05-01 by crossing visitas with v_pedidos_analitica. We are going to enrich it with customer and cart context, taking care that everything is information available at the instant of the decision:

CREATE OR REPLACE TABLE `alpinashop-datos.alpinashop_analitica.automl_carritos` AS
SELECT
  v.sesion_id,
  v.fecha,
  -- Session context
  v.dispositivo,
  v.canal,
  v.paginas_vistas,
  v.minutos_sesion,
  v.productos_vistos,
  -- Cart context
  v.unidades_carrito,
  ROUND(v.importe_carrito, 2)                                AS importe_carrito,
  ROUND(v.importe_carrito / NULLIF(v.unidades_carrito,0), 2) AS precio_medio_articulo,
  c.categoria_principal,
  -- Temporal context
  EXTRACT(DAYOFWEEK FROM v.fecha)                            AS dia_semana,
  EXTRACT(HOUR FROM v.hora_inicio)                           AS hora_dia,
  -- Customer context, computed BEFORE this session
  IFNULL(h.pedidos_previos, 0)                               AS pedidos_previos,
  IFNULL(ROUND(h.ticket_medio_previo, 2), 0)                 AS ticket_medio_previo,
  IFNULL(h.dias_desde_ultimo_pedido, 999)                    AS dias_desde_ultimo,
  -- Label
  v.compro
FROM `alpinashop-datos.alpinashop_analitica.ml_sesiones` v
LEFT JOIN `alpinashop-datos.alpinashop_analitica.carrito_categorias` c
  ON c.sesion_id = v.sesion_id
LEFT JOIN `alpinashop-datos.alpinashop_analitica.hist_cliente_diario` h
  ON h.cliente_hash = v.cliente_hash
 AND h.fecha        = DATE_SUB(v.fecha, INTERVAL 1 DAY)
WHERE v.unidades_carrito > 0;

The detail that decides whether this model is any use is in the last JOIN: h.fecha = DATE_SUB(v.fecha, INTERVAL 1 DAY). The customer's history is taken from the day before the session, not from today. If it were joined with the current history, every training row would carry information about orders placed after the session being predicted. The model would learn that "customers with many orders buy", which is true and also useless: in production, at the moment of deciding, that future order does not exist.

This kind of table — a snapshot of the state of each entity on each date — is called a historical feature table and it is what prevents almost every temporal leak. It is exactly what a Feature Store stores for you (05-01, section 7).

Creating the dataset and launching training:

gcloud ai datasets create \
  --project=alpinashop-datos --region=europe-west1 \
  --display-name=carritos-conversion \
  --metadata-schema-uri="gs://google-cloud-aiplatform/schema/dataset/metadata/tabular_1.0.0.yaml" \
  --metadata="{\"inputConfig\":{\"bigquerySource\":{\"uri\":\"bq://alpinashop-datos.alpinashop_analitica.automl_carritos\"}}}"

From the console, the flow asks for: target column (compro), objective type (binary classification), columns to exclude (sesion_id and fecha — the first is an identifier, the second defines the split and must not go in as a feature), split method, budget and optimisation metric.

In Python it comes out more explicit and, above all, reproducible:

from google.cloud import aiplatform

aiplatform.init(project="alpinashop-datos", location="europe-west1")

ds = aiplatform.TabularDataset("projects/.../datasets/1234567890")

job = aiplatform.AutoMLTabularTrainingJob(
    display_name="carritos-conversion-v1",
    optimization_prediction_type="classification",
    optimization_objective="maximize-au-prc",   # PR-AUC, not accuracy
)

model = job.run(
    dataset=ds,
    target_column="compro",
    predefined_split_column_name="conjunto",     # our own temporal split
    budget_milli_node_hours=2000,                # 2 node-hours
    model_display_name="automl-carritos-v1",
    disable_early_stopping=False,
)

Two choices that deserve justification:

  • maximize-au-prc instead of accuracy. With imbalanced classes, the area under the precision-recall curve reflects what you care about — finding the positives — far better than the percentage of correct answers. We come back to this in section 7.
  • disable_early_stopping=False: if the model stops improving, AutoML halts and does not charge you the remaining budget. Turning it off only makes sense in very specific cases.

  1. The temporal split: why a random one would be cheating

By default AutoML splits into 80 % training, 10 % validation and 10 % test, at random. For AlpinaShop's problem that is wrong, and it is worth understanding properly because it is the most repeated conceptual mistake.

Imagine that on 14 February AlpinaShop launches a campaign and conversions shoot up. With a random split, some 14 February sessions end up in training and others in test. The model, while training, sees how that particular day behaved and learns to recognise it. When evaluated on the other sessions of the same day, it gets them spectacularly right.

That result will never reproduce in production, because in production the model predicts on days it has never seen. A random split measures "can it interpolate within the known?" when the real question is "can it extrapolate to the unknown?".

The rule: if the problem has a temporal dimension, the split must be temporal. You train with the past, validate with the immediate past and evaluate with the future.

CREATE OR REPLACE VIEW `alpinashop-datos.alpinashop_analitica.automl_carritos_split` AS
SELECT
  *,
  CASE
    WHEN fecha <  DATE '2025-11-01' THEN 'TRAIN'
    WHEN fecha <  DATE '2026-01-01' THEN 'VALIDATE'
    ELSE                                 'TEST'
  END AS conjunto
FROM `alpinashop-datos.alpinashop_analitica.automl_carritos`
WHERE fecha BETWEEN DATE '2024-01-01' AND DATE '2026-03-31';

Vertex AI accepts this column with predefined_split_column_name="conjunto", and the values must be exactly TRAIN, VALIDATE and TEST.

There is a side effect worth anticipating: the metrics will drop. That is normal and it is good. If with a random split the AUC was 0.91 and with a temporal split it is 0.83, the 0.83 is the real number. The 0.91 was an illusion, and finding that out now is far cheaper than finding it out after deployment.

One further nuance: the test period must cover at least one complete business cycle. In a mountain gear shop, two winter months do not represent the summer. If you can, it is worth also evaluating over a seasonally different period.

  1. Training budget and cost

The budget is expressed in node-hours and limits how much AutoML explores. It is given in thousandths: budget_milli_node_hours=2000 is 2 node-hours.

Situation Indicative budget Comment
First trial, is there signal? 1 node-hour Cheap and enough to rule it out
Working tabular model 3–6 node-hours The usual point of diminishing returns
Large, complex dataset 10–20 node-hours Only if the short trial looked promising
Image Several hours Higher hourly cost than tabular

The price per node-hour varies by data type and by region and changes over time: always check the current official documentation. As an order of magnitude, a tabular training run of a few hours runs to tens of euros; an image one with many hours can reach several hundred.

Three pieces of advice that genuinely save money:

  1. Always start with 1 node-hour. If with that budget the AUC is 0.52 — that is, chance — spending twenty times more will not fix it: the problem is in the data.
  2. Leave early stopping enabled. You do not pay for what is not used.
  3. Before spending, compare against BigQuery ML. A logistic regression costs cents and on simple tabular problems comes surprisingly close. If AutoML improves on BigQuery ML by two AUC points, the legitimate question is whether those two points are worth the difference in cost and in opacity.

  1. Reading the results without fooling yourself

This is where the person who uses AutoML is separated from the person who understands it.

Suppose the cart model's result on the test set, with 10,000 sessions of which 1,200 ended in a purchase (12 %):

Predicted purchase Predicted no purchase
Actually bought 780 (TP) 420 (FN)
Did not buy 890 (FP) 7,910 (TN)

The metrics that come out of it:

Metric Formula Value What it means here
Accuracy (TP+TN)/total 86.9 % Misleading: saying "nobody buys" would give 88 %
Precision TP/(TP+FP) 46.7 % Of every 100 I give the incentive to, 47 were going to buy
Recall TP/(TP+FN) 65.0 % I detect 65 out of every 100 real buyers
F1 harmonic mean 54.5 % Balance between the previous two
ROC-AUC — 0.83 Probability of ranking a positive above a negative correctly
PR-AUC — 0.51 The honest metric with imbalanced classes

Accuracy is the metric that misleads most, and by a wide margin. In this case, a model that always answered "no purchase" would have 88 % accuracy — more than ours — and would be completely useless. Every time you see someone boasting about accuracy on an imbalanced problem, ask what the percentage of the majority class is.

ROC-AUC and PR-AUC are not interchangeable. ROC-AUC incorporates the true negatives, which here are 7,910 and extremely abundant, and that inflates it: 0.83 sounds good. PR-AUC only looks at what happens with the positives, and its 0.51 is a far more faithful description of the real difficulty. With imbalanced classes, always look at PR-AUC.

And the reference you must never forget: how much would a silly rule give? If ranking sessions by cart amount already identified 55 % of the buyers, the model contributes ten points, not sixty-five.

  1. The decision threshold is a business decision

A classifier does not return "yes" or "no": it returns a probability between 0 and 1. Turning it into a decision requires a threshold, and that threshold is not chosen by the model. You choose it, and it is an economic decision.

For AlpinaShop, the incentive is a 10 % discount. The numbers per session:

  • True positive: I give the discount to someone who was going to buy → I lose 10 % of the margin unnecessarily. On an average cart of €95, about –€4.
  • False positive: I give the discount to someone who was not going to buy → if the discount convinces a fraction of them, I gain; if not, it costs nothing because there is no sale. Expected value slightly positive.
  • False negative: I do not give the discount to someone who was going to abandon → I lose a sale I could have recovered, about –€20 of expected margin.
  • True negative: I do not give a discount to someone who was not going to buy → €0.

With those numbers, the expensive error is the false negative, and that pushes the threshold downwards: it pays to be generous handing out incentives. The trade-off table:

Threshold Sessions with incentive Precision Recall Business reading
0.20 4,100 26 % 89 % Almost everyone gets a discount; margin is given away
0.35 2,400 38 % 76 % A reasonable compromise
0.50 1,670 47 % 65 % The default, not necessarily the best
0.70 620 68 % 35 % Very selective; two thirds are lost

The right way to choose is to compute the expected profit of each threshold with the business values, not to look at which one gives the best F1:

WITH escenarios AS (
  SELECT
    umbral,
    COUNTIF(prob >= umbral AND compro = 1)      AS vp,
    COUNTIF(prob >= umbral AND compro = 0)      AS fp,
    COUNTIF(prob <  umbral AND compro = 1)      AS fn
  FROM `alpinashop-datos.alpinashop_analitica.predicciones_test`,
       UNNEST([0.2, 0.3, 0.35, 0.4, 0.5, 0.6, 0.7]) AS umbral
  GROUP BY umbral
)
SELECT
  umbral, vp, fp, fn,
  ROUND(vp * -4.0 + fp * 0.5 + fn * -20.0, 0) AS beneficio_estimado_eur
FROM escenarios
ORDER BY beneficio_estimado_eur DESC;

This query turns a technical discussion into a table that management understands. And it makes explicit something that usually stays implicit: the coefficients -4, 0.5 and -20 are business hypotheses, not truths. Writing them down forces you to discuss them, which is exactly what has to happen.

One final warning about the threshold: it can be different per segment. For new customers, where the acquisition value is higher, a lower threshold may be worth it. That is no longer a model decision, it is product design.

  1. Feature importance and explainability

AutoML returns two levels of explanation, and they serve different purposes.

Global importance: which features weigh most in the model as a whole. For the cart model, a plausible result:

Feature Importance Interpretation
pedidos_previos 0.24 Whoever has bought before buys again
minutos_sesion 0.19 Long sessions indicate intent
importe_carrito 0.15 Expensive carts are abandoned more
dias_desde_ultimo 0.12 Recency as a signal of engagement
canal 0.10 Brand search converts more than display
dispositivo 0.07 Mobile converts less than desktop
hora_dia 0.05 Marginal
dia_semana 0.03 Almost irrelevant

How to read this table, and how not to. Importance is association, not causation. The fact that minutos_sesion weighs a lot does not mean that artificially lengthening the session increases purchases: it means that whoever is going to buy tends to spend more time. Confusing the two produces absurd product decisions.

What you should do with this table is a sanity check. If a feature turned up with an importance of 0.80 and all the others at zero, that is an alarm signal: that column probably leaks the outcome. And if all the importances were similar and low, that is a sign that no feature carries real signal.

Local attributions: why the model predicted what it predicted for one specific row. They are requested in the prediction call with explain instead of predict, and they return each feature's contribution to that individual prediction. They are indispensable when a decision has to be justified to a person, and they will return in 05-07 as part of the responsible ML checklist.

  1. Case 2: classifying the catalogue photos

The second problem is different in nature: unstructured input, 60 GB of images in alpinashop-catalogo under productos/<sku>/original|web|thumb/.

The goal: classify each photo into AlpinaShop's own taxonomy — backpack, ice axe, helmet, jacket, boot, crampon, rope, sleeping bag — in order to detect product pages with wrongly assigned images and to enrich the catalogue metadata.

Why the Cloud Vision API is not enough. The pre-trained API (lesson 05-05) recognises "backpack" perfectly, but it does not tell a 60-litre trekking backpack from a 25-litre summit pack, and it will call a technical ice axe a "hatchet". AutoML Vision exists precisely for when the taxonomy is your own and the general vocabulary does not cover it.

The input data is declared in an index CSV in Cloud Storage:

TRAIN,gs://alpinashop-catalogo/productos/MOC-4471/web/01.jpg,backpack
TRAIN,gs://alpinashop-catalogo/productos/PIO-1120/web/01.jpg,ice-axe
VALIDATE,gs://alpinashop-catalogo/productos/CAS-8802/web/02.jpg,helmet
TEST,gs://alpinashop-catalogo/productos/BOT-3391/web/01.jpg,boot

Three columns: set, image URI and label. If the first column is omitted, Vertex AI splits for you.

Here the split can be random — there is no temporal dimension — but with one critical condition: all the photos of the same SKU must fall in the same set. A product usually has five or six almost identical photos; if some go to training and others to test, the model memorises that particular product and the evaluation lies. Grouping by entity is to the image problem what the temporal split is to the tabular one.

job = aiplatform.AutoMLImageTrainingJob(
    display_name="catalogo-categorias-v1",
    prediction_type="classification",
    multi_label=False,
    model_type="CLOUD",          # served on an endpoint; EDGE to export
    base_model=None,
)

image_model = job.run(
    dataset=image_ds,
    budget_milli_node_hours=8000,   # 8 node-hours
    training_filter_split=None,
    model_display_name="automl-catalogo-v1",
    disable_early_stopping=False,
)

model_type="CLOUD" produces a model optimised for serving on Vertex AI. "EDGE" generates an exportable variant (TensorFlow Lite, container) to run outside, smaller and somewhat less accurate. AlpinaShop does not need edge: its images are already in the cloud.

  1. Labelling well: the cost nobody budgets for

This section is the most boring in the lesson and the one that sinks the most projects.

To train AutoML Vision you need labelled images, and AlpinaShop's 60 GB are not. The real numbers of labelling:

  • 8 categories × 300 images = 2,400 images to label.
  • At around 6 seconds per image with a good interface, that is about 4 hours of human work.
  • Plus the time to agree the taxonomy, which always takes longer than expected.

And the problem is not the time, it is the consistency. The questions that come up half an hour in:

  • A photo showing a backpack with an ice axe strapped to it: what is it?
  • A climbing helmet and a ski helmet: are they the same label?
  • A close-up of the stitching on a jacket: is that "jacket" or is it discarded?
  • A pack of rope + carabiners: which label does it get?

If those decisions are not written down beforehand and are not consistent, the model learns the inconsistency. And a model trained with contradictory labels cannot exceed the quality of its labels: it is a hard ceiling.

The minimum protocol, which costs an afternoon and saves the project:

  1. Write the labelling guide with the definition of each category and at least one borderline case resolved per category.
  2. An otros label for whatever does not fit. Without it, people force labels and contaminate the set.
  3. Double-labelling of a sample: two people label the same 200 images. If they agree on less than 90 %, the taxonomy is ambiguous and has to be fixed before going on.
  4. Start with the web/ photos, which are normalised, before the original/ ones.

Vertex AI offers labelling workflows in the console and managed labelling services with human labellers. Be careful with the latter if the images contain anything sensitive: sending customer photos to an external labelling service has privacy implications that must be assessed beforehand. AlpinaShop's catalogue photos do not have them; the ones customers upload with their reviews do.

  1. Deploying: endpoint or batch prediction

The same dilemma as in 05-01, and the answer is once again the same for the same reason.

For the image model, the task is to classify 60 GB of photos once, and after that only the new ones that come in. That is a textbook case of batch prediction:

gcloud ai batch-prediction-jobs create \
  --project=alpinashop-datos --region=europe-west1 \
  --display-name=clasifica-catalogo-inicial \
  --model=MODEL_ID \
  --input-paths-uri=gs://alpinashop-datalake/entradas/catalogo_imagenes.jsonl \
  --input-format=jsonl \
  --output-uri-prefix=gs://alpinashop-datalake/salidas/catalogo-clases/ \
  --output-format=jsonl

And the results are loaded into the governed dataset to cross them with the catalogue:

CREATE OR REPLACE TABLE `alpinashop-datos.alpinashop_analitica.imagenes_clasificadas` AS
SELECT
  REGEXP_EXTRACT(instance.content, r'productos/([^/]+)/') AS sku,
  instance.content                                        AS uri_imagen,
  prediction.displayNames[OFFSET(0)]                      AS categoria_predicha,
  ROUND(prediction.confidences[OFFSET(0)], 3)             AS confianza
FROM `alpinashop-datos.alpinashop_analitica.raw_predicciones_imagen`;
-- Product pages with a possibly wrong image
SELECT i.sku, p.categoria AS categoria_ficha,
       i.categoria_predicha, i.confianza, i.uri_imagen
FROM `alpinashop-datos.alpinashop_analitica.imagenes_clasificadas` i
JOIN `alpinashop-datos.alpinashop_analitica.productos` p USING (sku)
WHERE i.categoria_predicha != p.categoria
  AND i.confianza > 0.85
ORDER BY i.confianza DESC;

This query is the real product of all the work. It is not a deployed model: it is a list of product pages a person can review, ordered by confidence. A model that produces a prioritised work list contributes more value with less risk than one that acts on its own.

For new photos, classification is triggered from the imagenes-subidas Pub/Sub topic, with the implementation that arrives in 06-03 with Cloud Functions.

  1. The three traps: leakage, imbalance and overfitting

Data leakage. A feature contains information that in production will not be available at prediction time. Unmistakable symptom: metrics that are too good. At AlpinaShop, the candidates are metodo_pago (it only exists if there was a purchase), direccion_envio_confirmada, codigo_descuento_aplicado and any customer history computed as of today instead of as of the session date. The smell test is always the same question: did this piece of data exist and was it known at the exact instant of the decision? If the answer is "more or less", then it did not.

Class imbalance. When a class is rare, the model learns to ignore it. With 12 % conversion the problem is moderate; with 0.3 % fraud it is serious. Remedies: choose PR-AUC as the objective, weight the classes, and in extreme cases undersample the majority class — never the test set, which must keep the real proportion so the metrics mean something. And sometimes the right remedy is to reframe: instead of classifying, rank by risk and review the top N.

Overfitting. The model memorises instead of generalising. Symptom: excellent on training, mediocre on test. AutoML controls it fairly well with regularisation and early stopping, but it cannot protect you from two causes that depend on you: too little data for the complexity of the problem, and leakage between sets — the same SKU or the same customer in training and in test.

Trap Symptom Usual cause What to do
Leakage Suspicious AUC > 0.97 A column from the future Audit every column with the instant question
Imbalance Tiny recall Rare class PR-AUC, weights, reframe as ranking
Overfitting Test much worse than train Little data or leakage Group by entity, more data, less complexity

  1. When AutoML is not the answer

Criterion Pre-trained API BigQuery ML AutoML Custom training
Own data needed None Yes, in BigQuery Yes, labelled Yes, labelled
Time to result Minutes Hours Hours or days Weeks
Knowledge required Calling an API SQL Interface + judgement ML and engineering
Control over the model None Low Low Total
Training cost 0 Very low Medium-high High
Cost per prediction Per unit Very low Medium Variable
Quality ceiling The provider's Medium High The highest
When to choose it The problem is generic Tabular and already in BigQuery Own taxonomy, no ML team Requirements nothing else covers

AutoML is not the answer when:

  • The problem is generic. "Does this image contain a person?" or "is this text positive?" are solved by the pre-trained APIs (05-04, 05-05) with no data, no training and for cents.
  • You have no labelled data and you are not going to invest in labelling it. Without labels there is no AutoML.
  • The data is tabular and already lives in BigQuery. Try BigQuery ML first: it is hours against days and cents against tens of euros.
  • You need your own architecture or loss function. That is the case of the two-tower recommender in 05-03.
  • A heuristic already solves 80 %. The DA-003 lesson from 05-01 still stands.
  • You need to run the model outside the cloud with strict requirements. EDGE mode helps, but your own model gives you more control over size and latency.

  1. Bias, automated decision-making and the legal framework

The two models in this lesson have very different risk profiles, and seeing them side by side is the best way to understand the framework.

The image model is low risk. It classifies inanimate objects. The worst possible error is labelling an ice axe as a hatchet, and the remedy is for a person to correct it in a review list.

The cart model does process personal data and does affect people. And there you have to think before deploying:

Bias. The model learns from what happened. If historically customers coming in from mobile convert less — because the mobile site is worse — the model will learn not to offer them incentives, so they will convert even less, so the model will confirm itself. It is a feedback loop: the model does not describe reality, it manufactures it. The way to detect it is to measure performance by segment — device, channel, country, new or returning customer — and not only in aggregate. A global AUC of 0.83 can hide a 0.86 on desktop and a 0.61 on mobile.

Personal data and GDPR. Even if it is trained over v_pedidos_analitica with the customer pseudonymised, pseudonymisation is not anonymisation: it is still personal data for GDPR purposes and the legal basis, minimisation, purpose limitation and retention periods still apply. The rights of access, objection and erasure apply too, which means being able to withdraw a customer's data from the next cycle's training set.

Automated decision-making. Showing or hiding a commercial discount is not the same as denying credit. But the boundary of Article 22 of the GDPR — decisions based solely on automated processing with significant effects — is not always obvious, and systematic price differentiation between customers can come closer to it than it seems.

The EU AI Act. The AI Regulation classifies systems by risk level, and from that classification follow obligations on risk management, data quality, technical documentation, record-keeping, transparency and human oversight. A commercial recommendation system usually sits at a low level, but the classification depends on the specific use, and the use can drift over time without anyone reviewing it.

Express recommendation. Before deploying the cart model, a compliance professional or the DPO must determine the legal basis for the processing, whether the intended use constitutes automated decision-making under Article 22, what information has to be given to customers, the retention period for the training set, and the system's classification under the AI Act with the associated obligations. This lesson describes technical controls and does not constitute legal advice. All data is fictitious.

And one technical measure that is worth more than many declarations: evaluate by segment and publish that table alongside the global metric. If nobody looks at performance by group, the bias is not detected.

Common Mistakes and Tips

Trusting accuracy. With imbalanced classes it is the metric that misleads most. Look at PR-AUC, precision and recall, and always compare them against the percentage of the majority class.

Leaving the split random on a temporal problem. It produces inflated metrics that do not reproduce in production.

Letting photos of the same product fall into different sets. Pure leakage: the model memorises the product and the evaluation lies. Group by entity.

Putting the identifier in as a feature. sesion_id or sku as a column makes the model memorise. AutoML usually discards high-cardinality columns, but do not rely on it: exclude them yourself.

Accepting the 0.5 threshold without thinking. It is a default value, not a recommendation. The threshold is computed from the costs of each type of error.

Labelling with no written guide. Inconsistent labels put a ceiling on model quality that no training budget lifts.

Spending the big budget on the first attempt. One node-hour tells you whether there is signal. If there is not, twenty will not find it either.

Tip: keep the test predictions along with their real label. That is what lets you recompute thresholds, analyse by segment and compare against the next model without retraining.

Tip: put the result in a person's hands before putting it in a process's hands. The list of product pages with a suspicious image contributes value from day one and with no risk.

Exercises

Exercise 1

Lucía trains a model with AutoML to predict whether an order will be returned. She gets an accuracy of 94.2 % and wants to deploy it. You know that 6 % of AlpinaShop's orders are returned. What do you ask her before giving your approval, and what metrics do you ask for?

Exercise 2

Design the training set for a model that predicts, at the moment an order is placed, whether that order will arrive late (more than 5 days). List the features you would use, mark the ones that would be data leakage and explain how you would split train/validation/test.

Exercise 3

The image classification model gets 91 % global accuracy. Broken down by category, crampon has 34 % recall and backpack has 98 %. Diagnose the possible causes and propose a three-action plan.

Solutions

Solution 1

The key question, before any other: what accuracy would a model that always said "not returned" have? Answer: 94 %. Lucía's model contributes 0.2 points over doing nothing. In all likelihood it has learned to say "no" almost always.

Metrics to ask for:

  1. The complete confusion matrix over the test set. If the true positives are almost zero, the diagnosis is confirmed.
  2. Precision and recall of the "returned" class, which is the only one that matters. Global accuracy is irrelevant here.
  3. PR-AUC, not ROC-AUC, because of the 94/6 imbalance.
  4. Comparison against a baseline: what do you get with the rule "textile category + size at the end of the range"? Many returns of technical clothing are about sizing, and a simple rule can capture a good part of them.

Questions about the data:

  • Which columns went in? If motivo_devolucion, fecha_devolucion or estado_final appear, there is leakage and the accuracy means nothing.
  • How was it split? If it was random and there is seasonality — and with returns there is, with the January peak — the metrics are inflated.
  • How many returns are there in the test set? With 6 % of 2,000 rows that is 120 cases, a small number to estimate anything with precision.

And the business question that frames it all: what is going to be done with the prediction? If it is warning the customer to check the size before confirming, a model with moderate recall already contributes. If it is blocking orders, it is not deployed: a false positive means losing a legitimate sale and an angry customer.

Recommendation: do not deploy. Retrain optimising PR-AUC with auto_class_weights, audit the columns, split temporally and evaluate again against the baseline.

Solution 2

Objective: binary classification llego_tarde (delivery > 5 calendar days), decided at the instant the order is confirmed.

Valid features (all known at confirmation time):

Group Features
Order Units, amount, number of lines, estimated total weight, volume
Product Dominant category, whether any item is out of stock, whether it comes from an external supplier
Destination Province, whether it is a rural area, whether it is the Balearics or the Canaries, grouped postcode
Shipping Assigned carrier, service type, whether it is pick-up at a collection point
Temporal Day of the week, hour, whether it is the eve of a public holiday, week of the year, whether it is a campaign
Historical Carrier's average delay to that province over the last 30 days, external supplier's average delay

Data leakage — features that CANNOT go in:

Feature Why it is leakage
fecha_entrega It is the label in disguise
dias_transito Same
numero_incidencias Incidents happen during shipping
estado_actual It only exists afterwards
veces_reintentado_reparto Subsequent to shipping
retraso_medio_transportista computed over the whole history It includes periods later than this order

The last one deserves a pause: the feature is valid, it is the computation that leaks. It must be a rolling window over the 30 days before the order date, materialised in a daily historical table, exactly like the hist_cliente_diario in section 4.

CREATE OR REPLACE TABLE `alpinashop-datos.alpinashop_analitica.ml_entregas` AS
SELECT
  p.pedido_id,
  p.fecha_pedido,
  p.unidades, p.n_lineas, ROUND(p.peso_kg,2) AS peso_kg,
  p.provincia, p.transportista, p.tipo_servicio,
  EXTRACT(DAYOFWEEK FROM p.fecha_pedido) AS dia_semana,
  EXTRACT(WEEK      FROM p.fecha_pedido) AS semana,
  t.retraso_medio_30d,                      -- window BEFORE the order
  IF(DATE_DIFF(p.fecha_entrega, p.fecha_pedido, DAY) > 5, 1, 0) AS llego_tarde,
  CASE
    WHEN p.fecha_pedido < DATE '2025-10-01' THEN 'TRAIN'
    WHEN p.fecha_pedido < DATE '2025-12-15' THEN 'VALIDATE'
    ELSE                                         'TEST'
  END AS conjunto
FROM `alpinashop-datos.alpinashop_analitica.pedidos_entrega` p
LEFT JOIN `alpinashop-datos.alpinashop_analitica.hist_transportista_diario` t
  ON t.transportista = p.transportista
 AND t.provincia     = p.provincia
 AND t.fecha         = DATE_SUB(p.fecha_pedido, INTERVAL 1 DAY)
WHERE p.fecha_entrega IS NOT NULL;

Split: temporal, without exception. Logistics is seasonal in the worst sense: the Christmas campaign concentrates the delays. With a random split, the model would see December orders in training and would learn the pattern of that particular December.

An important caution about the test period: if TEST falls entirely within the Christmas campaign, the metrics will be pessimistic and not representative of the rest of the year. The correct approach is to evaluate over two periods — one campaign period and one normal one — and look at both. A single metric over an atypical period is a misleading metric.

Final note: fecha_entrega IS NOT NULL excludes orders still in transit. That is correct for training, but it introduces a bias if very delayed orders take so long that they fall outside the window. It is worth checking how many there are.

Solution 3

Diagnosis. The four causes, by probability:

1. Too few crampon examples (the most likely). It is a niche product: if there are 40 crampon photos against 600 backpack ones, the model has barely seen the class. Besides, the 91 % global accuracy is dominated by the large classes, so the failure hides.

2. Confusion with visually close classes. A crampon and an ice axe's metalwork share metal, points and straps. You have to look at the per-class confusion matrix to see where the errors go: if 50 % of crampons are classified as "ice axe", the problem is discrimination between similar classes, not lack of data.

3. Unrepresentative photos. If crampons are always photographed on a white background and from above, and in test they appear mounted on a boot, the model does not recognise the object in context.

4. Inconsistent labelling. Is a pack of crampon + case a "crampon"? And a photo of the boot with the crampon fitted? If those decisions were not taken in writing, there is noise in the label.

-- How many examples per class and where the errors go
SELECT etiqueta_real, COUNT(*) AS ejemplos
FROM `alpinashop-datos.alpinashop_analitica.imagenes_etiquetadas`
GROUP BY etiqueta_real ORDER BY ejemplos;

SELECT etiqueta_real, categoria_predicha, COUNT(*) AS casos
FROM `alpinashop-datos.alpinashop_analitica.eval_imagenes`
WHERE etiqueta_real = 'crampon'
GROUP BY 1,2 ORDER BY casos DESC;

A three-action plan, in order:

Action 1 — Balance and expand crampon. Go from 40 to at least 200 images, and with deliberate variety: different models, backgrounds, angles, fitted and loose, with and without the case. If there are not 200 photos of your own, you can use every available variant from the catalogue (original, web, thumb count as different photos only if they genuinely differ; if they are the same image rescaled, they contribute nothing). It is the action with the highest return.

Action 2 — Clarify the taxonomy. If the confusion matrix shows that crampons and ice axes get mixed up, there are two ways out: merge them into a material_progresion_hielo label if the distinction adds no business value, or separate them better with an explicit guide and examples of every borderline case. Here, the right decision depends on what the classification is used for in the shop, not on the model.

Action 3 — Change the tracking metric and the use. Replace global accuracy with macro-average recall per class, which treats the rare category and the majority one alike, and always publish the per-class table. And in the meantime, use the model with a confidence threshold: automatically accept predictions above 0.90 and send the rest to human review. With that, the bad class stops being a production problem and becomes a short work queue, while more examples are collected — which, incidentally, come already labelled by the reviewer, feeding Action 1.

What you should NOT do: raise the training budget. The problem is the data, not the compute time, and spending more node-hours on 40 images does not create information that does not exist.

Conclusion

You have seen what AutoML does underneath — feature engineering, architecture search, hyperparameter tuning and ensembling — and, above all, what it does not do: it does not choose the problem, it does not obtain data, it does not detect that you have slipped in a column from the future and it does not know what the consequences of being wrong are. That is still human work, and it is where the value sits.

You know the supported problem types and the difference between the technical minimum and the reasonable minimum of data, with the rule that variety matters more than quantity.

You have built AlpinaShop's tabular case — predicting a cart's conversion — taking care of the only thing that really decides the result: that every feature was available at the instant of the decision, joining the customer's history with the previous day and not with today. And you have understood why a temporal split is compulsory when time is involved, and why the metrics drop when you apply it: because the previous ones were a lie.

You know how to read a confusion matrix and you no longer trust accuracy: with 12 % conversion, saying "nobody buys" gives 88 %. You tell ROC-AUC from PR-AUC and you know which one to look at when the classes are imbalanced. And you have turned the choice of threshold into what it really is: a table of expected profit with the costs of each type of error, up for discussion by management and not by the technical team. You know how to read feature importance as association and not as causation, and to use it as a sanity check.

You have applied AutoML Vision to the catalogue's 60 GB, understanding why the pre-trained API is not enough when the taxonomy is your own, why all the photos of one SKU must fall in the same set, and what labelling well really costs: not the hours, but the consistency. And the result you have put into production is not a model that decides on its own, but a prioritised list of suspicious product pages that a person reviews — more value and less risk.

You have the three traps identified — leakage, imbalance and overfitting — with their symptom and their remedy, and the table that decides between a pre-trained API, BigQuery ML, AutoML and custom training. And you are clear about the responsibility framework: measuring performance by segment to detect bias before the feedback loop entrenches it, treating pseudonymisation as what it is — personal data — and going through compliance before deploying any model that affects people, with the GDPR and the AI Act on the table.

But AutoML has a ceiling, and AlpinaShop is about to hit it. The recommender Marta wants — one that learns at the same time what customers are like and what products are like, and that knows how to place them in the same space to measure affinity — is not a tabular classification problem. There is no target column to predict: a representation has to be learned. That requires a specific architecture, a specific loss function and control over how the data is read.

In 05-03, TensorFlow on GCP, we go down to the code. You will see the minimum of TensorFlow and Keras needed to follow along without being an expert, you will build AlpinaShop's recommender with customer and product embeddings following the two-tower intuition, you will learn to read efficiently from Cloud Storage with tf.data and TFRecord, you will package the training for Vertex AI choosing between CPU, GPU and TPU with judgement, you will use checkpoints to survive Spot VMs, you will follow the training with managed TensorBoard and Experiments, and you will end up saving a SavedModel in the Model Registry ready to serve.

And you will still have to beat thirty lines of SQL.

Google Cloud Platform (GCP) Course

Module 1: Introduction to Google Cloud Platform

Module 2: Core GCP Services

Module 3: Networking and Security

Module 4: Data and Analytics

Module 5: Machine Learning and AI

Module 6: DevOps and Monitoring

Module 7: Advanced GCP Topics

Module 8: Final Project

© Copyright 2026. All rights reserved