In AlpinaShop's data bucket there are 8,000 customer reviews written in free text, already de-identified with Sensitive Data Protection in 04-07. Five years of people saying what they thought of a backpack, whether the jacket is warm, whether the sizing came up small, whether the delivery arrived broken.

Nobody has read them. Not one.

Not out of neglect: reading 8,000 reviews at thirty seconds each is almost seventy hours of work, and at the end you would have a general impression you could not cross-reference with anything. The reviews exist, they are shown on the product page, they generate the average star rating, and there their useful life ends. All the information about why a product is liked or disliked is locked inside a text that no process touches.

The previous lesson ended by building a model from scratch, with containers, a GPU and three weeks of A/B testing. This one starts with the opposite principle, which is the one that saves most money in practice: do not train what is already trained.

Contents

  1. Pre-trained APIs versus your own models
  2. What the Natural Language API offers exactly
  3. score and magnitude: where everybody goes wrong
  4. Entity analysis and entity sentiment
  5. Content classification and syntactic analysis
  6. A first call from Python
  7. Processing the 8,000 reviews: batches, quotas and errors
  8. Real cost and how to estimate it before starting
  9. From BigQuery to opiniones_sentimiento and back
  10. The business question: does bad sentiment predict returns?
  11. Languages: Spanish, Catalan and what you have to check
  12. The real limits: irony, negation and mountain jargon
  13. Validating with a hand-labelled sample
  14. An honest comparison: API, Gemini or your own classifier
  15. The full family: Translation, Speech-to-Text, Document AI
  16. Privacy: what gets sent, what gets stored and what the GDPR says

  1. Pre-trained APIs versus your own models

A pre-trained API is a model somebody has already trained, with far more data and resources than you are ever going to have, exposed as a service. You send text and you receive an answer. There is no training data, no labelling, no training, no deployment, no drift monitoring.

Criterion Pre-trained API Your own model (AutoML or custom)
Data needed None Hundreds or thousands of labelled examples
Time to production An afternoon Weeks
Up-front cost 0 Labelling + training
Cost per use Per unit processed Endpoint or batch
Quality on generic tasks High Similar, with far more effort
Quality on your own jargon Medium or low High if there is data
Maintenance None Retraining and watching
Control None Total

The practical rule can be phrased as one question: is my problem essentially different from any other company's? Detecting whether a text in Spanish is positive or negative is not: it is exactly the same problem for AlpinaShop, for a pizzeria and for a bank. Knowing whether a review is about the harness fit or the stiffness of the sole is specific.

And the resulting order of work is clear: always start with the pre-trained API. If it solves the problem, you have finished in an afternoon. If it does not, you now know exactly where it fails, and that knowledge is precisely what you need to decide whether training something of your own is worth it.

  1. What the Natural Language API offers exactly

The Cloud Natural Language API exposes five functions. It is worth knowing all of them because three are systematically ignored and are very useful.

Function What it takes What it returns Use at AlpinaShop
Sentiment analysis A text score and magnitude for the text and for each sentence How positive each review is
Entity analysis A text People, places, products, quantities, with salience What is mentioned in the reviews
Entity sentiment A text Each entity with its own sentiment "The backpack fine, the delivery terrible"
Content classification A text (at least around 20 words) Categories from a general taxonomy Classifying long texts by topic
Syntactic analysis A text Tokens, lemmas, parts of speech, dependencies Normalising and extracting patterns

The third one — entity sentiment — is the hidden gem and the one with most value for a shop. A real review is almost never plainly "good" or "bad": it is "the jacket is fantastic but it took three weeks to arrive". The overall sentiment of that text comes out close to zero, which is both true and useless. Per-entity sentiment separates the fact that the product is liked from the fact that the logistics are not.

  1. score and magnitude: where everybody goes wrong

This is the most important section in the lesson, and the one that generates most misunderstandings in real projects.

Sentiment analysis returns two numbers, not one:

  • score: between −1.0 and +1.0. It indicates the overall emotional orientation: negative, neutral or positive.
  • magnitude: from 0.0 to infinity. It indicates the total amount of emotional charge in the text, and it is not normalised by length: a long text accumulates more magnitude.

The universal mistake is looking only at the score. And the problem is that a score close to 0 has two completely opposite meanings, which can only be told apart by looking at the magnitude.

score magnitude Interpretation Example
+0.8 0.9 Clearly positive "Perfect backpack, very comfortable."
−0.7 0.8 Clearly negative "The stitching came apart in a week. Awful."
0.0 0.1 Genuinely neutral "40-litre backpack, blue."
0.0 3.4 Mixed, highly emotional "The product is superb, but the customer service has been appalling and the delivery took three weeks."

The last two rows are the key. Both have a score of 0. The first is a description with no emotion. The second is an intense review with strong emotions in both directions that cancel out when averaged.

And they are cases that demand opposite actions. The neutral review requires nothing. The mixed review contains a serious customer service and logistics problem that has to be dealt with today. An analysis that only looks at the score will classify both as "neutral", will ignore the second, and the team will conclude that "most reviews are neutral, there is nothing here".

The right way to classify uses both dimensions:

CREATE OR REPLACE VIEW `alpinashop-datos.alpinashop_analitica.v_opiniones_clasificadas` AS
SELECT
  opinion_id, sku, fecha, puntuacion, score, magnitude,
  CASE
    WHEN score >=  0.25                     THEN 'positiva'
    WHEN score <= -0.25                     THEN 'negativa'
    WHEN magnitude >= 1.5                   THEN 'mixta'      -- high emotion, low score
    ELSE                                         'neutra'
  END AS clasificacion
FROM `alpinashop-datos.alpinashop_analitica.opiniones_sentimiento`;

The 0.25 and 1.5 thresholds are not sacred. They are a reasonable starting point that has to be calibrated against a hand-labelled sample — section 13 — exactly like the decision threshold in 05-02: it is a business adjustment, not a universal constant.

magnitude is not comparable across texts of very different lengths. A 300-word review will have more magnitude than a 20-word one even if it is less intense. If you need to compare, it is worth normalising by dividing by the number of sentences, or at least segmenting the analysis by length bands.

And one clarification that avoids absurd conclusions: sentiment does not measure truthfulness or severity. A very polite customer who writes "sadly the carabiner gave way on a route" may give a moderate score, and that is a safety incident. Sentiment ranks and prioritises; it does not replace reading what matters.

  1. Entity analysis and entity sentiment

Entity analysis identifies what is mentioned in the text and assigns it a type (PERSON, LOCATION, ORGANIZATION, CONSUMER_GOOD, EVENT, NUMBER, PRICE, OTHER) and a salience between 0 and 1 that indicates its importance within the text.

Entity sentiment combines both things. For the review:

"The jacket really is waterproof and very light, but the zip jams and the delivery took three weeks."

A plausible result:

Entity Type Salience score magnitude
jacket CONSUMER_GOOD 0.52 +0.8 1.2
zip CONSUMER_GOOD 0.24 −0.6 0.7
delivery OTHER 0.18 −0.7 0.9

This is actionable and the overall sentiment was not. The overall sentiment of that review would hover around 0 with high magnitude — "mixed" — correct but mute. The per-entity breakdown says three concrete things: the product delivers on its main promise, there is a component defect the supplier must correct, and there is a logistics problem.

Aggregating 8,000 reviews by entity gives you what no star rating does: what exactly fails and how often.

-- What gets mentioned with the worst sentiment in backpack reviews
SELECT
  entidad,
  COUNT(*)                                  AS menciones,
  ROUND(AVG(score), 3)                      AS score_medio,
  ROUND(AVG(salience), 3)                   AS relevancia_media,
  COUNTIF(score < -0.3)                     AS menciones_negativas
FROM `alpinashop-datos.alpinashop_analitica.opiniones_entidades` e
JOIN `alpinashop-datos.alpinashop_analitica.productos` p USING (sku)
WHERE p.categoria = 'mochila'
GROUP BY entidad
HAVING menciones >= 20
ORDER BY score_medio ASC
LIMIT 20;

The HAVING menciones >= 20 matters: an entity mentioned three times with dreadful sentiment is an anecdote, not a pattern. It is the same discipline as soporte in the heuristic recommender from 05-01.

  1. Content classification and syntactic analysis

Content classification assigns the text to categories from a general, hierarchical taxonomy (/Sports/Outdoors/Climbing & Mountaineering). It requires texts of a certain length — of the order of twenty words or more — and its taxonomy is general, not your shop's. For AlpinaShop it is of little value on reviews, which are short, but it is valuable for automatically classifying blog articles or long descriptions.

Syntactic analysis breaks the text down into tokens with their lemma, part of speech and dependency relation. It is rarely used as-is, but it has two very concrete practical applications:

  • Lemmatisation: grouping "weighs", "weighed" and "weighing" under the lemma "weigh" before counting frequencies. Without this, a word count scatters across variants.
  • Extracting adjective-noun patterns: "uncomfortable strap", "hard-wearing fabric", "loose stitching". Counting those pairs over 8,000 reviews produces a very dense summary of defects, with no trained model at all.

  1. A first call from Python

Preparation beforehand: enable the API with gcloud services enable language.googleapis.com --project=alpinashop-datos and create the sa-nlp-opiniones service account. Calling the API does not require a specific role; the permissions to grant are read and write on BigQuery (roles/bigquery.dataEditor).

And the call:

from google.cloud import language_v2

client = language_v2.LanguageServiceClient()

text = ("The jacket really is waterproof and very light, "
        "but the zip jams and the delivery took three weeks.")

document = language_v2.Document(
    content=text,
    type_=language_v2.Document.Type.PLAIN_TEXT,
    language_code="en",          # forcing the language avoids wrong detections
)

r = client.analyze_sentiment(
    request={"document": document,
             "encoding_type": language_v2.EncodingType.UTF8})

print(f"Overall -> score {r.document_sentiment.score:+.2f} "
      f"magnitude {r.document_sentiment.magnitude:.2f}")

for sentence in r.sentences:                 # per-sentence breakdown, no extra cost
    print(f"  [{sentence.sentiment.score:+.2f}] {sentence.text.content}")

e = client.analyze_entities(                 # entities: a second billed call
    request={"document": document,
             "encoding_type": language_v2.EncodingType.UTF8})

Three details that matter more than they look:

An explicit language_code. If it is not given, the API detects the language. In short texts detection fails frequently, and mistaking one language for a close one — Spanish for Portuguese, or Catalan for Spanish — degrades the result with no warning at all. If you know the language, declare it.

encoding_type=UTF8. It determines how character positions are counted in the response. With accents and ñ — that is, always in Spanish — using the wrong value shifts the offsets the API returns and breaks any logic that extracts fragments.

The per-sentence analysis (r.sentences) is free: it comes in the same response and in the same call. Storing it lets you locate exactly which sentence of a long review is the negative one, without calling again.

  1. Processing the 8,000 reviews: batches, quotas and errors

The API processes one document per call. Eight thousand reviews are eight thousand calls, and that is where a naive script crashes.

The three problems to solve: quotas (there is a limit of requests per minute which, if exceeded, returns a 429 error), transient errors (503, network drops) and cost (every call counts).

import time, random
from concurrent.futures import ThreadPoolExecutor
from google.cloud import language_v2
from google.api_core import exceptions

client = language_v2.LanguageServiceClient()

def analyse(review, max_attempts=5):
    """Analyses one review with retries and exponential backoff."""
    doc = language_v2.Document(
        content=review["texto"][:20000],    # trim anomalous texts
        type_=language_v2.Document.Type.PLAIN_TEXT,
        language_code=review.get("idioma", "en"),
    )
    for attempt in range(max_attempts):
        try:
            r = client.analyze_sentiment(
                request={"document": doc,
                         "encoding_type": language_v2.EncodingType.UTF8})
            return {
                "opinion_id": review["opinion_id"],
                "sku":        review["sku"],
                "score":      round(r.document_sentiment.score, 4),
                "magnitude":  round(r.document_sentiment.magnitude, 4),
                "n_frases":   len(r.sentences),
                "estado":     "OK",
            }
        except exceptions.ResourceExhausted:            # 429: quota
            wait = (2 ** attempt) + random.random()
            time.sleep(wait)
        except exceptions.ServiceUnavailable:           # 503: transient
            time.sleep(1 + attempt)
        except exceptions.InvalidArgument as err:       # text cannot be processed
            return {"opinion_id": review["opinion_id"], "sku": review["sku"],
                    "score": None, "magnitude": None, "n_frases": 0,
                    "estado": f"ERROR: {err.message[:120]}"}
    return {"opinion_id": review["opinion_id"], "sku": review["sku"],
            "score": None, "magnitude": None, "n_frases": 0,
            "estado": "ERROR: retries exhausted"}


with ThreadPoolExecutor(max_workers=8) as pool:
    results = list(pool.map(analyse, reviews))

The decisions in this code, one by one:

Exponential backoff with jitter (2 ** attempt + random()). If eight threads receive a 429 at the same time and all retry at exactly one second, they collide again. The random component desynchronises them. It is the same retry pattern as Pub/Sub in 04-04.

Distinguishing error types. ResourceExhausted is quota and gets retried with a longer wait. ServiceUnavailable is transient and gets retried quickly. InvalidArgument is a text the API cannot process — empty, in an unsupported language, too long — and retrying it is time and money wasted: it always fails the same way.

Recording the failure instead of aborting. The result always brings a row, with estado. A process over 8,000 documents dying at number 7,400 because one came in empty is unacceptable. At the end you count how many failed and why.

max_workers=8, not 100. With too much concurrency you exhaust the quota constantly, everything goes into retries and the process ends up being slower as well as more expensive. Eight threads is a conservative starting point that is worth raising while measuring.

Trimming the text. The API has a size limit per document. An anomalous 500 KB text — and they exist: somebody pastes something by mistake — fails and consumes quota. Cutting at 20,000 characters is a cheap defence.

  1. Real cost and how to estimate it before starting

The Natural Language API bills in units of 1,000 characters, rounding up per document and for each function requested.

The two practical consequences:

  • A 150-character review bills 1 unit, the same as a 990-character one. Rounding per document penalises short texts.
  • Asking for sentiment and entities on the same text is two billed calls, not one.

An estimate before spending anything, computed in SQL over the real data:

SELECT
  COUNT(*)                                             AS documentos,
  ROUND(AVG(LENGTH(texto)), 0)                         AS long_media,
  MAX(LENGTH(texto))                                   AS long_maxima,
  SUM(CEIL(LENGTH(texto) / 1000))                      AS unidades_sentimiento,
  SUM(CEIL(LENGTH(texto) / 1000)) * 2                  AS unidades_con_entidades
FROM `alpinashop-datos.alpinashop_analitica.v_opiniones_analitica`;

A typical result for AlpinaShop: 8,000 reviews, an average length of around 240 characters, almost all below 1,000. That is about 8,000 units for sentiment, or 16,000 if entities are also requested.

With prices of the order of one or two euros per 1,000 units — an order of magnitude, the current price must always be checked in the official documentation, and there is usually a monthly free tier — the full analysis of AlpinaShop's history costs of the order of tens of euros, once only.

Put that number next to the cost of training your own classifier: labelling 2,000 reviews by hand, training, deploying, maintaining. The comparison answers itself, and it is exactly the argument in section 1.

And the incremental flow, which is what makes the spend sustainable: the history is processed once; from then on only the new reviews are analysed, a few dozen a day. The recurring cost is practically zero.

-- Only what has not been analysed yet
SELECT o.opinion_id, o.sku, o.texto
FROM `alpinashop-datos.alpinashop_analitica.v_opiniones_analitica` o
LEFT JOIN `alpinashop-datos.alpinashop_analitica.opiniones_sentimiento` s
  USING (opinion_id)
WHERE s.opinion_id IS NULL;

  1. From BigQuery to opiniones_sentimiento and back

The architecture of the process, with the pieces you already know from module 4:

flowchart LR
    A[v_opiniones_analitica<br/>BigQuery] --> B[Python process<br/>batches + retries]
    B --> C[Natural Language API]
    C --> B
    B --> D[opiniones_sentimiento<br/>BigQuery]
    D --> E[Cross with orders<br/>and returns]
    E --> F[Looker Studio]
    G[Cloud Scheduler] --> H[Workflows] --> B

The results table, with partitioning and clustering as 04-01 requires:

CREATE TABLE IF NOT EXISTS `alpinashop-datos.alpinashop_analitica.opiniones_sentimiento` (
  opinion_id      STRING  NOT NULL,
  sku             STRING  NOT NULL,
  fecha           DATE,
  score           FLOAT64,
  magnitude       FLOAT64,
  n_frases        INT64,
  idioma          STRING,
  estado          STRING,
  modelo_version  STRING,          -- traceability: what analysed this
  procesado_en    TIMESTAMP
)
PARTITION BY fecha
CLUSTER BY sku;

The two columns almost nobody includes and that always end up being needed are modelo_version and procesado_en. Pre-trained APIs update on their own: the model analysing your texts today is not necessarily the one from a year ago. If in six months' time the average score values shift, the first question will be "did the model change or did the customers change?", and without those two columns it is impossible to answer.

Writing is done with insert_rows_json in batches of around 500 rows, adding modelo_version and procesado_en to each result before inserting, and always checking the error list the call returns (which does not raise an exception: if you do not look at it, the rejected rows are lost in silence).

The periodic run hooks into the orchestrator AlpinaShop already chose in 04-06: Cloud Scheduler fires a Workflow that runs the incremental process every night.

  1. The business question: does bad sentiment predict returns?

This is where the analysis stops being an exercise and starts being worth money.

WITH sentimiento_sku AS (
  SELECT
    sku,
    COUNT(*)                                       AS n_opiniones,
    ROUND(AVG(score), 3)                           AS score_medio,
    ROUND(AVG(magnitude), 3)                       AS magnitude_media,
    COUNTIF(score < -0.25)                         AS opiniones_negativas,
    ROUND(100 * COUNTIF(score < -0.25) / COUNT(*), 1) AS pct_negativas
  FROM `alpinashop-datos.alpinashop_analitica.opiniones_sentimiento`
  WHERE estado = 'OK'
  GROUP BY sku
  HAVING n_opiniones >= 15
),
devoluciones_sku AS (
  SELECT
    sku,
    COUNT(*)                                       AS unidades_vendidas,
    COUNTIF(devuelto)                              AS unidades_devueltas,
    ROUND(100 * COUNTIF(devuelto) / COUNT(*), 2)   AS tasa_devolucion
  FROM `alpinashop-datos.alpinashop_analitica.lineas_pedido`
  WHERE fecha_pedido >= DATE_SUB(CURRENT_DATE(), INTERVAL 2 YEAR)
  GROUP BY sku
  HAVING unidades_vendidas >= 50
)
SELECT
  s.sku, p.nombre, p.categoria,
  s.n_opiniones, s.score_medio, s.pct_negativas,
  d.unidades_vendidas, d.tasa_devolucion,
  ROUND(d.unidades_devueltas * p.precio_medio, 0) AS coste_devoluciones_eur
FROM sentimiento_sku s
JOIN devoluciones_sku d USING (sku)
JOIN `alpinashop-datos.alpinashop_analitica.productos` p USING (sku)
ORDER BY d.tasa_devolucion DESC
LIMIT 30;

And the global correlation, which is what answers the question:

SELECT
  COUNT(*)                                    AS skus_analizados,
  ROUND(CORR(score_medio, tasa_devolucion), 3) AS correlacion
FROM sentimiento_sku JOIN devoluciones_sku USING (sku);

How to read the result, without fooling yourself.

If the correlation comes out around −0.4, it means that the worse the sentiment, the higher the return rate, with a moderate relationship. That is already useful: sentiment works as an early warning signal. A new product with fifteen bad-sentiment reviews is probably going to generate returns before the returns show up in the data.

If it comes out close to 0, the conclusion is not that the analysis is worthless: it is that AlpinaShop's returns are explained by something else — most likely, with technical clothing, sizing — and that is also a valuable finding that changes where you have to act.

And the three compulsory cautions:

  1. Correlation is not causation. There may be an underlying variable — the category — that explains both: if clothing is returned more and also generates more critical reviews, the correlation appears without one thing causing the other. The check is to repeat the analysis within each category.
  2. Selection bias in who writes reviews. Only a fraction of customers write, and very happy and very angry people are usually over-represented. The average sentiment of the reviews is not the average sentiment of the customers.
  3. The n_opiniones >= 15 filter leaves out most of the catalogue. With 2,400 SKUs and 8,000 reviews, the average is three reviews per product: the per-SKU analysis is only reliable for the best sellers. For the rest, aggregating by category or by family is the only honest reading.

The practical output is a work list, exactly as in 05-02: the products with the worst sentiment and the highest return cost, ordered, so that somebody looks at the actual reviews and decides whether to talk to the supplier, correct the product page or change the size chart.

  1. Languages: Spanish, Catalan and what you have to check

The API supports a broad set of languages, but not all functions support the same languages, nor with the same quality. It is a common source of surprises.

Aspect Practical situation
Spanish Full support and high quality across all functions
Catalan Lower coverage than Spanish; it has to be verified function by function in the official documentation
Automatic detection Unreliable on short texts, which is exactly the case with reviews
Mixed texts A review that alternates Spanish and Catalan gives inconsistent results

AlpinaShop is a Catalan small business with customers all over Spain, so this point is not theoretical: some of the reviews will be in Catalan, and some will mix both languages in the same sentence.

The sensible strategy, in three steps:

  1. Store the language declared by the customer in the review form, if it exists. It is the most reliable information and it is free.
  2. If it does not exist, detect it once and store it, instead of letting the API guess on every call.
  3. Check the real coverage before processing in bulk: take 50 reviews in Catalan, analyse them, and compare against human judgement. If the result is not reliable, there are two reasonable ways out: translate into Spanish with Cloud Translation before analysing — it adds cost and some noise, but it works — or use Gemini (05-06), which handles Catalan comfortably.

And before anything else, a GROUP BY idioma over v_opiniones_analitica to know how much there is of each: it stops you building a complex solution for 2 % of the cases.

  1. The real limits: irony, negation and mountain jargon

An honest lesson has to say where the tool fails.

Irony and sarcasm. "Great, the tent held up perfectly… for three hours." The model sees "great" and "perfectly" and probably returns a positive score. Irony requires understanding the contradiction between form and content, and classic sentiment models do not capture it reliably.

Complex negations. "I wouldn't say it's a bad backpack, though it's not like I loved it either." Double negation plus qualification. The result will be something close to zero, which is not wrong but loses all the nuance.

Domain jargon. This is the most treacherous one, because it fails silently and on perfectly ordinary texts.

Phrase Generic reading Mountain reading
"The backpack is heavy" Negative It depends: on an expedition pack it is expected; on a trail pack, a serious defect
"The sleeping bag is only just good for 0 degrees" Ambiguous A warning: it does not do what it promises
"The boots are stiff" Negative Positive on crampon boots: rigidity is the requirement
"The fabric crackles" Neutral Negative: it indicates a stiff, noisy membrane
"Technical" Neutral Positive: a specialised product

The fourth and fifth rows are the problem in its pure form. The same word has the opposite sign depending on the product. A generic model knows nothing about crampons, and "stiff" always sounds bad to it.

This does not invalidate the tool: it bounds it. Generic sentiment analysis is good for prioritising and aggregating — which categories are doing worse, which products concentrate complaints, how sentiment evolves month by month. It is not good for deciding automatically about a specific product without anyone reading anything.

  1. Validating with a hand-labelled sample

Everything above leads to a single operational conclusion: you have to measure how often the API gets it right on your texts. Not on texts in general: on yours.

The protocol, which costs about three hours and is worth the whole lesson:

  1. Sample 200 reviews at random, stratifying by product category and by star rating, so that the sample is not all backpacks or all five stars.
  2. Label them by hand into four classes: positive, negative, neutral, mixed. Ideally two people independently, to measure how much they agree with each other: if two humans only agree 75 % of the time, you cannot demand 95 % from the machine.
  3. Compare with the API's classification using the thresholds from section 3.
  4. Adjust the thresholds to maximise agreement. It may be that in your texts the right cut-off is 0.15 and not 0.25.
  5. Repeat the validation every six months, because the underlying model changes without notice.
-- Agreement matrix between the human labelling and the API
SELECT
  m.etiqueta_humana,
  c.clasificacion                            AS etiqueta_api,
  COUNT(*)                                   AS casos
FROM `alpinashop-datos.alpinashop_analitica.muestra_etiquetada` m
JOIN `alpinashop-datos.alpinashop_analitica.v_opiniones_clasificadas` c
  USING (opinion_id)
GROUP BY 1, 2
ORDER BY 1, casos DESC;

A plausible result over AlpinaShop's reviews: good agreement on clear positives and negatives, and most of the disagreements concentrated on the boundary between "neutral" and "mixed" — exactly where section 3 warned — and on the reviews with technical jargon.

This exercise produces two things of permanent value. One is the number you can tell management: "the automatic classification agrees with human judgement in 82 % of cases". The other is a set of 200 labelled reviews that, if one day you decide to train your own classifier with AutoML (05-02), is already your starting point.

  1. An honest comparison: API, Gemini or your own classifier

Three ways of solving the same problem in 2026:

Criterion Natural Language API Gemini (05-06) AutoML Text (05-02)
Own data None None Hundreds of labelled examples
Getting started Hours Hours Days or weeks
Output Fixed: score, magnitude, entities Whatever you ask for in the prompt Your labels
Domain jargon Does not understand it It can be explained in the prompt It learns it from the data
Irony Badly Considerably better Depends on the examples
Catalan Limited coverage Good Depends on the data
Cost per document Low and very predictable Higher, per token Low after training
Latency Low Higher Low
Determinism High: same input, same output Variable High

When to use each, for AlpinaShop:

Natural Language API for the bulk processing of the history and the daily flow. It is cheap, fast, predictable and gives a number that is comparable over time. It is the choice for the 8,000 reviews.

Gemini when the question is richer than "positive or negative?". For example: "classify this review into product / sizing / delivery / customer service, give the sentiment of each aspect and extract the specific defect if there is one, returning JSON". That, which the API cannot do, Gemini does with a prompt and without training anything. And it understands irony and Catalan far better. In exchange, it costs more per document, it is slower and it is not deterministic: the same review can give slightly different answers. We see it in 05-06.

AutoML Text only if one very specific condition holds: that there is a stable taxonomy of your own — "stitching defect", "sizing problem", "logistics delay", "description error" — and that labelling a thousand examples pays off. With the 200 from section 13 you are already halfway there.

And a combination that is often the best of the three: use the API to process everything cheaply, and reserve Gemini for the interesting subset — the negative and the mixed reviews, which will be a few hundred — asking it for a detailed aspect-by-aspect analysis. Low cost over the volume, high quality where it matters.

  1. The full family: Translation, Speech-to-Text, Document AI

The Natural Language API is one of several pre-trained APIs that share a philosophy, pay-per-use billing and a total absence of training.

API What it does Possible use at AlpinaShop
Cloud Translation Translates between languages Translating product pages into Catalan and English; homogenising reviews before analysing them
Speech-to-Text Transcribes audio Transcribing customer service calls and analysing them like the reviews
Text-to-Speech Generates speech Marginal here
Document AI Extracts data from documents Supplier invoices and delivery notes (detailed in 05-05)

Cloud Translation has two editions: the basic one, which translates plain text, and the advanced one, which allows glossaries and batch translation. The glossary matters a great deal in this domain: without it, "ice axe" or "harness" can be translated incorrectly or inconsistently across product pages. With a glossary, the terminology is pinned down.

Speech-to-Text opens up an interesting possibility: AlpinaShop has a customer service phone line whose conversations leave no analysable data at all. Transcribing them and putting them through the same sentiment analysis would give a second source on the same problems. With a first-order legal warning: recording and transcribing calls requires prior notice, a legal basis and a clear retention policy, and the transcript contains personal data in abundance. Before touching anything, go through compliance and apply the de-identification from 04-07.

  1. Privacy: what gets sent, what gets stored and what the GDPR says

What gets sent. The complete text of the document travels to the API. Even though the infrastructure is inside Google Cloud, the content leaves your project and enters a managed service.

Where it is processed. Pre-trained APIs offer regional endpoints for certain functions and languages. If data residency in the EU is a requirement, you have to verify in the official documentation which regional endpoints are available for the specific function and language, and configure them explicitly. Do not assume that because your project is in europe-west1 the processing happens there.

What gets stored. Google Cloud's current policy for prediction APIs is not to use customer data to train its models, and there are contractual commitments to that effect. It is a point that has to be verified in the current terms and documented in the record of processing activities: it is not something that can be taken as read.

The specific risk of free text. This is the critical point and it deserves saying clearly: a free-text field is where the personal data nobody expected ends up. A customer writes their name, their phone number so they can be called back, the address where the parcel was left, or somebody else's name. No form validation prevents it.

AlpinaShop already did the right thing in 04-07: the reviews are de-identified with Sensitive Data Protection before becoming available for analysis, replacing each finding with its type ([EMAIL_ADDRESS], [PERSON_NAME]), which preserves the structure and the sentiment of the text. And the analysis in this lesson runs over v_opiniones_analitica, the de-identified view, never over the original table. That is the control that makes the rest defensible.

With the usual warning: automatic de-identification is not infallible. A customer who writes "I'm the one who bought the blue backpack on Tuesday at the Sabadell shop" will not be detected by any predefined type and remains potentially re-identifiable.

Express recommendation. Before processing customer text in bulk, a compliance professional or the DPO must determine: the legal basis for the analytical processing of the reviews and whether the purpose is covered by the information given when they were collected; whether sending the text to the API constitutes a disclosure of data requiring additional information; the data residency requirements and which regional endpoints meet them; the retention period for the texts and for the results; and whether the de-identification applied is sufficient. This lesson describes technical controls and does not constitute legal advice. All data is fictitious.

On the EU AI Act: analysing the sentiment of product reviews to improve the catalogue reasonably sits at a low risk level. It is worth pointing out, however, that the regulation contemplates specific restrictions for emotion recognition systems applied to people in contexts such as the workplace or education. Analysing the tone of a text about a backpack is not that, but the boundary matters: if tomorrow somebody proposed applying the same analysis to the customer service team's communications to assess their performance, the legal framing would change completely. Any use aimed at people rather than at products requires prior legal review.

Common Mistakes and Tips

Looking only at the score. The mistake of this lesson. A score of 0 with a magnitude of 3.4 is an intense, mixed review, not a neutral one, and those are the ones that contain the most information.

Not declaring the language. Automatic detection fails on short texts and degrades the result with no warning.

Comparing magnitude across texts of very different lengths. It is not normalised: a long text accumulates more magnitude even if it is less intense.

Retrying InvalidArgument errors. A text the API cannot process will always fail the same way. Retrying it costs time and quota.

Launching 100 threads. It exhausts the quota, triggers retries and ends up slower and more expensive than going with eight.

Requesting every function "just in case". Each function is billed separately. Ask for sentiment and entities if you are going to use both; if not, only one.

Analysing SKUs with three reviews. With 8,000 reviews and 2,400 products, the average is three per SKU. For most of the catalogue, the only honest reading is by category.

Forgetting modelo_version and procesado_en. Without them it is impossible to tell a change in the customers from a change in the provider's model.

Tip: always store the per-sentence sentiment. It comes free in the same response and lets you locate the problematic sentence of a long review without calling again.

Tip: validate with 200 hand-labelled reviews before showing management a single number. It is the difference between "the average sentiment is 0.31" and "the classification agrees with human judgement in 82 % of cases, and it fails mainly on irony and on technical jargon".

Exercises

Exercise 1

Lucía tells management that "71 % of AlpinaShop's reviews are neutral" and concludes that customers are indifferent. Diagnose what may have happened, what query you would run to check it, and what you would propose she present instead.

Exercise 2

Design the complete process for analysing the new reviews every night: architecture, destination table, cost control and what to do when a review fails. State which services from previous modules you reuse.

Exercise 3

The analysis shows that product BOT-3391 (a stiff mountain boot) has an average sentiment of −0.18, among the worst in the catalogue, but its return rate is 2.1 %, well below the 6.4 % average. Explain the apparent contradiction and propose how to investigate it.

Solutions

Solution 1

Diagnosis: almost certainly, she is classifying only by score and ignoring magnitude.

Section 3 explains the exact mechanism. Mixed reviews — "the product great, the delivery terrible" — have the two sentiments cancelling each other out and produce a score close to 0. If the criterion is "neutral if the score is between −0.25 and +0.25", all those intense reviews fall into the neutral bucket.

A possible secondary cause: if the language was not declared and some of the reviews are in Catalan, detection may have failed and returned poorly discriminating values.

Query to check it:

SELECT
  CASE WHEN ABS(score) < 0.25 THEN 'score_neutro' ELSE 'score_polarizado' END AS grupo,
  CASE WHEN magnitude < 0.6 THEN 'baja'
       WHEN magnitude < 1.5 THEN 'media'
       ELSE                      'alta' END                                   AS emocion,
  COUNT(*)                                     AS opiniones,
  ROUND(AVG(puntuacion), 2)                    AS estrellas_medias,
  ROUND(AVG(LENGTH(texto)), 0)                 AS long_media
FROM `alpinashop-datos.alpinashop_analitica.opiniones_sentimiento` s
JOIN `alpinashop-datos.alpinashop_analitica.v_opiniones_analitica` o USING (opinion_id)
WHERE estado = 'OK'
GROUP BY grupo, emocion
ORDER BY grupo, emocion;

What to expect. If inside score_neutro there is a large block with high emotion, there is the answer: they are not neutral, they are mixed. And the figure that confirms it independently is estrellas_medias: if that group has an average of 2.8 stars and long texts, it is obvious they are not indifferent customers.

What to present instead:

  1. The four-class classification — positive, negative, neutral, mixed — with the view from section 3. It is very likely that the real "neutrals" drop from 71 % to 25-30 %.
  2. The breakdown of the mixed ones by entity, which is the actionable information: "in the mixed reviews, the product has an average sentiment of +0.6 and the delivery −0.5". That does not say customers are indifferent: it says the product is liked and the logistics are failing, which is a conclusion with actions attached.
  3. The validation with 200 hand-labelled reviews, with the corresponding sentence of honesty: "the automatic classification agrees with human judgement in X % of cases".

The underlying lesson: a misinterpreted number leads to the most mistaken business conclusion possible — "customers don't care" — when the data says exactly the opposite.

Solution 2

Architecture, reusing what module 4 gave us:

flowchart LR
    A[Cloud Scheduler<br/>03:30 daily] --> B[Workflows]
    B --> C[Incremental query<br/>in BigQuery]
    C --> D[Cloud Run Job<br/>Python process]
    D --> E[Natural Language API]
    E --> D
    D --> F[opiniones_sentimiento]
    D --> G[opiniones_fallidas]
    B --> H[Quality check]
    H -->|failure| I[Alert to gcp-datos@]

Reused services: Cloud Scheduler and Workflows as the orchestrator (04-06, a decision already taken over Composer), BigQuery as source and destination (04-01) with the de-identified view from 04-07, Cloud Run as the process runner (detail in 07-02), Cloud Monitoring for the alert (06-04) and Secret Manager if any credential were needed (03-06).

Why a Cloud Run Job and not a function: the process can take several minutes with retries, and a Cloud Run job accepts long executions without a function's time restrictions.

Incremental selection, the query from section 8, with a safety cap:

SELECT o.opinion_id, o.sku, o.fecha, o.texto, o.idioma
FROM `alpinashop-datos.alpinashop_analitica.v_opiniones_analitica` o
LEFT JOIN `alpinashop-datos.alpinashop_analitica.opiniones_sentimiento` s
  USING (opinion_id)
WHERE s.opinion_id IS NULL
LIMIT 5000;   -- safety cap: if something runs wild, the cost does not explode

The LIMIT 5000 is deliberate. If through a data error 400,000 unprocessed reviews appeared, the cap avoids an unexpected invoice. Whatever is left over will be processed the following night, and the quality check will warn that the queue is growing.

Destination table: the one from section 9, partitioned by fecha and clustered by sku, with modelo_version and procesado_en.

What to do when a review fails:

Failure type Action
429 quota Retry with exponential backoff and jitter, up to 5 times
503 transient Quick retry
InvalidArgument Do not retry. Row with estado='ERROR: ...'
Empty text or < 10 characters Filter before calling; it does not consume quota
Retries exhausted Row with an error status; the next run does not retry it automatically

The key: every review read produces a row, with a result or with an error. That way opinion_id IS NULL still means "not processed" and not "it failed and will be retried forever", which is how you build an infinite cost loop.

Cost control:

  1. Filter out trivial texts before calling (LENGTH(texto) >= 10).
  2. LIMIT 5000 as a hard cap.
  3. Request sentiment only in the daily flow; entity analysis is reserved for the negative and mixed reviews, which are a fraction.
  4. The centro-coste:analitica billing label and a budget alert (01-04).

Quality check in the Workflow, which blocks if something goes wrong:

SELECT
  COUNTIF(estado = 'OK')                                        AS ok,
  COUNTIF(estado != 'OK')                                       AS fallidas,
  SAFE_DIVIDE(COUNTIF(estado != 'OK'), COUNT(*))                AS tasa_fallo
FROM `alpinashop-datos.alpinashop_analitica.opiniones_sentimiento`
WHERE DATE(procesado_en) = CURRENT_DATE();

If tasa_fallo > 0.05, the Workflow marks a failure and alerts gcp-datos@. It is the same quality discipline that blocks the nightly pipeline in 04-07.

Solution 3

There is no contradiction. There are two plausible explanations, and both are good news badly told.

Explanation 1 (the most likely): domain jargon. It is exactly the case in the table in section 12. A stiff mountain boot generates reviews like "they're incredibly stiff", "it's hard work walking on the flat in them", "very rigid", "they're heavy". The generic model reads "stiff", "hard work", "heavy" and returns negative sentiment. For a crampon-compatible boot, rigidity is the functional requirement, not a defect. The customer who writes "they're incredibly stiff" is describing with satisfaction that the product does what it promises.

And the returns confirm it: whoever buys this product knows what they are buying and keeps it.

Explanation 2: the complaint is about the breaking-in process, not about the product. Stiff boots require an uncomfortable breaking-in period. The reviews can be negative about that experience — "the first few days were torture" — and positive about the final result. The overall sentiment picks up the negative part.

How to investigate it, in three steps:

Step 1 — Read. Pull the 20 most negative reviews for the SKU and read them. It costs fifteen minutes and very often settles the case. No automatic analysis replaces this.

SELECT o.texto, s.score, s.magnitude, o.puntuacion
FROM `alpinashop-datos.alpinashop_analitica.opiniones_sentimiento` s
JOIN `alpinashop-datos.alpinashop_analitica.v_opiniones_analitica` o USING (opinion_id)
WHERE s.sku = 'BOT-3391' AND s.estado = 'OK'
ORDER BY s.score ASC LIMIT 20;

Step 2 — Contrast with the stars. The numeric rating is a human label that already exists and is free:

SELECT
  ROUND(AVG(o.puntuacion), 2)              AS estrellas_medias,
  ROUND(AVG(s.score), 3)                   AS score_medio,
  ROUND(CORR(o.puntuacion, s.score), 3)    AS correlacion_sku,
  COUNT(*)                                 AS n
FROM `alpinashop-datos.alpinashop_analitica.opiniones_sentimiento` s
JOIN `alpinashop-datos.alpinashop_analitica.v_opiniones_analitica` o USING (opinion_id)
WHERE s.sku = 'BOT-3391' AND s.estado = 'OK';

This is the decisive step. If the average stars are 4.3 and the average score is −0.18, the verdict is unambiguous: the problem is with the analysis, not with the product. Customers are happy and the model is reading it backwards.

Step 3 — Compare within the category. Repeat the previous query grouping by categoria and subcategoria with HAVING opiniones >= 30, ordering by correlation ascending. If every stiff boot in the catalogue shows the same mismatch between stars and score, it is not an isolated case: it is a systematic bias of the model in that category. The categories with low or negative correlation between stars and score are the ones where the generic analysis cannot be trusted.

What to do with the finding:

  1. Short term: do not use the absolute score to compare across categories. Compare each product against its category's average, not against the global average. A score of −0.18 in a category whose average is −0.22 is actually among the best.
  2. Medium term: use Gemini for the problematic categories (05-06), explaining the context in the prompt: "These are reviews of stiff mountain boots. In this context, 'stiff', 'rigid' and 'heavy' can be positive descriptions of technical performance. Classify the customer's real sentiment." That fixes the problem without training anything.
  3. Long term: if the problem is widespread, train your own classifier with AutoML Text (05-02) using the stars as a weak label and a hand-validated sample. There it would pay off.

And the general lesson: the best detector of failures in sentiment analysis is a human label you already have. The stars were in the database from the start. Crossing a model's output with an independent human signal is the cheapest and most effective check there is.

Conclusion

You have applied the principle that saves most money in applied machine learning: do not train what is already trained. Eight thousand reviews that had gone five years without anyone reading them are now analysed, crossed with sales and available in BigQuery, for a few tens of euros and without a single minute of training.

You know the five functions of the Natural Language API — sentiment, entities, entity sentiment, content classification and syntax — and why the third is the most valuable for a shop: a real review is almost never plainly good or bad, it is "the product great, the delivery terrible", and only the per-entity breakdown turns that into something actionable.

And above all you understand score and magnitude, which is where almost everybody goes wrong. A score of 0 with a magnitude of 0.1 is a descriptive text with no emotion; a score of 0 with a magnitude of 3.4 is an intense review with a serious problem inside. Confusing them leads to the most mistaken conclusion possible — "customers don't care" — exactly when the data says the opposite.

You have built the processing of the 8,000 reviews with everything a process of this kind needs to survive reality: exponential backoff with jitter for the quotas, a distinction between retryable and definitive errors, one result row per document read even when it fails, moderate concurrency and defensive trimming of anomalous texts. You know how to estimate the cost before spending it with a SQL query, and why rounding per document and per function is what really determines the invoice.

The results live in opiniones_sentimiento, partitioned and clustered, with modelo_version and procesado_en — the two columns that make it possible to tell, in six months' time, whether the model changed or the customers did. And the nightly incremental process hooks into Cloud Scheduler and Workflows, the orchestrator AlpinaShop already chose in 04-06, with a safety cap and a quality check that blocks if the failure rate shoots up.

You have answered the business question by crossing sentiment with returns, and you know how to read the result with the three cautions: correlation is not causation, whoever writes reviews is not a representative sample, and with three reviews per SKU on average the only honest reading for most of the catalogue is by category.

And you know the real limits, said without embellishment: irony escapes it, complex negations get diluted, and mountain jargon plays nasty tricks on it — "stiff" is a defect in a pair of trainers and a virtue in crampon boots. That is why the protocol of validating with 200 hand-labelled reviews is not optional: it is what turns "the average sentiment is 0.31" into "it agrees with human judgement in 82 % of cases, and it fails on irony and technical jargon". And you have the cheapest trick of all: crossing the model's output with the stars the customers had already given.

You have the honest comparison between the API, Gemini and your own classifier, with the combination that usually wins — the API for the volume, Gemini for the interesting subset — and you know that Translation, Speech-to-Text and Document AI share the same philosophy. And you have the privacy framework clear: the analysis runs over the de-identified view from 04-07, free text is where the personal data nobody expected ends up, data residency has to be verified per function and language, and any use aimed at people instead of products changes the legal framing completely.

There is one question from module 4 still unanswered, and it is the bulkiest: 60 GB of images in alpinashop-catalogo. AutoML classified them by category in 05-02, but a great deal more can be extracted from each photo than what is in it: which objects appear and where, what text is printed on it, which colours dominate, whether it is fit to publish. And there is a new problem: customers are starting to upload their own photos with their reviews, and nobody is moderating them.

In 05-05, the Vision API, you apply the same principle to images. You will see Cloud Vision's functions — labels, objects with bounding boxes, OCR, logos, landmarks, dominant colours, SafeSearch and web detection — you will process the 60 GB in bulk with asyncBatchAnnotate while controlling the cost per thousand images, and you will turn them into four concrete things: the accessibility alt text that does not exist today, the shop's colour filter, automatic moderation of customer photos and the detection of product pages that do not meet the catalogue standard. With the event-driven architecture over the imagenes-subidas topic, and with the module's most serious legal warning, which arrives when a face appears in an image.

Google Cloud Platform (GCP) Course

Module 1: Introduction to Google Cloud Platform

Module 2: Core GCP Services

Module 3: Networking and Security

Module 4: Data and Analytics

Module 5: Machine Learning and AI

Module 6: DevOps and Monitoring

Module 7: Advanced GCP Topics

Module 8: Final Project

© Copyright 2026. All rights reserved