In the alpinashop-catalogo bucket there are 60 GB of images. Around 2,400 product pages, each with between five and ten photos, organised into productos/<sku>/original/, productos/<sku>/web/ and productos/<sku>/thumb/.

About those images, AlpinaShop knows exactly one thing: which folder they are in. Nothing else.

And that translates into four concrete problems that are costing money right now:

The site is not accessible. No image has an alt attribute. A screen reader announces "image" and that is it. As well as being a real barrier for customers with a visual impairment, it is a search ranking problem and a possible compliance issue.

There is no colour filter. Marketing has spent a year asking for the catalogue to be filterable by colour. The color field on the product page is filled in by the supplier whenever they feel like it, with values such as "azul", "azul marino", "navy", "blau" and in many cases empty.

Nobody moderates customer photos. For a few months now, reviews have accepted photos. They are published directly, with no review. It is a matter of time before somebody uploads something inappropriate and it appears on a product page.

The catalogue is visually inconsistent. Some photos have a white background, others the floor of the supplier's warehouse. Nobody has audited 20,000 images.

All four are solved with the same tool and without training anything. It is the same principle as the previous lesson — do not train what is already trained — now applied to pixels.

Contents

  1. What the Cloud Vision API offers
  2. A first call and reading the response
  3. Bulk processing with asyncBatchAnnotate
  4. From the output JSON to alpinashop_analitica
  5. Cost per thousand images and how to control it
  6. Application 1: accessibility alt text
  7. Application 2: dominant colours and the shop's filter
  8. Application 3: moderating customer photos with SafeSearch
  9. Application 4: auditing the catalogue's visual standard
  10. Event-driven architecture with imagenes-subidas
  11. OCR: text in images and documents
  12. Limits and biases: where it fails with mountain gear
  13. When to make the jump to AutoML Vision
  14. Document AI for invoices and delivery notes
  15. Faces, biometrics and the AI Act: the serious warning

  1. What the Cloud Vision API offers

A single API with several functions that are requested separately. Each function requested is billed independently, so it is worth knowing which one is good for what.

Function What it returns Use at AlpinaShop
Label detection (LABEL_DETECTION) Concepts present in the image, with confidence The basis of the alt text
Object detection (OBJECT_LOCALIZATION) Objects with their bounding box and confidence Checking framing and detecting stray objects
OCR (TEXT_DETECTION) Text in the image with its position Reading sizes, brands and codes in the photos
Document OCR (DOCUMENT_TEXT_DETECTION) Text structured into pages, blocks and paragraphs Scanned documents, not product photos
Logos (LOGO_DETECTION) Recognised commercial brands Verifying that the brand matches the product page
Landmarks (LANDMARK_DETECTION) Famous places Marginal: mountain scenery photos
Image properties (IMAGE_PROPERTIES) Dominant colours in RGB with pixel fraction The colour filter
SafeSearch (SAFE_SEARCH_DETECTION) Probability of adult, violent, medical content, etc. Moderating customer photos
Web detection (WEB_DETECTION) Similar images on the web and associated entities Detecting unauthorised use of your own photos
Crop hints (CROP_HINTS) Recommended crops by aspect ratio Generating well-framed thumbnails
Faces (FACE_DETECTION) Position and attributes of faces See section 15 before using it

Two important clarifications from the outset.

Label detection and object detection are not the same thing. The first says what concepts are there ("backpack", "mountain", "blue", "equipment"). The second says what objects there are and where, with coordinates. To describe an image the first is enough; to check whether the product is centred and fills the frame correctly, you need the second.

Web detection is the least known and has a very concrete use for a shop: finding out whether your catalogue photos are being used by competitors. It is not the goal of this lesson, but it is worth knowing it exists.

  1. A first call and reading the response

After enabling the API (gcloud services enable vision.googleapis.com --project=alpinashop-datos):

from google.cloud import vision

client = vision.ImageAnnotatorClient()

image = vision.Image()
image.source.image_uri = "gs://alpinashop-catalogo/productos/MOC-4471/web/01.jpg"

response = client.annotate_image({
    "image": image,
    "features": [
        {"type_": vision.Feature.Type.LABEL_DETECTION,      "max_results": 10},
        {"type_": vision.Feature.Type.OBJECT_LOCALIZATION,  "max_results": 5},
        {"type_": vision.Feature.Type.IMAGE_PROPERTIES},
        {"type_": vision.Feature.Type.SAFE_SEARCH_DETECTION},
    ],
})

if response.error.message:
    raise RuntimeError(response.error.message)

A key note: the image is referenced by its Cloud Storage URI, it is not uploaded in the request. The API reads it directly from the bucket. That means there is no need to download 60 GB anywhere, but it also means the service account needs read permission on the bucket. It is configuration failure number one in this lesson.

The response, field by field:

# --- Labels ---
for label in response.label_annotations:
    print(f"{label.description:25} score={label.score:.3f}  mid={label.mid}")
# Backpack         score=0.968  mid=/m/0n1cs
# Luggage          score=0.944
# Equipment        score=0.911
# Blue             score=0.887
  • description: the concept, in the language you ask for. It is controlled with image_context.language_hints=["en"].
  • score: confidence between 0 and 1. Below 0.6 the labels are usually noise.
  • mid: a stable identifier for the concept in Google's knowledge graph. It is more reliable than the text for grouping: /m/0n1cs is always the same thing even if the translation changes.

The objects arrive in response.localized_object_annotations, each with its name, its score and a bounding_poly.normalized_vertices — for example, Backpack 0.921 x[0.18-0.83] y[0.09-0.94]. The coordinates are normalised: they run from 0 to 1 relative to the image's width and height, not in pixels. That makes them independent of resolution, which is very convenient because the same threshold works for the original photo and for the thumbnail. With those four numbers you work out what fraction of the frame the product occupies and whether it is centred, which is exactly what is needed in section 9.

The dominant colours arrive in response.image_properties_annotation.dominant_colors.colors, each with its RGB, its pixel_fraction and its score. They are two fields that get confused: pixel_fraction is the proportion of pixels of that colour and score is the estimated visual relevance. In a studio photo, the first result is usually RGB(245,245,244) with a fraction of 0.61 — the white background — and the product's real colour comes behind. That is exactly the problem in section 7.

And SafeSearch is read in response.safe_search_annotation, with the fields adult, violence, racy, medical and spoof:

s = response.safe_search_annotation
print(s.adult.name, s.violence.name, s.racy.name, s.medical.name, s.spoof.name)
# VERY_UNLIKELY VERY_UNLIKELY UNLIKELY VERY_UNLIKELY VERY_UNLIKELY

SafeSearch does not return a boolean: it returns one of six levels (UNKNOWN, VERY_UNLIKELY, UNLIKELY, POSSIBLE, LIKELY, VERY_LIKELY) for each of five categories. Where you put the cut-off is an editorial policy decision, not a technical one — section 8.

  1. Bulk processing with asyncBatchAnnotate

Calling annotate_image 20,000 times works, but it is slow, it handles quotas badly and it saturates the local process. For volume there is asyncBatchAnnotate: an asynchronous job that processes a batch of images and writes the results straight into Cloud Storage, without your process having to wait.

from google.cloud import vision

client = vision.ImageAnnotatorClient()

def build_request(uri):
    return vision.AnnotateImageRequest(
        image=vision.Image(source=vision.ImageSource(image_uri=uri)),
        features=[
            vision.Feature(type_=vision.Feature.Type.LABEL_DETECTION, max_results=10),
            vision.Feature(type_=vision.Feature.Type.IMAGE_PROPERTIES),
            vision.Feature(type_=vision.Feature.Type.OBJECT_LOCALIZATION, max_results=5),
        ],
        image_context=vision.ImageContext(language_hints=["en"]),
    )

output = vision.OutputConfig(
    gcs_destination=vision.GcsDestination(
        uri="gs://alpinashop-datalake/vision/catalogo/"),
    batch_size=100,           # images per output JSON file
)

operation = client.async_batch_annotate_images(
    requests=[build_request(u) for u in batch_uris],   # up to 2000 per operation
    output_config=output,
)

result = operation.result(timeout=3600)

The four points to understand:

It really is asynchronous. async_batch_annotate_images returns a long-running operation. You can wait for it with operation.result() or keep the operation's name and check it later. For 20,000 images, the sensible thing is the latter, letting a Workflow supervise it.

The results go to Cloud Storage, not to your process. They are written as JSON files under the given prefix. Your memory never sees 20,000 responses.

batch_size groups the responses into files. With 100, each JSON contains 100 annotations. Very small files generate thousands of objects that then have to be listed; very large ones are awkward to process. Between 50 and 200 is reasonable.

There is a cap on requests per operation (of the order of 2,000). For AlpinaShop's 20,000 images you have to split into several operations, launching them in a staggered way.

And the image list is prepared with a gcloud storage ls --recursive "gs://alpinashop-catalogo/productos/**/web/*.jpg", applying two filters that save money:

Only web/. The original/, web/ and thumb/ folders contain the same image at three resolutions. Analysing all three multiplies the cost by three while contributing absolutely nothing. This single filter divides the invoice by three.

Only one or two photos per SKU for the initial analysis. If the goal is to describe the product and extract its colour, the main photo is enough. The others are only processed if they are needed for the framing audit.

  1. From the output JSON to alpinashop_analitica

The JSON files in Cloud Storage are no use until they are queryable. The bridge is a bq load --source_format=NEWLINE_DELIMITED_JSON --autodetect over the output prefix. In practice, the output of asyncBatchAnnotate comes nested as a responses array, so a flattening step beforehand is usually needed: for moderate volumes, a Python script that reads the JSON and writes flat rows; for large volumes, a Dataflow job (04-02).

The normalised destination table:

CREATE OR REPLACE TABLE `alpinashop-datos.alpinashop_analitica.imagenes_vision`
PARTITION BY fecha_analisis
CLUSTER BY sku AS
SELECT
  REGEXP_EXTRACT(uri, r'productos/([^/]+)/')            AS sku,
  uri,
  ARRAY(SELECT AS STRUCT descripcion, score, mid
        FROM UNNEST(etiquetas) WHERE score >= 0.60)     AS etiquetas,
  colores,
  objetos,
  safesearch_adult, safesearch_violence, safesearch_racy,
  CURRENT_DATE()                                        AS fecha_analisis,
  'vision-api-v1'                                       AS modelo_version
FROM `alpinashop-datos.alpinashop_analitica.raw_vision_catalogo`;

The score >= 0.60 filter inside the ARRAY is deliberate: labels below that confidence are almost always generic noise ("product", "object", "photograph") that pollutes any later analysis.

And once again modelo_version and fecha_analisis, for the same reason as in 05-04: pre-trained APIs update on their own, and without those columns it is impossible to know whether a change in the results comes from the model or from the images.

The immediate check is a GROUP BY over UNNEST(etiquetas) counting images per description. That listing answers at a glance "what does the machine see in my catalogue?" and usually holds surprises: generic labels dominating the ranking, or unexpected concepts that reveal wrongly classified photos.

  1. Cost per thousand images and how to control it

The Cloud Vision API bills per unit, where a unit is one image for each function requested. It is the same logic as in 05-04 and it has the same consequence: asking for four functions on one image is four units.

The rates are structured in tiers — there is a monthly free tier and the price per thousand drops as volume increases — and they change over time: always check the current official pricing calculator. As an order of magnitude, the cost per thousand units sits around a few euros for the usual functions.

The estimate for AlpinaShop, made before spending:

Scenario Images Functions Units
The whole bucket, all functions 20,000 5 100,000
Only web/, all functions ~7,000 5 35,000
Only web/, main photo, 3 functions 2,400 3 7,200
New photos per month ~200 3 600

The difference between the first and the third row is a factor of almost 14. And the third option covers the four problems from the start of the lesson. That is the real work of this section: not negotiating the price, but not asking for what you do not need.

The four control measures:

  1. Filter by folder. Only web/. Divides by three.
  2. One photo per SKU for labels and colour. Divides by as much again.
  3. Request only the necessary functions. SafeSearch makes no sense on the supplier's catalogue photos: it is reserved for the photos customers upload.
  4. Do not reprocess. Store the result and analyse only what is new, with the same incremental query as in 05-04.

The incremental query is the same idea as in 05-04: a LEFT JOIN of the image inventory against imagenes_vision, filtering v.uri IS NULL and carpeta = 'web', with a LIMIT 2000 as a safety cap.

And the billing labels (centro-coste:analitica) with their budget alert, as throughout the module.

  1. Application 1: accessibility alt text

The problem: 20,000 images with no alt. The solution using labels is direct, but it has to be done properly.

CREATE OR REPLACE TABLE `alpinashop-datos.alpinashop_analitica.imagenes_alt` AS
WITH etiquetas_utiles AS (
  SELECT
    v.sku, v.uri,
    ARRAY_AGG(e.descripcion ORDER BY e.score DESC LIMIT 3) AS conceptos
  FROM `alpinashop-datos.alpinashop_analitica.imagenes_vision` v,
       UNNEST(v.etiquetas) e
  WHERE e.score >= 0.75
    AND LOWER(e.descripcion) NOT IN
        ('product','object','photograph','image','background','white')
  GROUP BY v.sku, v.uri
)
SELECT
  t.sku, t.uri,
  CONCAT(p.nombre, '. ',
         ARRAY_TO_STRING(t.conceptos, ', '),
         '. Brand ', p.marca, '.')            AS alt_propuesto,
  ARRAY_LENGTH(t.conceptos)                   AS n_conceptos,
  FALSE                                       AS revisado
FROM etiquetas_utiles t
JOIN `alpinashop-datos.alpinashop_analitica.productos` p USING (sku);

The three decisions in this query:

Start with the product name, not with the labels. The most useful alt for a screen reader is "Trek 40L Blue Backpack. Backpack, mountain equipment, blue. Brand Alpina." The information the shop already has is more precise than what the model guesses; the API complements it, it does not replace it.

A blacklist of generic labels. "Product", "object" or "photograph" contribute nothing to anybody and turn up constantly.

revisado = FALSE by default. It is a proposal, not published text. And here comes the important part.

An incorrect alt is worse than no alt at all. A blind person who hears "jacket, textile, red" over a photo of a blue backpack receives false information, with no way of detecting the error. The absence of alt is a gap; a wrong alt is deception.

The correct flow is staged: automatically publish only the high-confidence proposals — the ones where the main label matches the product page's category — and send the rest to human review. With 2,400 product pages, if 80 % match, around 480 are left to review: an afternoon's work versus a project of weeks, and the result is correct.

And a note on scope: the alt this technique generates is descriptive and correct, but flat. More natural, richer wording is achieved by combining the labels with Gemini, which sees the image and writes the sentence. That is 05-06.

  1. Application 2: dominant colours and the shop's filter

The color field on the product pages is a mess of free-form values. The API's dominant colours are objective and consistent, but they have an obvious problem: the background is usually the dominant colour.

CREATE OR REPLACE TABLE `alpinashop-datos.alpinashop_analitica.productos_color` AS
WITH dominante AS (
  SELECT
    v.sku, c.red, c.green, c.blue, c.pixel_fraction,
    ROW_NUMBER() OVER (PARTITION BY v.sku ORDER BY c.score DESC) AS pos
  FROM `alpinashop-datos.alpinashop_analitica.imagenes_vision` v,
       UNNEST(v.colores) c
  WHERE NOT (c.red > 225 AND c.green > 225 AND c.blue > 225)   -- white background
    AND NOT (c.red < 30  AND c.green < 30  AND c.blue < 30)    -- black background
    AND c.pixel_fraction >= 0.05
)
SELECT
  sku, red, green, blue, ROUND(pixel_fraction, 3) AS fraccion,
  CASE
    WHEN red > 150 AND green < 90  AND blue < 90        THEN 'red'
    WHEN blue > 130 AND red < 110 AND green < 130       THEN 'blue'
    WHEN green > 120 AND red < 120 AND blue < 120       THEN 'green'
    WHEN red > 190 AND green > 140 AND blue < 90        THEN 'orange'
    WHEN red > 190 AND green > 190 AND blue < 120       THEN 'yellow'
    WHEN red < 80  AND green < 80  AND blue < 80        THEN 'black'
    WHEN red > 190 AND green > 190 AND blue > 190       THEN 'white'
    WHEN ABS(red-green) < 28 AND ABS(green-blue) < 28   THEN 'grey'
    ELSE 'other'
  END AS color_comercial
FROM dominante
WHERE pos = 1;

The two filters at the start are the key to the whole section. Without them, 90 % of the catalogue's products would come out classified as "white", because the studio background occupies more pixels than the product. Discarding the extremes of brightness and requiring at least 5 % of the pixels leaves the colours that genuinely belong to the object.

Translating RGB into a commercial name is approximate and it is a business decision. The thresholds above are a reasonable starting point; "navy blue" and "electric blue" will both fall into "blue", and that is probably fine for a shop filter. If more subtlety were needed, the correct conversion goes through the HSV space, where hue is separated from saturation and brightness.

And the indispensable validation, which costs ten minutes: a GROUP BY p.color, c.color_comercial joining productos_color with productos over the product pages that do have a declared colour. Comparing against the colours that are declared on the page is the honest check: if "blue" on the page matches "blue" as detected in most cases, the method works. It is exactly the same idea as crossing sentiment with the stars in 05-04: using a human label that already exists to validate the automatic output.

  1. Application 3: moderating customer photos with SafeSearch

This is the riskiest of the four cases, and the one you are most grateful to have solved before you need it.

SafeSearch returns five categories with six levels each. AlpinaShop's moderation policy:

Category What it detects Blocking threshold Review threshold
adult Explicit sexual content LIKELY POSSIBLE
violence Violent or gory content LIKELY POSSIBLE
racy Suggestive content VERY_LIKELY LIKELY
medical Medical or surgical content — LIKELY
spoof Manipulated image or meme — LIKELY
flowchart TD
    A[Customer uploads photo in review] --> B[Store in quarantine]
    B --> C[SafeSearch]
    C -->|adult or violence LIKELY+| D[Automatic rejection<br/>+ notice to customer]
    C -->|category at POSSIBLE| E[Human review queue]
    C -->|all VERY_UNLIKELY/UNLIKELY| F[Check for a product object]
    F -->|product detected| G[Publish]
    F -->|no product| E
    E --> H[Human decision recorded]

The five design decisions in this architecture:

Quarantine by default. The photo is stored in a private bucket and is not publicly accessible until it passes the filter. Publishing first and moderating afterwards means the problematic content was visible.

Three outcomes, not two. Automatic rejection, human review and publication. The middle band is where most of the ambiguity lives, and forcing a binary decision guarantees errors in both directions.

medical and spoof never block automatically. In a mountain shop, a photo of chafing from a harness or of a blister is legitimate and useful content in a review about a pair of boots. Blocking it would be absurd. It goes to review.

It is also checked that a product appears. A perfectly harmless photo that shows no product at all — a landscape, a screenshot, an accidental photo — should not be published on a product page either. OBJECT_LOCALIZATION solves this and contributes more value than SafeSearch day to day.

Every decision is recorded. With the model version, the levels returned and, if there was a review, who decided and when. It is what you show if a customer complains that their photo was unfairly rejected.

LEVEL = {"UNKNOWN": 0, "VERY_UNLIKELY": 1, "UNLIKELY": 2,
         "POSSIBLE": 3, "LIKELY": 4, "VERY_LIKELY": 5}

def decide(safesearch):
    s = {k: LEVEL[getattr(safesearch, k).name]
         for k in ("adult", "violence", "racy", "medical", "spoof")}
    if s["adult"] >= 4 or s["violence"] >= 4 or s["racy"] >= 5:
        return "RECHAZADA", s
    if max(s.values()) >= 3:
        return "REVISION", s
    return "APTA", s

And the section's final warning, which is managerial and not technical: no automatic filter is infallible in either direction. There will be problematic content that gets through and harmless content that gets blocked. That is why, as well as the filter, you need a reporting channel for users and a quick takedown procedure. Automatic moderation reduces the volume of human work; it does not eliminate it, and it does not transfer responsibility to the machine.

  1. Application 4: auditing the catalogue's visual standard

AlpinaShop's standard says the main photo must have a uniform light background, the product centred filling between 50 % and 85 % of the frame, and no stray objects. Nobody has ever verified it.

With OBJECT_LOCALIZATION and IMAGE_PROPERTIES it checks itself:

CREATE OR REPLACE VIEW `alpinashop-datos.alpinashop_analitica.v_auditoria_catalogo` AS
WITH metricas AS (
  SELECT
    v.sku, v.uri,
    -- Area of the main object (normalised coordinates 0-1)
    (o.x_max - o.x_min) * (o.y_max - o.y_min)      AS area_producto,
    ABS(((o.x_min + o.x_max) / 2) - 0.5)           AS desvio_h,
    ABS(((o.y_min + o.y_max) / 2) - 0.5)           AS desvio_v,
    (SELECT MAX(c.pixel_fraction) FROM UNNEST(v.colores) c
      WHERE c.red > 215 AND c.green > 215 AND c.blue > 215) AS fondo_claro,
    ARRAY_LENGTH(v.objetos)                        AS n_objetos
  FROM `alpinashop-datos.alpinashop_analitica.imagenes_vision` v,
       UNNEST(v.objetos) o
  WHERE o.principal
)
SELECT
  sku, uri, ROUND(area_producto, 3) AS area, n_objetos,
  ARRAY_TO_STRING(ARRAY(
    SELECT problema FROM UNNEST([
      IF(area_producto < 0.50,               'producto_pequeno',  NULL),
      IF(area_producto > 0.85,               'producto_recortado',NULL),
      IF(desvio_h > 0.12 OR desvio_v > 0.12, 'descentrado',       NULL),
      IF(IFNULL(fondo_claro, 0) < 0.35,      'fondo_no_estandar', NULL),
      IF(n_objetos > 2,                      'objetos_ajenos',    NULL)
    ]) AS problema WHERE problema IS NOT NULL), ', ') AS problemas
FROM metricas;

The work list comes out of joining that view with productos and with the last twelve months' sales, filtering problemas != '' and ordering by units sold descending.

And here is the pattern that repeats throughout the module, now for the fourth time. The result is not a system that rejects photos automatically. It is a work list prioritised by business impact: fifty product pages, the best-selling ones, with the specific problem of each. Somebody looks at it, decides and asks the supplier for new photos where appropriate.

An automatic system that rejected photos would be more "intelligent" and much worse: it would get legitimate cases wrong, it would block product listings and it would end up switched off within a fortnight.

  1. Event-driven architecture with imagenes-subidas

Everything above is processing the history. For new images an automatic flow is needed, and AlpinaShop already has the central piece: the Pub/Sub topic imagenes-subidas (04-04).

flowchart LR
    A[Upload to Cloud Storage] -->|notification| B[Topic imagenes-subidas]
    B --> C[Push subscription]
    C --> D[Cloud Function 06-03]
    D --> E[Vision API]
    E --> D
    D --> F[BigQuery<br/>imagenes_vision]
    D -->|customer photo| G{SafeSearch}
    G -->|fit| H[Publish]
    G -->|doubtful| I[Review queue]
    G -->|rejected| J[Flag and notify]
    D -->|error| K[Dead letter topic]

The Cloud Storage to Pub/Sub notification is configured with gcloud storage buckets notifications create gs://alpinashop-catalogo --topic=imagenes-subidas --event-types=OBJECT_FINALIZE --object-prefix=productos/. OBJECT_FINALIZE fires when an object's upload finishes, and the prefix avoids generating events for files that are of no interest.

Four things you already know from module 4 that apply here unchanged:

Idempotency. The same event can be delivered more than once. The business key is the object's URI: if a row already exists for that URI with the same model version, it is not reprocessed. Without this, a retry costs money twice.

Dead letter topic. If processing fails repeatedly — a corrupt image, an unsupported format — the message ends up in the failed message queue instead of being retried forever. It is exactly the pattern from 04-04.

OIDC authentication on the push subscription, so that only Pub/Sub can invoke the endpoint.

Decoupling. The image upload does not wait for the analysis. The supplier uploads the photo and is done; the analysis happens afterwards.

The function's implementation arrives in 06-03, with Cloud Functions. Here the architecture and the guarantees it must meet are defined.

  1. OCR: text in images and documents

The API distinguishes two text recognition modes:

Mode Designed for Returns
TEXT_DETECTION Sparse text in photos The full text + each word with its position
DOCUMENT_TEXT_DETECTION Dense scanned documents Structure in pages, blocks, paragraphs and words

For AlpinaShop's product photos, TEXT_DETECTION has three concrete and useful applications:

  • Verifying the brand printed on the product against the brand declared on the product page. It detects photo assignment errors.
  • Reading sizes and capacities visible in the image ("40L", "XL", "8000 mm"), useful for checking that the photo corresponds to the right variant.
  • Detecting watermarks or promotional text from the supplier that should not appear in the catalogue ("50% OFF", the distributor's logo). It is a real problem and hard to find by hand.

It is requested with features=[{"type_": vision.Feature.Type.TEXT_DETECTION}] and read in response.full_text_annotation.text. The first element of text_annotations contains all the detected text; the following ones, each word with its bounding box. And language_hints=["en","es"] matters: in technical gear, English and Spanish get mixed constantly.

  1. Limits and biases: where it fails with mountain gear

As in 05-04, it is time for the honest part.

The labels are generic by design. The model was trained with a broad, general vocabulary. For AlpinaShop, that produces results like these:

Real product Typical labels returned Problem
Technical ice-fall axe "Hatchet", "Tool", "Metal" Wrong category
12-point crampons "Metal", "Footwear", "Tool" It does not recognise it
Membrane jacket "Outerwear", "Jacket", "Blue" Correct but superficial
60 L backpack "Backpack", "Luggage", "Bag" Correct
Climbing harness "Belt", "Strap", "Rope" Partial description
HMS carabiner "Metal", "Tool", "Ring" Unrecognisable

The pattern is clear: the more common the object is in everyday life, the better it identifies it. A backpack, perfect. An HMS carabiner, "a metal ring". The reason is simple and has no remedy inside the pre-trained API: the model saw millions of backpacks during its training and very few carabiners.

Biases inherited from training. The model reflects what was in its data: it is more accurate with brands and product styles common in over-represented markets, and it can associate gender labels with technical garments in debatable ways. For AlpinaShop the effect is minor — we are talking about objects, not people — but it deserves a check: if the colour filter or the labels behave systematically worse in some category, you need to know before a customer notices.

And the operational limit: the labels are not stable over time. The model updates and the labels can change between runs. That is why modelo_version and fecha_analisis are not optional.

  1. When to make the jump to AutoML Vision

The table in section 12 draws the boundary precisely.

Need Tool
Describing an image in general terms Vision API
Extracting dominant colours Vision API
Moderating content Vision API (SafeSearch)
Reading text Vision API (OCR)
Checking framing and composition Vision API (objects)
Classifying into the shop's own taxonomy AutoML Vision (05-02)
Telling a trekking backpack from a summit pack AutoML Vision
Detecting a specific manufacturing defect AutoML Vision (object detection)
Describing the image in rich natural language Gemini (05-06)

The criterion in one sentence: if your vocabulary is the world's, use the Vision API; if your vocabulary is your own, train with AutoML Vision.

And the two complement each other well. In 05-02, AlpinaShop trained AutoML Vision with its eight-category taxonomy precisely because the generic API could not tell its products apart. The API still contributes what AutoML does not give: colours, text, moderation, framing. One does not replace the other: different functions are requested over the same image.

A reasonable combined flow for each new catalogue photo: the Vision API for colours, OCR and framing; AutoML Vision for the shop's own category; and Gemini, when it arrives, to write the description.

  1. Document AI for invoices and delivery notes

There is one AlpinaShop problem that is not about product images: supplier invoices and delivery notes arrive as PDFs and somebody types them in by hand into the ERP.

DOCUMENT_TEXT_DETECTION extracts the text from a scanned PDF, but it returns plain text: you have to write regular expressions to find the invoice number, the amount or the line items, and those expressions break with every supplier and with every template change.

Document AI is a different product and it solves the problem in another way: it uses specialised processors per document type that return structured fields.

Aspect Vision OCR Document AI
Output Plain text with positions Named fields with values
Invoices Requires your own parsing Dedicated invoice processor
Tables You have to reconstruct them Native table extraction
Your own documents Not applicable Trainable custom processor
Cost per page Lower Higher

For AlpinaShop, the natural use case would be an invoice processor that extracts supplier, number, date, net amount, VAT, total and line items, and dumps the result into BigQuery to reconcile it with the purchase orders. With the usual cautions: an invoice contains personal and financial data, and automatic extraction requires human validation before anything is posted to the accounts.

It is a product that deserves its own project and it falls outside the scope of this lesson. What is relevant here is knowing it exists and not trying to do with OCR and regular expressions what a specific tool already handles.

  1. Faces, biometrics and the AI Act: the serious warning

This section is short and it is the most important in the lesson.

The Cloud Vision API offers face detection (FACE_DETECTION). It returns the position of each face in the image, facial landmarks and expression probabilities (joy, surprise, anger). It does not perform facial recognition: it does not identify who the person is nor compare them against a database.

That distinction is technically important and it is not legally sufficient. It is worth understanding why.

Face detection. Determining that there is a face in an image and where it is. In itself it identifies nobody.

Facial recognition. Determining who that person is. It involves processing biometric data, which the GDPR classifies as a special category of data (Article 9), with a far stricter regime: a general prohibition with exhaustively listed exceptions, among them explicit consent.

Emotion recognition. Inferring emotional state from the face. The EU AI Act contemplates specific restrictions for emotion recognition systems, especially in the workplace and educational spheres, and establishes transparency obligations in other contexts. The expression attributes that face detection returns fall into territory that has to be assessed with legal judgement, not technical judgement.

Where this shows up at AlpinaShop, without anybody looking for it:

  • Photos customers upload with their reviews. A person photographs themselves wearing the backpack. Their face is in the image.
  • Catalogue photos with models. Garments are photographed on people.
  • Scenery photos of climbing or trekking sent by the supplier.

In none of the three cases does AlpinaShop need to analyse faces. And there lies the practical recommendation:

Do not request FACE_DETECTION if you do not have a clear, documented and legally validated business need. Do not ask for it "just in case" or because it comes included in a copied example. Every function you request is a decision about which data gets processed.

If at some point it were needed — for instance, to blur faces before publishing a customer photo, which is a protective and reasonable use — then you would have to document the purpose and the legal basis, apply minimisation by processing only what is indispensable, not retain the facial data beyond the processing, and inform users in the privacy policy and in the upload form itself.

Express recommendation. Before enabling any facial analysis function, a compliance professional or the DPO must determine: whether the processing involves biometric data under Article 9 of the GDPR; the applicable legal basis and whether explicit consent is required; the system's classification under the AI Act and the associated obligations, with particular attention to the restrictions on emotion recognition; the need for an impact assessment (DPIA); and the information requirements towards data subjects. This lesson describes technical controls and does not constitute legal advice. All data is fictitious.

And for the rest of the functions — labels, colours, OCR, SafeSearch over product images — the precautions are the usual ones: the image is sent to the API, you have to verify in the official documentation the availability of regional endpoints if data residency in the EU is a requirement, and customer photos must have a defined retention period and a route to erasure at the data subject's request.

Common Mistakes and Tips

Analysing original/, web/ and thumb/. It is the same image three times. It multiplies the cost by three while contributing nothing.

Forgetting read permission on the bucket. The API reads the image from Cloud Storage with the service account. Without roles/storage.objectViewer, every call fails.

Confusing pixel_fraction with relevance. The background usually wins on pixel fraction. Without filtering the extremes of brightness, the whole catalogue comes out "white".

Publishing the generated alt without review. An incorrect alt text is worse than none: it deceives whoever cannot verify it.

Treating SafeSearch as a boolean. It is five categories with six levels. Forcing a binary decision produces false blocks and false passes.

Blocking automatically on medical. In a mountain shop, a photo of chafing is legitimate content in a review.

Requesting every function on every image. Each function is billed separately. SafeSearch over the supplier's catalogue photos is money thrown away.

Expecting precision with technical gear. An HMS carabiner will be "a metal ring". If you need your own taxonomy, that is AutoML Vision.

Requesting FACE_DETECTION with no need. See section 15.

Tip: store the mid of each label, not just the text. It is a stable identifier that does not depend on the language or on translation changes.

Tip: always validate against a human data point you already have. The colour declared on the product page, the product's category, the brand. It is the cheapest check there is and the one that catches the silent failures.

Exercises

Exercise 1

Marta calculates that analysing the catalogue images will cost around €500 and thinks that is excessive. Her plan was: every image in the bucket, with the labels, objects, properties, SafeSearch, OCR and logos functions. Cut the cost by at least 90 % without losing any of the lesson's four applications, and justify each cut.

Exercise 2

Design the automatic check that detects product pages whose photo does not match the product described. State which signals you would use, how you would combine them, and what you would do with the results.

Exercise 3

A customer complains that their photo was unfairly rejected in a review about a pair of boots. Describe what information you need to answer them, what must have been recorded at the moment of rejection, and what you would change in the system if it turns out the rejection was incorrect.

Solutions

Solution 1

The original plan: 20,000 images × 6 functions = 120,000 units.

The four cuts, in order of impact:

Cut 1 — Only the web/ folder. The original/, web/ and thumb/ folders are the same image at three resolutions. Analysing all three is literally paying three times for the same result. From 20,000 to around 7,000 images. Reduction: 65 %.

Cut 2 — Only the main photo of each SKU for labels, colour and OCR. To describe the product, extract its colour and read its brand, the main photo is enough: the others are the same product from another angle. From 7,000 to 2,400 images. Cumulative reduction: 88 %.

Cut 3 — Remove unnecessary functions.

Function Needed? Justification
Labels Yes The basis of the alt (application 1)
Image properties Yes Colour filter (application 2)
Objects Yes Framing audit (application 4)
SafeSearch Not here These are the supplier's photos, not customers'. It is reserved for the reviews flow
OCR Not in bulk Only over the sample where a watermark is suspected
Logos No The brand is already on the product page; it adds nothing

From 6 functions to 3. Cumulative reduction: 94 %.

Cut 4 — Do not reprocess. Store the results in imagenes_vision and analyse only what is new with the incremental query from section 5. The recurring cost becomes around 200 images a month, that is, practically nothing.

Result: 2,400 images × 3 functions = 7,200 units against 120,000. A reduction of 94 %, and the lesson's four applications are still covered:

Application Covered? With what
Alt text Yes Labels over the main photo
Colour filter Yes Image properties
Moderating customer photos Yes SafeSearch in the event-driven flow, not in the history
Catalogue audit Yes Objects over the main photo

What is lost and has to be said: the framing audit only covers the main photo, not all the ones in the carousel. It is an acceptable limitation — the main one is what 90 % of customers see — and it can always be extended later to the best-selling product pages.

The underlying lesson: no price has been negotiated and no functionality has been given up. You have simply stopped asking for what you did not need. It is the same reasoning as the SELECT of specific columns in BigQuery (04-01).

Solution 2

Available signals, and none is conclusive on its own:

Signal Source Strength
AutoML Vision category vs. product page category 05-02 High
Vision API main label vs. category Vision API Medium
Detected colour vs. declared colour Vision API Medium
Brand read by OCR vs. product page brand Vision API High when there is text
Detected object vs. product type Vision API Medium
Total absence of a recognisable object Vision API High (unusable photo)

The combination: a points system, not a single rule.

CREATE OR REPLACE VIEW `alpinashop-datos.alpinashop_analitica.v_fotos_sospechosas` AS
WITH senyales AS (
  SELECT
    v.sku, v.uri, p.nombre, p.categoria, p.marca,
    IF(a.categoria_predicha != p.categoria AND a.confianza > 0.85, 3, 0) AS p_automl,
    IF(t.marca_detectada IS NOT NULL
       AND UPPER(t.marca_detectada) != UPPER(p.marca), 3, 0)             AS p_marca,
    IF(ARRAY_LENGTH(v.objetos) = 0, 2, 0)                                AS p_sin_objeto,
    IF(c.color_comercial != LOWER(p.color) AND p.color IS NOT NULL, 1, 0) AS p_color,
    IF(NOT EXISTS(SELECT 1 FROM UNNEST(v.etiquetas) e WHERE e.score > 0.80
         AND LOWER(e.descripcion) LIKE CONCAT('%', LOWER(p.categoria), '%')), 1, 0)
                                                                         AS p_etiqueta
  FROM `alpinashop-datos.alpinashop_analitica.imagenes_vision` v
  JOIN `alpinashop-datos.alpinashop_analitica.productos` p USING (sku)
  LEFT JOIN `alpinashop-datos.alpinashop_analitica.imagenes_clasificadas` a USING (uri)
  LEFT JOIN `alpinashop-datos.alpinashop_analitica.productos_color` c USING (sku)
  LEFT JOIN `alpinashop-datos.alpinashop_analitica.imagenes_texto` t USING (uri)
)
SELECT *, p_automl + p_marca + p_sin_objeto + p_color + p_etiqueta AS puntuacion
FROM senyales
WHERE p_automl + p_marca + p_sin_objeto + p_color + p_etiqueta >= 3;

Why the weights differ. The AutoML discrepancy (3 points) and the OCR brand one (3 points) are strong signals: AutoML is trained with the shop's own taxonomy and text printed on the product is hard to misread. Colour (1 point) is the weakest signal because a product can have several colour variants sharing a page, and because the RGB-to-name translation is approximate. The absence of an object (2 points) indicates an unusable photo rather than a wrong one.

A single point does not trigger anything. A threshold of 3 requires either one strong signal or the coincidence of several weak ones. This drastically reduces the false positives, which is what kills this kind of system: nobody reviews a list with 800 false alarms.

What to do with the results, at three levels:

  1. Score ≥ 6 (several strong signals): priority review, and if the page is active and selling, notify whoever manages the catalogue.
  2. Score 3-5: a review queue ordered by the last twelve months' sales. The impact of an error on a product that sells 400 units is not the same as on one that sells 2.
  3. Never unpublish automatically. A false positive would leave a product with no photo in the shop, which is worse than the problem you are trying to solve.

And an operational recommendation: every human review generates a label. After a few months there is a set of confirmed and dismissed cases that lets you calibrate the weights with data instead of with intuition — or, if the volume justifies it, train your own classifier.

Solution 3

Information needed to answer the customer:

  1. The photo's identifier and the exact moment of upload.
  2. The complete SafeSearch response: the five levels returned, not just the conclusion.
  3. The model version and the analysis date.
  4. The rule that was applied: which category and which threshold triggered the rejection.
  5. Whether there was a human review: who, when and on what criterion.
  6. The image itself, so it can be looked at.

What must have been recorded at the moment of rejection:

CREATE TABLE IF NOT EXISTS `alpinashop-datos.alpinashop_analitica.moderacion_imagenes` (
  imagen_id      STRING NOT NULL, opinion_id STRING, uri_cuarentena STRING,
  subida_en      TIMESTAMP,       analizada_en TIMESTAMP, modelo_version STRING,
  nivel_adult    STRING, nivel_violence STRING, nivel_racy STRING,
  nivel_medical  STRING, nivel_spoof    STRING,
  decision       STRING,   -- APTA | REVISION | RECHAZADA
  regla_aplicada STRING,   -- e.g. 'adult >= LIKELY'
  revisor        STRING,   -- NULL if it was automatic
  revisado_en    TIMESTAMP, motivo_revisor STRING,
  reclamacion    BOOL,      resolucion     STRING
)
PARTITION BY DATE(subida_en);

Without this table, the answer to the customer is "the system rejected it and we do not know why", which is unacceptable and, if the decision affects their rights, legally problematic.

Probable diagnosis. In a mountain shop, the usual suspect is medical or violence firing at a photo of a blister, some chafing or a bruise caused by the boots. It is exactly the kind of photo a customer uploads to back up a criticism about uncomfortable footwear: perfectly legitimate, informative and relevant content. The second suspect is racy at a photo where skin appears — a bare foot showing the chafing — with no connotation at all.

What I would change in the system:

1. Never automatic rejection on medical. As the policy in section 8 already provides, this category must only send to review. If the rejection happened there, there is an implementation error with respect to the defined policy.

2. Raise the racy threshold and lower the review one. In this context, it is preferable for more photos than strictly necessary to go to human review than to automatically reject legitimate content. The cost of a false positive — an angry customer and a lost review — is greater than that of a few minutes of review.

3. Context in the decision. Combine SafeSearch with object detection: if footwear, a backpack or sports gear is detected in the image, that is a sign the photo is about the product. A photo with medical at POSSIBLE and a boot detected is almost certainly legitimate chafing.

4. A complaints channel with a deadline. The customer must be able to request a review of a rejection, and that request must reach a person within a defined period. The fact that this exercise exists means the customer found a way to complain, which is already a good thing.

5. An honest rejection message. "Your photo is pending review" is correct and accuses nobody. "Your photo has been rejected for inappropriate content" over a blister is offensive, and probably false as well.

6. A periodic review of the rejection rate. If the percentage of automatically rejected photos exceeds a small threshold, something is badly calibrated. It is the same discipline as the Dataplex quality rules in 04-07: measuring the process itself, not just the result.

And the management conclusion: this exercise illustrates why automatic moderation does not transfer responsibility to the machine. AlpinaShop answered for that decision, not the API provider. Automation reduces human work; the responsibility stays where it was.

Conclusion

Sixty gigabytes of images that only had a folder path are now queryable data in alpinashop_analitica, and they have solved the four problems the lesson opened with, without training anything and for a figure that fits into an afternoon's budget.

You know the Cloud Vision API's functions and, more importantly, which one is good for what: labels to describe, objects with bounding boxes to check composition, OCR to read, image properties for colour, SafeSearch to moderate, web detection to keep an eye on the use of your photos. And you know how to read the response field by field: the mid as a stable identifier against translatable text, the normalised coordinates that work the same at any resolution, and the difference between pixel_fraction and score in the colours, which is what separates a useful colour filter from one that classifies the whole catalogue as white.

You have processed the history with asyncBatchAnnotate, understanding that it really is asynchronous, that the results go to Cloud Storage and not into your memory, and that you have to split because of the request cap. And you have taken them into BigQuery partitioned and clustered, with modelo_version and fecha_analisis, because pre-trained APIs update on their own and without those columns you cannot tell a model change from a change in the images.

On cost, you have applied the reasoning that holds for the whole module: the invoice is not negotiated, you stop asking for what you do not need. Analysing only web/, only the main photo and only three functions cuts the spend by 94 % without losing any application.

And the four applications are built. The alt text starting from the product name and completing it with the labels, with automatic publication only where there is a match and human review for the rest, because an incorrect alt deceives whoever cannot verify it. The colour filter with the two brightness filters that stop the studio background from taking over the classification, validated against the colour that was already on the product page. The moderation of customer photos with quarantine by default, three outcomes instead of two, medical and spoof with no automatic block because a photo of chafing is legitimate content in a mountain shop, and a full record of every decision. And the visual standard audit, which produces the same thing AutoML and the review analysis produced: a work list prioritised by sales, not a system that decides on its own.

You have the event-driven architecture defined over the imagenes-subidas topic, with idempotency by URI, a dead letter topic, OIDC authentication and decoupling — the same guarantees as 04-04 — ready for the implementation to arrive with Cloud Functions in 06-03.

And you know the limits without embellishment: the API recognises a backpack perfectly and calls an HMS carabiner "a metal ring", because it saw millions of the former and very few of the latter. That is the boundary with AutoML Vision: if your vocabulary is the world's, the API; if it is your own, train. And they do not compete, they complement each other over the same image.

Finally, the warning that matters most. Face detection is not facial recognition, but that technical distinction is not legally sufficient: biometric data is a special category under Article 9 of the GDPR, and the AI Act specifically restricts emotion recognition. AlpinaShop has faces in its customer photos, in the catalogue photos with models and in the scenery ones, and it does not need to analyse a single one. The recommendation is literal: do not request FACE_DETECTION without a documented and legally validated need. Every function you request is a decision about which data you process.

Two questions from module 4 remain unanswered, and both are about producing rather than classifying. The 2,400 product pages that nobody has written, with descriptions that today are the supplier's spec sheet copied verbatim. And the shop's search, which only finds what matches literally: a customer who types "backpack for a three-day trek" gets nothing, because no product page contains those exact words.

In 05-06, generative AI on Vertex AI, the nature of the problem changes: from classifying to producing. You will work with the Gemini family from the SDK, understanding what temperature, top_p and system instructions really do, and how to force JSON output. You will generate the 2,400 product pages in batches with per-token cost control and compulsory human review before publishing. You will summarise each product's reviews into three useful sentences, comparing it with the sentiment from 05-04. You will understand what text embeddings are and you will build the catalogue's semantic search with Vector Search or with VECTOR_SEARCH in BigQuery. You will see the RAG pattern for the customer service assistant. And you will face what generative AI brings with it and the previous APIs did not have: hallucinations, safety filters, grounding, evaluation, transparency of generated content, intellectual property and the AI Act.

Google Cloud Platform (GCP) Course

Module 1: Introduction to Google Cloud Platform

Module 2: Core GCP Services

Module 3: Networking and Security

Module 4: Data and Analytics

Module 5: Machine Learning and AI

Module 6: DevOps and Monitoring

Module 7: Advanced GCP Topics

Module 8: Final Project

© Copyright 2026. All rights reserved