In 05-05 an explicit loose end was left. AlpinaShop designed an event-driven architecture to process automatically every product image uploaded to the alpinashop-catalogo bucket: a message to the imagenes-subidas topic, a call to the Vision API, the labels written to alpinashop_analitica and a thumbnail generated. The design was clear. The implementation was postponed because one piece was missing: something that runs code when an event happens, without a server sitting there waiting.

That piece is Cloud Functions, and it is what we are going to build here from start to finish.

But first it is worth understanding why the problem is not trivial. The "obvious" solution would be to add an endpoint to the Flask catalogue that receives the notification and processes the image. It is a bad idea for three reasons: processing an image takes seconds and would block a gunicorn worker that should be serving customer requests; the catalogue would scale according to the image load instead of the web traffic; and a processing failure would affect the shop. Reacting to events wants its own life cycle, and that is exactly the gap FaaS fills.

Contents

  1. What FaaS is and what the deployment unit is
  2. Cloud Functions today: the 2nd generation on top of Cloud Run and Eventarc
  3. HTTP functions: your first function and its deployment
  4. Authentication: why they should almost never be public
  5. Event-driven functions: CloudEvents and Eventarc
  6. The pending case: procesar-imagen-producto from start to finish
  7. Idempotency, retries and dead letter
  8. The real life cycle: cold starts and concurrency
  9. Configuration, secrets and identity
  10. Connecting to the VPC to reach Cloud SQL
  11. Limits and cost
  12. Local testing and deployment from Cloud Build
  13. When a function and when a service

  1. What FaaS is and what the deployment unit is

Run through the abstraction continuum from 02-07 in your head, now with one more box:

Model Deployment unit What you manage What you pay for
IaaS (Compute Engine) The virtual machine OS, patches, scaling, everything The machine while it is on
CaaS (GKE) The container The cluster and the manifests The nodes
PaaS (App Engine) The application The code and its configuration Instances
FaaS (Cloud Functions) The function Only the function's code Invocations and compute time

Function as a Service takes abstraction to its practical extreme: you write a function — a unit of code with a specific signature —, the platform packages it, deploys it, runs it when something triggers it and shuts it down when it is not needed.

The four properties that define the model:

  • It is triggered by an event. An HTTP request, a Pub/Sub message, a new file in a bucket, a modified document in Firestore. The function does not wait: it gets woken up.
  • It scales to zero. With no events, there are no instances and no compute is paid for. A radical difference from a MIG VM, which costs the same at four in the morning as it does in the middle of a campaign.
  • It scales up on its own. If a thousand events arrive at once, the platform starts instances. There is no autoscaler to configure.
  • It is ephemeral and stateless. An instance processes and may disappear. Nothing saved in memory or on local disk survives reliably.

That last property is the hardest to internalise and the one that produces the most design errors. State lives outside: in Cloud SQL, in Firestore, in Cloud Storage, in BigQuery. A function that accumulates something in a global variable and expects to find it on the next invocation works in testing — because it reuses the same instance — and fails in production, intermittently and inexplicably.

  1. Cloud Functions today: the 2nd generation on top of Cloud Run and Eventarc

This is the section you need to get right so as not to be misled by the old documentation still doing the rounds.

Cloud Functions was born in 2016 with a platform of its own. In 2022 the 2nd generation appeared, and the change is not cosmetic:

A 2nd generation Cloud Function is a Cloud Run service with a container built automatically by Google from your code, and its event triggers are Eventarc subscriptions.

Google builds the container for you using buildpacks, deploys it to Cloud Run and connects Eventarc so that events arrive as HTTP requests. Everything Cloud Run knows how to do, the function inherits.

flowchart TD
    A[Your code:<br/>main.py + requirements.txt] --> B[Cloud Build:<br/>buildpacks]
    B --> C[Image in<br/>Artifact Registry]
    C --> D[Cloud Run service<br/>managed by Cloud Functions]
    E[Pub/Sub, Cloud Storage,<br/>Firestore, Audit Logs] --> F[Eventarc]
    F -->|CloudEvent over HTTP POST| D
    G[Direct HTTP request] --> D

The practical differences, which are the ones that matter when designing:

Feature 1st generation 2nd generation
Platform Its own Cloud Run + Eventarc
Concurrency per instance 1 request Up to 1,000, configurable
Maximum time (HTTP) 9 minutes 60 minutes
Maximum time (events) 9 minutes 9 minutes (via Eventarc)
Maximum memory 8 GB 32 GB
CPU Tied to the memory Configurable, up to 8 vCPU
Traffic splitting between revisions No Yes, like Cloud Run
Event sources A handful of native ones More than 130 via Eventarc
Minimum instances Limited Yes
Base cost Lower on minimal workloads Slightly higher, more capacity

Concurrency is the most important difference. In the 1st generation, each instance served one request: 100 simultaneous requests meant 100 instances, with 100 cold starts and 100 database connections. In the 2nd, one instance with concurrency 80 serves 80 requests at a time, which drastically reduces cold starts, cost and pressure on Cloud SQL — a very real problem: the number of Cloud SQL connections is limited, and a badly sized 1st generation function exhausts the pool with astonishing ease.

The rule in 2026: always use the 2nd generation. --gen2 is mandatory in this lesson's examples. The 1st generation only shows up in old functions nobody has migrated.

And the natural question: if a 2nd generation function is a Cloud Run, why not use Cloud Run directly? It is an excellent question and it has an answer, but we leave it to section 13, when you have the context to weigh it up.

  1. HTTP functions: your first function and its deployment

The Functions Framework is the library that turns an ordinary Python function into an HTTP service. It is installed like any other dependency and uses decorators.

The minimum structure of a function is two files:

funcion-salud/
├── main.py
└── requirements.txt
# main.py
import functions_framework
from flask import jsonify

@functions_framework.http
def comprobar_stock(request):
    """Returns the stock for a SKU. HTTP entry point."""
    # request is a Flask Request object: the same API you already know
    sku = request.args.get("sku")
    if not sku:
        return jsonify({"error": "the sku parameter is missing"}), 400

    units = look_up_stock(sku)               # implementation omitted
    if units is None:
        return jsonify({"error": "unknown sku"}), 404

    return jsonify({"sku": sku, "units": units, "available": units > 0}), 200
# requirements.txt
functions-framework==3.*
google-cloud-firestore==2.*

Three details worth noting:

  • The @functions_framework.http decorator marks the entry point. The name of the Python function (comprobar_stock) is what you pass to --entry-point.
  • request is a Flask object. If you wrote the catalogue with Flask, the API is familiar: request.args, request.get_json(), request.headers.
  • The return value follows Flask's conventions: a string, a (body, code) tuple or a complete response.

The deployment:

gcloud functions deploy comprobar-stock \
  --gen2 \
  --region=europe-west1 \
  --runtime=python312 \
  --source=. \
  --entry-point=comprobar_stock \
  --trigger-http \
  --no-allow-unauthenticated \
  --service-account=sa-funcion-stock@alpinashop-prod.iam.gserviceaccount.com \
  --memory=256Mi \
  --timeout=30s \
  --max-instances=20 \
  --project=alpinashop-prod

Each option, and why it is there:

Option What it does Why this value
--gen2 Uses the 2nd generation Always, because of section 2
--region Where it is deployed europe-west1, next to everything else
--runtime Language version Pinned explicitly, not implicitly
--source Where the code comes from . locally; also gs:// or a repository
--entry-point Name of the Python function It must match exactly
--trigger-http Triggered over HTTP As opposed to event triggers
--no-allow-unauthenticated Requires IAM authentication Section 4
--service-account Execution identity Never the default account
--memory Memory per instance It also determines the CPU allocated
--timeout Maximum time per invocation Short: if it takes longer, something is wrong
--max-instances Scaling ceiling It protects what is behind it

--max-instances deserves an explanation, because its absence causes real incidents. With no ceiling, a traffic spike — or an infinite loop, or an attack — can start hundreds of instances that open hundreds of connections to alpinashop-pedidos and bring the database down. Infinite scaling is not a virtue if what is behind it does not scale the same way. Setting a limit turns a total outage into a partial degradation, which is infinitely preferable.

  1. Authentication: why they should almost never be public

It is tempting to deploy with --allow-unauthenticated because "it is easier to test". That is exactly how data gets leaked.

A public function has a URL that is guessable from a known pattern, it is exposed to the whole internet, anybody can invoke it as many times as they like — and you pay for every invocation — and it usually has IAM permissions over databases and buckets in your project. It is a privileged endpoint with no door.

With --no-allow-unauthenticated, invoking the function requires the roles/run.invoker role — Cloud Run's, consistent with section 2:

# Only the catalogue website can call this function
gcloud functions add-invoker-policy-binding comprobar-stock \
  --region=europe-west1 \
  --member="serviceAccount:[email protected]" \
  --project=alpinashop-prod

And from the catalogue, the call is authenticated with an OIDC token:

import google.auth.transport.requests
import google.oauth2.id_token

STOCK_URL = "https://europe-west1-alpinashop-prod.cloudfunctions.net/comprobar-stock"

def remote_stock_lookup(sku: str) -> dict:
    # The "audience" must be exactly the URL of the function
    auth_request = google.auth.transport.requests.Request()
    token = google.oauth2.id_token.fetch_id_token(auth_request, STOCK_URL)

    response = requests.get(
        STOCK_URL,
        params={"sku": sku},
        headers={"Authorization": f"Bearer {token}"},
        timeout=5,
    )
    response.raise_for_status()
    return response.json()

The library obtains the token from the environment's identity — the service account attached to the GKE pod — with no key at all. It is the same mechanism as the push subscriptions with OIDC from 04-04.

The three cases in which a function can be public, and what is needed in each:

Case What else you need
Third-party webhook (payment gateway) Verify the provider's signature in the request body
A genuinely public endpoint (web form) Cloud Armor in front (03-05), request rate limiting and strict validation
Development and testing That it be in alpinashop-dev and touch nothing real

The webhook case deserves a note: a public function that receives payment notifications must validate the HMAC signature the provider includes in the header. Without that validation, anybody can send a POST saying "order 1234 is paid". It is a failure that shows up with depressing frequency.

  1. Event-driven functions: CloudEvents and Eventarc

An event-driven function does not receive an HTTP request: it receives a CloudEvent, an industry-standard format — not a Google invention — with common metadata (id, source, type, time, subject) and a payload specific to the event type.

In Python it is declared with @functions_framework.cloud_event:

import functions_framework

@functions_framework.cloud_event
def process(event):
    print("id:",     event["id"])       # unique identifier of the event
    print("type:",   event["type"])     # google.cloud.pubsub.topic.v1.messagePublished
    print("source:", event["source"])   # the resource that generated it
    data = event.data                   # the payload, depending on the type

Eventarc is GCP's universal event router. It receives events from more than 130 sources, normalises them to CloudEvents and delivers them to the destination. The most useful triggers:

Source Event type It fires when Typical use at AlpinaShop
Pub/Sub ...pubsub.topic.v1.messagePublished Something is published to a topic imagenes-subidas (section 6)
Cloud Storage ...storage.object.v1.finalized An object is created or overwritten Direct alternative to the topic
Cloud Storage ...storage.object.v1.deleted An object is deleted Cleaning up orphaned references
Firestore ...firestore.document.v1.written A document is written Reacting to a cart
Firestore ...document.v1.created / .deleted Creation or removal Functional auditing
Audit logs google.cloud.audit.log.v1.written Any logged operation Alerting on firewall changes
Cloud Scheduler Via Pub/Sub At the scheduled time Light periodic tasks
BigQuery Via audit logs A job finishes Chaining analytical processes

The audit log trigger is the least known and one of the most powerful: it lets you react to any GCP API operation. For example, running a function every time somebody modifies a firewall rule in alpinashop-prod:

gcloud functions deploy alertar-cambio-firewall \
  --gen2 --region=europe-west1 --runtime=python312 \
  --entry-point=alertar --source=. \
  --trigger-event-filters="type=google.cloud.audit.log.v1.written" \
  --trigger-event-filters="serviceName=compute.googleapis.com" \
  --trigger-event-filters="methodName=v1.compute.firewalls.insert" \
  --service-account=sa-alertas-seguridad@alpinashop-prod.iam.gserviceaccount.com \
  --project=alpinashop-prod

An important note about Cloud Storage: you can trigger directly from the bucket (storage.object.v1.finalized) or by publishing to a Pub/Sub topic and triggering from there. AlpinaShop chose the second option in 05-05, and the reason is solid: with Pub/Sub in the middle, several consumers can react to the same event. Today it only processes images; tomorrow, a second subscriber can update the search index without touching anything that already exists. With the direct trigger, every new consumer forces you to reconfigure the bucket.

  1. The pending case: procesar-imagen-producto from start to finish

Here the loose end from 05-05 is tied off. Let us recall the flow designed back then:

flowchart LR
    A[Dani uploads a photo to<br/>gs://alpinashop-catalogo] -->|notification| B[Pub/Sub topic<br/>imagenes-subidas]
    B -->|CloudEvent via Eventarc| C[Function<br/>procesar-imagen-producto]
    C --> D[Vision API:<br/>labels, colours, SafeSearch]
    C --> E[400px thumbnail<br/>to Cloud Storage]
    D --> F[BigQuery<br/>imagenes_vision]
    C -.->|error after 5 attempts| G[imagenes-subidas-dlq]

Step 1: notify the topic from the bucket. A single command, which is probably already done since 05-05:

gcloud storage buckets notifications create gs://alpinashop-catalogo \
  --topic=imagenes-subidas \
  --event-types=OBJECT_FINALIZE \
  --object-prefix=productos/ \
  --project=alpinashop-prod

The --object-prefix=productos/ avoids processing objects that are not product photos — logos, temporary files — and it is the cheapest way to filter: the event is not even generated.

Step 2: the function. It is commented block by block because each one solves a specific problem:

# main.py
import base64
import json
import os
from datetime import datetime, timezone

import functions_framework
from google.cloud import bigquery, storage, vision
from PIL import Image
import io

# The clients are created ONCE, outside the function.
# They are reused for as long as the instance lives: it saves hundreds of ms per invocation.
vision_client = vision.ImageAnnotatorClient()
storage_client = storage.Client()
bq_client = bigquery.Client()

PROYECTO_DATOS = os.environ["PROYECTO_DATOS"]        # alpinashop-datos
TABLE = f"{PROYECTO_DATOS}.alpinashop_analitica.imagenes_vision"
BUCKET_MINIATURAS = os.environ["BUCKET_MINIATURAS"]  # alpinashop-catalogo
ANCHO_MINIATURA = 400


@functions_framework.cloud_event
def procesar_imagen(event):
    """Processes a new image: Vision API + thumbnail + BigQuery."""

    # 1. Decode the Pub/Sub message. The payload arrives base64-encoded.
    message = base64.b64decode(event.data["message"]["data"]).decode("utf-8")
    notice = json.loads(message)
    bucket_name = notice["bucket"]
    path = notice["name"]
    generation = notice["generation"]   # identifies the exact VERSION of the object

    # 2. Guard: ignore what must not be processed.
    #    Without this, the thumbnail we write below would fire another event
    #    and we would have an infinite loop that also costs money.
    if path.startswith("productos/miniaturas/"):
        print(f"Thumbnail ignored: {path}")
        return
    if not path.lower().endswith((".jpg", ".jpeg", ".png", ".webp")):
        print(f"Unsupported format ignored: {path}")
        return

    # 3. Idempotency: if this generation has already been processed, do not repeat it.
    image_id = f"gs://{bucket_name}/{path}#{generation}"
    if already_processed(image_id):
        print(f"Already processed, skipping: {image_id}")
        return

    # 4. Vision API over the object in Cloud Storage, without downloading it.
    image = vision.Image(source=vision.ImageSource(
        gcs_image_uri=f"gs://{bucket_name}/{path}"))
    response = vision_client.annotate_image({
        "image": image,
        "features": [
            {"type_": vision.Feature.Type.LABEL_DETECTION, "max_results": 10},
            {"type_": vision.Feature.Type.IMAGE_PROPERTIES},
            {"type_": vision.Feature.Type.SAFE_SEARCH_DETECTION},
        ],
    })
    if response.error.message:
        # API error: raise an exception so that Pub/Sub retries
        raise RuntimeError(f"Vision API: {response.error.message}")

    # 5. Generate the thumbnail and upload it to the prefix that step 2 ignores
    blob = storage_client.bucket(bucket_name).blob(path)
    original = Image.open(io.BytesIO(blob.download_as_bytes()))
    original.thumbnail((ANCHO_MINIATURA, ANCHO_MINIATURA))
    output = io.BytesIO()
    original.convert("RGB").save(output, format="JPEG", quality=82)

    thumb_path = f"productos/miniaturas/{os.path.basename(path)}"
    storage_client.bucket(BUCKET_MINIATURAS).blob(thumb_path).upload_from_string(
        output.getvalue(), content_type="image/jpeg")

    # 6. Write to BigQuery, with the identifier that guarantees idempotency
    row = {
        "id_imagen": image_id,
        "ruta_gcs": f"gs://{bucket_name}/{path}",
        "ruta_miniatura": f"gs://{BUCKET_MINIATURAS}/{thumb_path}",
        "etiquetas": [
            {"descripcion": l.description, "puntuacion": round(l.score, 4)}
            for l in response.label_annotations
        ],
        "color_dominante": dominant_colour(response),
        "safesearch_adulto": response.safe_search_annotation.adult.name,
        "procesada_en": datetime.now(timezone.utc).isoformat(),
    }
    errors = bq_client.insert_rows_json(TABLE, [row], row_ids=[image_id])
    if errors:
        raise RuntimeError(f"BigQuery: {errors}")

    print(f"Processed successfully: {image_id}")

The three details that make this work in production, and that separate a tutorial example from real code:

The clients outside the function. Creating an ImageAnnotatorClient means resolving credentials and establishing connections: hundreds of milliseconds. By creating it at module level, it is created once per instance and reused across every invocation that instance serves. In a function with concurrency 20 and a thousand events, the difference is enormous.

The guard against the infinite loop (step 2). It is the classic mistake with functions triggered by Cloud Storage and it deserves a slow explanation: the function writes the thumbnail into the same bucket, which generates an OBJECT_FINALIZE event, which fires the function, which generates another thumbnail, which fires the function… Each iteration costs invocations, Vision API calls — which are billed — and rows in BigQuery. A loop like that discovered on a Monday morning may have burned through an entire budget over the weekend. The three possible defences are the prefix in the notification filter, the guard in the code and writing to a different bucket; use at least two.

The row_ids in insert_rows_json. BigQuery deduplicates by that identifier within a time window. It is a second net beneath the idempotency check in step 3.

Step 3: deploy it with its identity and its permissions.

# Its own service account with minimum permissions
gcloud iam service-accounts create sa-procesar-imagen \
  --display-name="Function procesar-imagen-producto" --project=alpinashop-prod

[email protected]

gcloud projects add-iam-policy-binding alpinashop-prod \
  --member="serviceAccount:${SA}" --role=roles/storage.objectAdmin
gcloud projects add-iam-policy-binding alpinashop-datos \
  --member="serviceAccount:${SA}" --role=roles/bigquery.dataEditor
gcloud projects add-iam-policy-binding alpinashop-prod \
  --member="serviceAccount:${SA}" --role=roles/eventarc.eventReceiver

gcloud functions deploy procesar-imagen-producto \
  --gen2 --region=europe-west1 --runtime=python312 \
  --source=. --entry-point=procesar_imagen \
  --trigger-topic=imagenes-subidas \
  --service-account="${SA}" \
  --set-env-vars=PROYECTO_DATOS=alpinashop-datos,BUCKET_MINIATURAS=alpinashop-catalogo \
  --memory=1Gi \
  --timeout=120s \
  --max-instances=50 \
  --retry \
  --project=alpinashop-prod

--memory=1Gi is not a whim: opening a 12-megapixel product photo with Pillow consumes a fair amount of memory, and in the 2nd generation the allocated CPU grows with the memory, so it also runs faster. --retry enables retries, which is the subject of the next section.

  1. Idempotency, retries and dead letter

This is the part that separates a function that works from a function you can trust. And it starts from a fact already established in 04-04:

Pub/Sub guarantees at least once delivery. Your function will receive duplicate messages. It is not a remote possibility: it is a statistical certainty.

Duplicates arrive for perfectly normal reasons: the function processed correctly but took longer than the acknowledgement deadline; there was a network failure while acknowledging; the instance restarted after processing and before acknowledging; or the event was redelivered because a previous invocation raised an exception.

Idempotent means that processing the same event N times produces the same result as processing it once. Let us analyse our function's three actions:

Action Idempotent by nature? What would happen with a duplicate
Calling the Vision API Yes in result, no in cost You pay twice for the same thing
Writing the thumbnail Yes: same path, it is overwritten No harm, just wasted compute
Inserting into BigQuery No Duplicate row: the statistics lie

The third is the dangerous one, and that is why the design includes three layers of defence:

Layer 1, the stable identifier. gs://bucket/path#generation identifies a specific version of an object. Notice that the path alone would not be enough: if Dani uploads a corrected photo with the same name, it is a different object that does need reprocessing, and the generation distinguishes it.

Layer 2, the up-front check with a lightweight control table:

def already_processed(image_id: str) -> bool:
    """Checks in a control table whether this id has already been processed."""
    query = f"""
        SELECT 1 FROM `{PROYECTO_DATOS}.alpinashop_analitica.imagenes_procesadas`
        WHERE id_imagen = @id
          AND procesada_en > TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 7 DAY)
        LIMIT 1
    """
    job = bq_client.query(query, job_config=bigquery.QueryJobConfig(
        query_parameters=[bigquery.ScalarQueryParameter("id", "STRING", image_id)]))
    return next(job.result(), None) is not None

The 7-day window is a deliberate compromise: Pub/Sub duplicates arrive within seconds or minutes, not days, and limiting the query lets you take advantage of the table's partitioning instead of scanning the full history on every invocation. It is the same cost logic as in 04-01.

Layer 3, BigQuery's row_ids, which deduplicates automatically even if the previous two fail because of a race condition.

Retries. With --retry, if the function raises an exception, Pub/Sub receives no acknowledgement and redelivers the message. And here there is a design decision that is constantly got wrong:

Type of error Example Retry? What to do
Transient Vision API returning 503, network timeout Yes Raise an exception
Permanent Corrupt image, unsupported format No Log it and return normally
Configuration An IAM permission is missing It does not help Log as a serious error and alert

A permanent error that gets retried is a function failing forever on the same message, consuming quota and filling the logs. The rule: only raise an exception if trying again has some chance of working.

    try:
        original = Image.open(io.BytesIO(blob.download_as_bytes()))
    except UnidentifiedImageError:
        # PERMANENT error: retrying a thousand times will not fix a corrupt file
        print(f"PERMANENT ERROR: unreadable image {path}")
        record_permanent_failure(image_id, "unreadable_image")
        return          # normal return → Pub/Sub acknowledges and does not retry

Dead letter. Even with the distinction above, some message will fail indefinitely. The dead letter topic stops it blocking the queue:

gcloud pubsub topics create imagenes-subidas-dlq --project=alpinashop-prod

gcloud pubsub subscriptions update eventarc-europe-west1-procesar-imagen-sub \
  --dead-letter-topic=projects/alpinashop-prod/topics/imagenes-subidas-dlq \
  --max-delivery-attempts=5 \
  --project=alpinashop-prod

After five attempts, the message goes to the DLQ instead of being retried forever. And a DLQ nobody looks at is worse than having no DLQ, because it creates a false sense of control: you have to put an alert on its number of unacknowledged messages — exactly the kind of alert configured in 06-04.

  1. The real life cycle: cold starts and concurrency

A function instance goes through three phases, and understanding which one is expensive explains almost all the observed behaviour:

flowchart LR
    A[No instances<br/>cost 0] -->|an event arrives| B[COLD START<br/>create instance +<br/>load code +<br/>run the module]
    B --> C[Invocation<br/>WARM START]
    C -->|more events| C
    C -->|no events<br/>for a few minutes| A

The cold start is the time from the event arriving to your code starting to run: provisioning the instance, loading the runtime, importing the dependencies and running the module-level code. It ranges from a few hundred milliseconds to several seconds, and it depends above all on how heavy your imports are.

The five levers for mitigating it, ordered by real effectiveness:

Lever How Effect Cost
--min-instances Keep N instances always alive Removes cold starts for the baseline traffic You pay for the idle instance
Fewer dependencies Import only what is needed Reduces start-up a lot None
Deferred imports import inside the function Only if it is not always used Complicates the code
High concurrency --concurrency=80 Fewer instances, fewer starts Requires thread-safe code
More CPU --cpu=2 Faster start-up More expensive per second
# Critical function on the customer path: no cold starts
gcloud functions deploy comprobar-stock --gen2 --region=europe-west1 \
  --min-instances=2 --max-instances=50 --concurrency=80 \
  --project=alpinashop-prod

When cold starts matter and when they do not, which is the decision you really have to make:

Situation Does it matter? Recommendation
A function on the path of a customer request A lot --min-instances ≥ 1
Payment gateway webhook Yes: there are timeouts --min-instances=1
procesar-imagen-producto No --min-instances=0: two extra seconds bother nobody
Scheduled nightly task No 0

--min-instances costs money even with no traffic, and that is the price of partially giving up "scale to zero". Setting it by default on every function cancels out one of the model's economic advantages. Set it where a customer perceives the latency.

About concurrency, an important warning: with --concurrency=80, eighty requests run simultaneously in the same Python process. Any mutable global variable becomes shared:

# DANGEROUS with concurrency > 1
results = []                 # shared between simultaneous invocations!

@functions_framework.http
def process(request):
    results.append(request.args["id"])       # race condition
    return str(len(results))                 # returns anything at all

# CORRECT: global objects only for reusable, immutable clients
bq_client = bigquery.Client()   # safe: it is designed for concurrent use

@functions_framework.http
def process(request):
    local_results = []           # state inside the invocation
    ...

The rule: at module level, only clients and constants. All mutable state, inside the function.

  1. Configuration, secrets and identity

The three things every production function needs to have properly sorted.

Environment variables for non-sensitive configuration:

gcloud functions deploy procesar-imagen-producto --gen2 \
  --set-env-vars=PROYECTO_DATOS=alpinashop-datos,ENTORNO=prod,ANCHO_MINIATURA=400 \
  --region=europe-west1 --project=alpinashop-prod

Secrets from Secret Manager, never as an environment variable carrying the value:

gcloud functions deploy procesar-pago --gen2 \
  --set-secrets='API_KEY_PASARELA=api-key-pasarela-pago:latest' \
  --region=europe-west1 --project=alpinashop-prod

With that syntax, the function reads os.environ["API_KEY_PASARELA"] as normal, but the value is not in the deployment configuration: it is injected at run time from Secret Manager (03-06). Practical consequences: the value does not show up in the console, it is not left in the deployment history, it can be rotated without redeploying if you use :latest, and access is recorded in the audit logs. They can also be mounted as a file with --set-secrets='/etc/claves/api=secreto:latest', which is preferable for large values such as certificates.

And the detail that gets forgotten: the function's service account needs roles/secretmanager.secretAccessor on that specific secret, granted in the secret's policy.

Identity. By default, a function uses the project's Compute Engine service account, which usually has the Editor role — the same anti-pattern as in 06-01. Always --service-account with an account of its own per function. A function that only writes to BigQuery must not be able to delete buckets.

  1. Connecting to the VPC to reach Cloud SQL

By default, a function lives outside your VPC: it can go out to the internet but it cannot reach resources with a private IP. If alpinashop-pedidos only has a private IP — as it should, according to 03-01 — a function cannot get there.

The bridge is a Serverless VPC Access connector:

gcloud compute networks vpc-access connectors create conector-alpinashop \
  --region=europe-west1 \
  --network=alpinashop-vpc \
  --range=10.8.0.0/28 \
  --min-instances=2 --max-instances=4 \
  --machine-type=e2-micro \
  --project=alpinashop-prod

gcloud functions deploy sincronizar-stock --gen2 \
  --vpc-connector=projects/alpinashop-prod/locations/europe-west1/connectors/conector-alpinashop \
  --egress-settings=private-ranges-only \
  --region=europe-west1 --project=alpinashop-prod

Details that matter:

  • The /28 range must be free and must not overlap with sn-web-euw1 or sn-datos-euw1. It is a strict requirement and a common source of errors.
  • The connector is billed per instance and hour, whether it is in use or not. With --min-instances=2 you pay for two small machines permanently. It partially breaks "scale to zero", so it is shared between all the functions that need it.
  • --egress-settings=private-ranges-only sends only traffic to private ranges through the VPC; the rest goes straight out to the internet. With all-traffic, everything goes through the VPC and out via the Cloud NAT alpinashop-nat-euw1 from 03-01, which gives a fixed egress IP — useful if a third party demands IP allowlisting.
Need Solution Added cost
Cloud SQL with a public IP Cloud SQL connector over TLS None
Cloud SQL with a private IP VPC connector Connector instances
Fixed egress IP Connector + all-traffic + Cloud NAT Connector + NAT
Google APIs only Nothing: it already works None

Tip: do not add a connector "just in case". If the function only talks to Google APIs — Vision, BigQuery, Storage, like procesar-imagen-producto — it does not need one, and adding it is cost and complexity for nothing.

  1. Limits and cost

The 2nd generation limits worth keeping in mind — always check the current values in the official documentation:

Limit Approximate value What to do if it falls short
Maximum time (HTTP) 60 minutes Long work → Cloud Run Jobs or Dataflow
Maximum time (events) 9 minutes Split the work, chain it via Pub/Sub
Memory Up to 32 GB Heavy processing → Cloud Run or Batch
CPU Up to 8 vCPU Same
Size of the deployed code Tens of MB compressed Models and data in Cloud Storage
Size of the HTTP request 32 MB Upload straight to Cloud Storage + event
Simultaneous instances Thousands (quota can be raised) Request a quota increase
Environment variables A few KB in total Configuration in Firestore or Secret Manager

Cost has four components, and there is a generous monthly free tier:

Component You pay for Comment
Invocations Each call Cents per million
GB-second Memory × time The main lever
GHz-second CPU × time Tied to the memory
Network egress GB out to the internet Zero between services in the same region

An indicative calculation for procesar-imagen-producto. Let us assume 20,000 images a month, 1 GiB of memory, 3 seconds per image:

  • Invocations: 20,000 → covered many times over by the free tier.
  • GB-second: 20,000 × 3 s × 1 GiB = 60,000 GB-s, mostly within the free tier.
  • GHz-second: proportional, same order of magnitude.
  • Real compute cost: practically zero.

And now the figure that really matters: the Vision API for those 20,000 images, with three features each, costs an order of magnitude more than the entire execution of the function. The economic lesson is that in event-driven architectures compute is almost never the main cost: the costs are the APIs you call and the data you move. Optimising the function's memory while making redundant calls to a paid API is optimising what does not matter.

That said, three ways for the bill to genuinely explode:

  1. Infinite loops (section 6). The most expensive by a wide margin.
  2. --min-instances on functions that do not need it. It is a fixed 24×7 cost and it cancels out the model's advantage.
  3. A long timeout with hung errors. A function with --timeout=540s that sits waiting for a downed service pays nine minutes per failed invocation. Short timeouts, matched to reality.

  1. Local testing and deployment from Cloud Build

Locally, the Functions Framework brings up a server:

pip install functions-framework
functions-framework --target=procesar_imagen --signature-type=cloudevent --port=8080

And a Pub/Sub event is simulated with the real structure, base64 included:

DATA=$(printf '{"bucket":"alpinashop-catalogo","name":"productos/piolet-01.jpg","generation":"1712345678"}' | base64 -w0)

curl -X POST http://localhost:8080 \
  -H "Content-Type: application/json" \
  -H "ce-id: 1234" \
  -H "ce-source: //pubsub.googleapis.com/projects/alpinashop-prod/topics/imagenes-subidas" \
  -H "ce-type: google.cloud.pubsub.topic.v1.messagePublished" \
  -H "ce-specversion: 1.0" \
  -d "{\"message\":{\"data\":\"${DATA}\"}}"

Better still, unit tests that do not need anything brought up:

# test_main.py
import base64, json
from unittest.mock import patch, MagicMock
from cloudevents.http import CloudEvent
import main

def make_event(bucket, name, generation="1"):
    data = base64.b64encode(json.dumps(
        {"bucket": bucket, "name": name, "generation": generation}).encode())
    return CloudEvent(
        {"type": "google.cloud.pubsub.topic.v1.messagePublished",
         "source": "//pubsub.googleapis.com/", "id": "1"},
        {"message": {"data": data.decode()}})

def test_ignores_thumbnails():
    """The infinite loop guard is the MOST important test of all."""
    with patch.object(main, "vision_client") as vision_mock:
        main.procesar_imagen(make_event("alpinashop-catalogo",
                                        "productos/miniaturas/x.jpg"))
        vision_mock.annotate_image.assert_not_called()

def test_skips_if_already_processed():
    with patch.object(main, "already_processed", return_value=True), \
         patch.object(main, "vision_client") as vision_mock:
        main.procesar_imagen(make_event("alpinashop-catalogo", "productos/a.jpg"))
        vision_mock.annotate_image.assert_not_called()

The first test deserves a comment: it verifies that the Vision API is not called for a thumbnail. It is the test that protects against the most expensive possible failure of this architecture, and it costs six lines.

Deployment from Cloud Build, fitting in with 06-01 and 06-02:

# cloudbuild.yaml of the functions repository
steps:
  - name: 'python:3.12-slim'
    id: 'pruebas'
    entrypoint: 'bash'
    args:
      - '-c'
      - |
        pip install -r requirements.txt -r requirements-dev.txt -t /workspace/lib
        PYTHONPATH=/workspace/lib python -m pytest -q

  - name: 'gcr.io/google.com/cloudsdktool/cloud-sdk:slim'
    id: 'desplegar'
    waitFor: ['pruebas']
    args:
      - 'gcloud'
      - 'functions'
      - 'deploy'
      - 'procesar-imagen-producto'
      - '--gen2'
      - '--region=europe-west1'
      - '--runtime=python312'
      - '--source=.'
      - '--entry-point=procesar_imagen'
      - '--trigger-topic=imagenes-subidas'
      - '--service-account=sa-procesar-imagen@alpinashop-prod.iam.gserviceaccount.com'
      - '--set-env-vars=PROYECTO_DATOS=alpinashop-datos,BUCKET_MINIATURAS=alpinashop-catalogo'
      - '--memory=1Gi'
      - '--max-instances=50'
      - '--retry'
      - '--project=alpinashop-prod'

options:
  logging: CLOUD_LOGGING_ONLY

The Cloud Build service account needs roles/cloudfunctions.developer and roles/iam.serviceAccountUser over sa-procesar-imagen — this second permission is always forgotten and produces a permissions error that never mentions the word "function".

  1. When a function and when a service

We come back to the question left open in section 2: if a 2nd generation function is a Cloud Run, why choose a function?

Criterion Cloud Functions (2nd gen) Cloud Run
What you deploy Source code A container
Who builds the image Google, with buildpacks You, with your Dockerfile
Control of the environment The runtime Google offers Total
HTTP routes One function, one entry point Any complete web application
Event triggers Built into the deployment Eventarc configured separately
Learning curve Very low Medium
Migrating elsewhere Depends on the platform A container runs anywhere
System dependencies Those of the runtime Whatever you install

Choose Cloud Functions when:

  • The work is one single thing triggered by an event: procesar-imagen-producto is the perfect example.
  • You do not need system dependencies beyond the usual.
  • You want the shortest path between "I have an idea" and "it is working".
  • The team does not want to maintain Dockerfiles for every small piece.

Choose Cloud Run when:

  • It is an application with several routes, like the Flask catalogue.
  • You need to control the image: a specific version of a system library, a binary, a packaged model.
  • You already have a container — AlpinaShop has had one since 02-05.
  • You want real portability between GKE, Cloud Run and anywhere else.

For AlpinaShop, the allocation looks like this, consistent with decision DA-001 to move the catalogue to Cloud Run (developed in 07-02):

Piece Choice Reason
Flask catalogue website Cloud Run (07-02) A complete application, its own container, DA-001
procesar-imagen-producto Cloud Function One event, one action
Payment gateway webhook Cloud Function A single endpoint with signature verification
Nightly reports Workflows + Cloud Run Job (04-06) Long work, does not fit into 9 minutes

The warning: the distributed monolith of functions

There is an anti-pattern that shows up when people like functions too much, and it is worth recognising before falling into it:

flowchart LR
    A[funcion-validar] --> B[funcion-calcular-precio]
    B --> C[funcion-comprobar-stock]
    C --> D[funcion-reservar]
    D --> E[funcion-cobrar]
    E --> F[funcion-notificar]

Six chained functions to process an order. It looks modular. In reality it is a monolith chopped up with network latency between its lines of code, and it inherits all the drawbacks of both worlds:

  • Accumulated latency: six possible cold starts instead of one.
  • Very difficult debugging: following a request means correlating six sets of logs. It is exactly the problem that the distributed tracing in 06-06 solves, but which it is better not to have.
  • No transactions: if funcion-cobrar works and funcion-notificar fails, the system is left in an inconsistent state and has to be compensated by hand.
  • Changes that cross six deployments: modifying the flow means coordinating six functions.
  • Multiplied cost: six invocations and six times the overhead.

The warning sign is simple: if your functions call each other synchronously, they should probably be a single service. And if you really do need a multi-step stateful flow, the right tool is not chaining functions: it is Workflows, the orchestration AlpinaShop already chose in 04-06, which manages state, retries and errors explicitly.

Functions shine when they are leaves of the tree: they react to an event, do their job and finish. When they start becoming intermediate nodes coordinating others, they have been chosen badly.

Common Mistakes and Tips

The Cloud Storage infinite loop. A function triggered by a bucket that writes to that same bucket. It is the most expensive mistake in this lesson. Defend yourself with at least two of the three barriers: prefix in the notification filter, guard in the code and a different bucket for the outputs.

Assuming events arrive exactly once. They arrive at least once. Every event-driven function needs a stable identifier and an idempotency check. Without that, your data will have duplicates and you will not know why.

Retrying permanent errors. A corrupt image does not fix itself on the fifth attempt. Raise an exception only if retrying can work; for everything else, log it and return normally.

Creating the clients inside the function. Hundreds of milliseconds wasted on every invocation. Clients go at module level. But only clients and constants: any mutable global state is a race condition waiting for you to raise the concurrency.

Deploying with --allow-unauthenticated "just to test". That "just to test" sticks around. Use --no-allow-unauthenticated and grant run.invoker to whoever should be calling it.

Leaving the default service account. It is the Compute Engine one, with Editor. An account of its own per function, with the exact permissions.

Not setting --max-instances. A spike of events can open hundreds of connections to Cloud SQL and bring the database down. Unlimited scaling is not a virtue if what is behind it does not scale the same way.

Putting secrets into environment variables. Use --set-secrets with Secret Manager: it is not left in the configuration, it does not show up in the console and it is rotated without redeploying.

Adding a VPC connector "just in case". It costs money permanently. Only if you need to reach private IPs.

Using the 1st generation because you copied an old tutorial. --gen2 always: concurrency, more memory, more time and far more event sources.

And the tip that sums up section 13: if your functions call one another, stop and rethink. It is probably a service, or a Workflow.

Exercises

Exercise 1: design an idempotent function for orders

AlpinaShop wants a function triggered by the pedidos-nuevos topic that sends a confirmation email to the customer and adds a row to the billing table in BigQuery. Sending an email is not idempotent: the customer is annoyed if they receive three. Design the function explaining the idempotency mechanism, which errors you would retry and which you would not, and what deployment configuration you would use. Point out the exact place where a race condition could produce a duplicate email and how you would mitigate it.

Exercise 2: decide between a function, a service and a workflow

For each of these four AlpinaShop cases, choose Cloud Function, Cloud Run or Workflows, and justify the decision with criteria from this lesson: (a) generating the PDF invoice for an order, about 2 seconds per invoice, triggered after payment; (b) the internal administration panel, about 30 HTTP routes, used by 5 people during office hours; (c) the nightly process that recalculates recommendations, about 40 minutes, with 6 dependent sequential steps; (d) resizing images from customer reviews, with spikes of 500 images within a few minutes after a campaign.

Exercise 3: diagnose an unexpected bill

One Monday, Marta sees that the weekend's bill has been 40 times the usual. The data: procesar-imagen-producto recorded 1.2 million invocations in 48 hours against the usual 600; the imagenes_vision table has 1.2 million new rows, many of them with the same ruta_gcs; the alpinashop-catalogo bucket has grown by 300 GB; and the imagenes-subidas-dlq DLQ is empty. On Friday a "minor" change was deployed: also saving a black and white version of each photo. Diagnose the cause, explain why the empty DLQ is a clue rather than a reassurance, and detail the immediate containment and prevention measures.

Solutions

Solution 1

The core problem: the function has two effects with opposite properties. Inserting into BigQuery is controllable with row_ids; sending an email is irreversible. Once it is sent, there is no undo.

Idempotency mechanism with a marker before sending. The key is to record the intention before the irreversible action, with an atomic conditional write. Firestore fits better than BigQuery here because it offers transactions and low latency:

from google.cloud import firestore

db = firestore.Client()
bq_client = bigquery.Client()

@functions_framework.cloud_event
def confirmar_pedido(event):
    message = json.loads(base64.b64decode(event.data["message"]["data"]))
    id_pedido = message["id_pedido"]
    event_id = event["id"]              # unique identifier of the Pub/Sub message

    ref = db.collection("emails_sent").document(id_pedido)

    # 1. Atomic reservation: create() fails if the document already exists.
    #    It is a server-side atomic operation, not a "read and write".
    try:
        ref.create({
            "status": "sending",
            "event_id": event_id,
            "started_at": firestore.SERVER_TIMESTAMP,
        })
    except google.api_core.exceptions.AlreadyExists:
        doc = ref.get().to_dict()
        if doc["status"] == "sent":
            print(f"Email already sent for {id_pedido}, skipping")
            return                       # acknowledge without retrying
        # Status "sending": another invocation is on it, or died halfway
        if age_of(doc["started_at"]) < 300:
            print(f"Another invocation is processing {id_pedido}")
            return
        print(f"WARNING: orphaned reservation on {id_pedido}, retrying")

    # 2. Irreversible action
    try:
        send_confirmation_email(message)
    except TransientEmailError as e:
        ref.delete()                     # release the reservation so it can be retried
        raise                            # exception → Pub/Sub retries
    except PermanentEmailError as e:
        ref.update({"status": "failed", "error": str(e)})
        flag_for_review(id_pedido, e)
        return                           # do NOT retry

    ref.update({"status": "sent", "sent_at": firestore.SERVER_TIMESTAMP})

    # 3. Action idempotent by design: it can be repeated without harm
    bq_client.insert_rows_json(BILLING_TABLE, [row_from(message)],
                               row_ids=[id_pedido])

Error classification:

Error Retry? Handling
Mail server timeout Yes Release the reservation + exception
5xx from the mail provider Yes The same
Invalid email address No Mark as failed, notify customer support
A required field is missing from the message No Log and return; it is the sender's error
Permission denied in BigQuery It does not help Serious error + alert: it is a configuration failure

Deployment:

gcloud functions deploy confirmar-pedido --gen2 --region=europe-west1 \
  --runtime=python312 --entry-point=confirmar_pedido \
  --trigger-topic=pedidos-nuevos \
  --service-account=sa-confirmar-pedido@alpinashop-prod.iam.gserviceaccount.com \
  --set-secrets='API_KEY_CORREO=api-key-correo:latest' \
  --memory=512Mi --timeout=60s --max-instances=30 --min-instances=1 --retry \
  --project=alpinashop-prod

--min-instances=1 is justified here: a customer who has just paid is waiting for their confirmation, and two seconds of cold start at that moment are noticeable.

The race condition and its mitigation. The gap is between the create() and the send_confirmation_email(). If the instance dies right there, the reservation stays in the sending state forever and the customer never receives the email, because redeliveries will see the reservation and back off.

The code mitigates it with the 300-second timer: a sending reservation older than that is considered orphaned and is retried. The trade-off is explicit and has to be accepted: if the instance did not die but simply took a long time, two emails will be sent.

And here is the underlying lesson: with an irreversible external action there is no such thing as an exactly-once guarantee. You can only choose which side to fail on:

Strategy Risk When to choose it
Mark before sending It may not be sent When duplicating is worse (charges)
Mark after sending It may be sent twice When not sending is worse (notifications)
Mark before + timer Both, with low probability A reasonable compromise

For a confirmation email, an occasional duplicate is annoying but harmless, whereas not sending it generates a call to customer support. That is why the timer is the right choice here. For a card charge, the answer would be the opposite, and the appropriate solution would go through an idempotency key from the payment provider itself.

Solution 2

(a) Invoice PDF: Cloud Function. It is the canonical case: one event, one bounded action, 2 seconds of work, no exotic dependencies. Triggered by pedidos-nuevos or by a dedicated "payment confirmed" topic, with --memory=512Mi and a moderate --max-instances. Nuance: if generating the PDF needed corporate fonts or LaTeX, the system dependency would push towards Cloud Run with a container of its own.

(b) Administration panel: Cloud Run. Thirty HTTP routes are an application, not a function. A Cloud Function has one entry point, and putting a router inside it would be using the tool backwards. Besides, the usage pattern — five people during office hours — makes scaling to zero at night and at weekends ideal, and that is precisely Cloud Run. --min-instances=0, and if the cold start bothers people first thing in the morning, --min-instances=1 during working hours only. It is consistent with DA-001.

(c) 40-minute nightly recalculation: Workflows. Two reasons, each sufficient on its own. First, 40 minutes exceeds the 9-minute limit of an event-driven function. Second, and more important: six dependent sequential steps are orchestration, and chaining them with functions would be the distributed monolith from section 13. Workflows — already chosen in 04-06 — manages state, per-step retries and errors explicitly, and each step invokes whatever is appropriate: a Cloud Run Job, a BigQuery job or a Vertex AI pipeline.

(d) Resizing review images: Cloud Function. It is identical to procesar-imagen-producto: one event, one action, no state. The spikes of 500 images are exactly where the model shines — it scales on its own and goes back to zero afterwards. Configuration: --memory=1Gi, --max-instances=100 so that the spike is absorbed quickly, --min-instances=0 because nobody is waiting in real time, and --retry with idempotency by generation.

Case Choice Decisive criterion
(a) Invoice PDF Cloud Function One event, one action
(b) Administration panel Cloud Run An application with many routes
(c) Nightly recalculation Workflows Exceeds 9 min + it is orchestration
(d) Review images Cloud Function Event-driven, stateless, spiky

The criterion that unifies all four: count how many different things the piece does and how long it takes. One thing and little time → function. Many routes → service. Many coordinated steps → workflow. And when you hesitate between a function and a service for something that is already containerised, Cloud Run almost always wins on portability.

Solution 3

Diagnosis: an infinite loop, exactly the one from section 6.

Friday's "minor" change added writing a black and white version into the same bucket, and in all likelihood into a prefix that is not productos/miniaturas/, which is the only thing the guard in the code checks. Reconstructing the sequence:

  1. productos/piolet.jpg is uploaded → event → the function processes it.
  2. It writes productos/miniaturas/piolet.jpg → event → ignored by the guard. Correct.
  3. It writes productos/bn/piolet.jpg → event → the guard does NOT cover it → it is processed.
  4. Processing productos/bn/piolet.jpg writes productos/bn/bn/piolet.jpg → event → it is processed…

Each level generates the next one. The growth is exponential until something stops it. All four pieces of evidence fit without exception: 1.2 million invocations (recursion); repeated rows with the same ruta_gcs in imagenes_vision (each level writes a row, and idempotency does not protect you because each derived object is a different path); 300 GB of growth (the generated objects); and the runaway bill, dominated by the Vision API, not by the function.

Why the empty DLQ is a clue rather than a reassurance. Instinct says "the DLQ is empty, nothing has failed". It is exactly the other way round: the empty DLQ confirms that everything worked correctly. The function had no errors at all; it did perfectly what it was asked to do, one million two hundred thousand times. That is the danger of event-driven systems: the expensive failure does not produce errors, it produces successes. No alert based on error rate would have caught this. What would have caught it is an alert on the volume of invocations — a subject for 06-04 — and its absence is the real finding of the incident.

Immediate containment, in this order:

# 1. STOP IT NOW: max-instances to 0 halts processing without deleting anything
gcloud functions deploy procesar-imagen-producto --gen2 \
  --region=europe-west1 --max-instances=0 --project=alpinashop-prod

# 2. Drain the queue of pending events, which may be enormous
gcloud pubsub subscriptions seek eventarc-europe-west1-procesar-imagen-sub \
  --time=$(date -u +%Y-%m-%dT%H:%M:%SZ) --project=alpinashop-prod

# 3. Measure the scope before deleting anything
gcloud storage du -s gs://alpinashop-catalogo/productos/bn/ --project=alpinashop-prod

Setting --max-instances=0 instead of deleting the function is deliberate: it stops the bleeding immediately, it preserves the whole configuration for the diagnosis and it is reversible with one command.

Clean-up: delete the derived objects recursively (productos/bn/bn/... and onwards) keeping the first level if it turns out to be useful; remove from imagenes_vision the rows whose ruta_gcs contains /bn/; and check whether any downstream process — the search index, productos_color from 05-05 — consumed that data and needs to be redone.

Prevention, in five measures:

Measure What it prevents Cost
Write the outputs to a different bucket (alpinashop-derivadas) The recursion, at the root None
Notification filter with the prefix productos/originales/ The event being generated at all None
A guard by allowlist, not by blocklist The next forgotten variant None
Alert on invocations per hour for the function Catching it in 15 minutes, not 48 hours None
Budget alert on the project (01-04) Any cost leak recurring None

The third measure is the most transferable design lesson. The original guard was a blocklist: "if it starts with productos/miniaturas/, ignore it". A blocklist fails whenever a new case appears, and one will. The right thing is an allowlist:

ORIGINALS_PREFIX = "productos/originales/"

if not path.startswith(ORIGINALS_PREFIX):
    print(f"Ignored, not an original: {path}")
    return

With that version, Friday's change would have been harmless: the function only processes what is explicitly permitted. In an event-driven architecture, forbidding the known is fragile; permitting only the known is robust.

And the final reflection on the incident: there was no technical failure. There was a one-line change, reviewed by nobody, in a system where writing to a bucket means triggering code. Three things would have prevented it and all three are in this module: the code review from 06-02 — somebody would have asked where the new file gets written —, the unit test from section 12 — which verifies that Vision is not called for derived objects —, and the observability from 06-04 — a volume alert that warns you within minutes. The 40× bill is the price of having none of the three.

Conclusion

Module 5's loose end is tied off. procesar-imagen-producto exists, it fires from imagenes-subidas, it calls the Vision API, it generates the thumbnail, it writes to alpinashop_analitica and it wakes nobody up in the middle of the night.

You know what FaaS is and what its deployment unit is — the function —, with the four properties that define it: it is triggered by events, it scales to zero, it scales up on its own, and it is ephemeral and stateless, with the practical consequence that everything that has to survive lives outside.

You understand what a 2nd generation Cloud Function is today: a Cloud Run service with a container built by Google and triggers managed by Eventarc. And you know what that gives you — configurable concurrency instead of one request per instance, up to 60 minutes over HTTP, up to 32 GB, traffic splitting and more than 130 event sources — with the clear rule of always using --gen2.

You know how to write an HTTP function with the Functions Framework and deploy it with the options that matter, --max-instances included, because unlimited scaling is not a virtue if Cloud SQL does not scale the same way. And you know that almost no function should be public: --no-allow-unauthenticated, run.invoker to whoever needs it and an OIDC token in the caller, with the three legitimate exceptions and what each of them demands.

You know event-driven functions, CloudEvents and the table of what triggers what — Pub/Sub, Cloud Storage, Firestore and the audit logs — and why AlpinaShop goes through a topic instead of triggering straight from the bucket: so that tomorrow a second consumer can react without touching anything.

You have the complete case solved, with the three details that make it viable in production: clients created at module level, the guard against the infinite loop and BigQuery's row_ids. And you have the discipline that separates a function that works from one you can trust: idempotency with a stable identifier — path plus generation —, a time-bounded up-front check and deduplication at the destination; the distinction between transient errors that get retried and permanent ones that do not; and a dead letter with an alert, because a DLQ nobody looks at is worse than not having one.

You know what a cold start is, the five levers for mitigating it and — more importantly — when it matters and when it does not: --min-instances where a customer perceives it, zero in batch processing. You know the danger of mutable global variables with high concurrency, with the rule that only clients and constants go at module level. You know how to inject configuration with environment variables and secrets with --set-secrets from Secret Manager, give each function its own service account, and connect to the VPC with a connector when — and only when — you have to reach a private IP.

You have the limits and the cost with its indicative calculation, and the economic lesson that goes beyond the example: in an event-driven architecture, compute is almost never the main cost; the costs are the APIs you call and the data you move. You know how to test locally, write the six-line unit test that protects against the most expensive possible failure, and deploy from Cloud Build.

And you have the criterion from section 13: a function for one thing triggered by an event, a service for an application with many routes, a workflow for several coordinated steps, with the warning against the distributed monolith of functions and its warning sign — if your functions call each other synchronously, rethink it.

Now look at the state of AlpinaShop. There is a shop on GKE, a function processing images, data pipelines, models training themselves, a CI/CD pipeline and a network infrastructure holding all of that up. Lots of pieces. Many more than fitted three modules ago.

And there is still only one way to know whether they work: a customer writing an email to say the website is slow.

That is the problem in 06-04. It is time to stop looking at the console and start measuring: metrics, dashboards and alerts with Cloud Monitoring, so as to find out about problems before the customers do.

Google Cloud Platform (GCP) Course

Module 1: Introduction to Google Cloud Platform

Module 2: Core GCP Services

Module 3: Networking and Security

Module 4: Data and Analytics

Module 5: Machine Learning and AI

Module 6: DevOps and Monitoring

Module 7: Advanced GCP Topics

Module 8: Final Project

© Copyright 2026. All rights reserved