In 05-05 an explicit loose end was left. AlpinaShop designed an event-driven architecture to process automatically every product image uploaded to the alpinashop-catalogo bucket: a message to the imagenes-subidas topic, a call to the Vision API, the labels written to alpinashop_analitica and a thumbnail generated. The design was clear. The implementation was postponed because one piece was missing: something that runs code when an event happens, without a server sitting there waiting.
That piece is Cloud Functions, and it is what we are going to build here from start to finish.
But first it is worth understanding why the problem is not trivial. The "obvious" solution would be to add an endpoint to the Flask catalogue that receives the notification and processes the image. It is a bad idea for three reasons: processing an image takes seconds and would block a gunicorn worker that should be serving customer requests; the catalogue would scale according to the image load instead of the web traffic; and a processing failure would affect the shop. Reacting to events wants its own life cycle, and that is exactly the gap FaaS fills.
Contents
- What FaaS is and what the deployment unit is
- Cloud Functions today: the 2nd generation on top of Cloud Run and Eventarc
- HTTP functions: your first function and its deployment
- Authentication: why they should almost never be public
- Event-driven functions: CloudEvents and Eventarc
- The pending case:
procesar-imagen-productofrom start to finish - Idempotency, retries and dead letter
- The real life cycle: cold starts and concurrency
- Configuration, secrets and identity
- Connecting to the VPC to reach Cloud SQL
- Limits and cost
- Local testing and deployment from Cloud Build
- When a function and when a service
- What FaaS is and what the deployment unit is
Run through the abstraction continuum from 02-07 in your head, now with one more box:
| Model | Deployment unit | What you manage | What you pay for |
|---|---|---|---|
| IaaS (Compute Engine) | The virtual machine | OS, patches, scaling, everything | The machine while it is on |
| CaaS (GKE) | The container | The cluster and the manifests | The nodes |
| PaaS (App Engine) | The application | The code and its configuration | Instances |
| FaaS (Cloud Functions) | The function | Only the function's code | Invocations and compute time |
Function as a Service takes abstraction to its practical extreme: you write a function — a unit of code with a specific signature —, the platform packages it, deploys it, runs it when something triggers it and shuts it down when it is not needed.
The four properties that define the model:
- It is triggered by an event. An HTTP request, a Pub/Sub message, a new file in a bucket, a modified document in Firestore. The function does not wait: it gets woken up.
- It scales to zero. With no events, there are no instances and no compute is paid for. A radical difference from a MIG VM, which costs the same at four in the morning as it does in the middle of a campaign.
- It scales up on its own. If a thousand events arrive at once, the platform starts instances. There is no autoscaler to configure.
- It is ephemeral and stateless. An instance processes and may disappear. Nothing saved in memory or on local disk survives reliably.
That last property is the hardest to internalise and the one that produces the most design errors. State lives outside: in Cloud SQL, in Firestore, in Cloud Storage, in BigQuery. A function that accumulates something in a global variable and expects to find it on the next invocation works in testing — because it reuses the same instance — and fails in production, intermittently and inexplicably.
- Cloud Functions today: the 2nd generation on top of Cloud Run and Eventarc
This is the section you need to get right so as not to be misled by the old documentation still doing the rounds.
Cloud Functions was born in 2016 with a platform of its own. In 2022 the 2nd generation appeared, and the change is not cosmetic:
A 2nd generation Cloud Function is a Cloud Run service with a container built automatically by Google from your code, and its event triggers are Eventarc subscriptions.
Google builds the container for you using buildpacks, deploys it to Cloud Run and connects Eventarc so that events arrive as HTTP requests. Everything Cloud Run knows how to do, the function inherits.
flowchart TD
A[Your code:<br/>main.py + requirements.txt] --> B[Cloud Build:<br/>buildpacks]
B --> C[Image in<br/>Artifact Registry]
C --> D[Cloud Run service<br/>managed by Cloud Functions]
E[Pub/Sub, Cloud Storage,<br/>Firestore, Audit Logs] --> F[Eventarc]
F -->|CloudEvent over HTTP POST| D
G[Direct HTTP request] --> D
The practical differences, which are the ones that matter when designing:
| Feature | 1st generation | 2nd generation |
|---|---|---|
| Platform | Its own | Cloud Run + Eventarc |
| Concurrency per instance | 1 request | Up to 1,000, configurable |
| Maximum time (HTTP) | 9 minutes | 60 minutes |
| Maximum time (events) | 9 minutes | 9 minutes (via Eventarc) |
| Maximum memory | 8 GB | 32 GB |
| CPU | Tied to the memory | Configurable, up to 8 vCPU |
| Traffic splitting between revisions | No | Yes, like Cloud Run |
| Event sources | A handful of native ones | More than 130 via Eventarc |
| Minimum instances | Limited | Yes |
| Base cost | Lower on minimal workloads | Slightly higher, more capacity |
Concurrency is the most important difference. In the 1st generation, each instance served one request: 100 simultaneous requests meant 100 instances, with 100 cold starts and 100 database connections. In the 2nd, one instance with concurrency 80 serves 80 requests at a time, which drastically reduces cold starts, cost and pressure on Cloud SQL — a very real problem: the number of Cloud SQL connections is limited, and a badly sized 1st generation function exhausts the pool with astonishing ease.
The rule in 2026: always use the 2nd generation. --gen2 is mandatory in this lesson's examples. The 1st generation only shows up in old functions nobody has migrated.
And the natural question: if a 2nd generation function is a Cloud Run, why not use Cloud Run directly? It is an excellent question and it has an answer, but we leave it to section 13, when you have the context to weigh it up.
- HTTP functions: your first function and its deployment
The Functions Framework is the library that turns an ordinary Python function into an HTTP service. It is installed like any other dependency and uses decorators.
The minimum structure of a function is two files:
# main.py
import functions_framework
from flask import jsonify
@functions_framework.http
def comprobar_stock(request):
"""Returns the stock for a SKU. HTTP entry point."""
# request is a Flask Request object: the same API you already know
sku = request.args.get("sku")
if not sku:
return jsonify({"error": "the sku parameter is missing"}), 400
units = look_up_stock(sku) # implementation omitted
if units is None:
return jsonify({"error": "unknown sku"}), 404
return jsonify({"sku": sku, "units": units, "available": units > 0}), 200Three details worth noting:
- The
@functions_framework.httpdecorator marks the entry point. The name of the Python function (comprobar_stock) is what you pass to--entry-point. requestis a Flask object. If you wrote the catalogue with Flask, the API is familiar:request.args,request.get_json(),request.headers.- The return value follows Flask's conventions: a string, a
(body, code)tuple or a complete response.
The deployment:
gcloud functions deploy comprobar-stock \
--gen2 \
--region=europe-west1 \
--runtime=python312 \
--source=. \
--entry-point=comprobar_stock \
--trigger-http \
--no-allow-unauthenticated \
--service-account=sa-funcion-stock@alpinashop-prod.iam.gserviceaccount.com \
--memory=256Mi \
--timeout=30s \
--max-instances=20 \
--project=alpinashop-prodEach option, and why it is there:
| Option | What it does | Why this value |
|---|---|---|
--gen2 |
Uses the 2nd generation | Always, because of section 2 |
--region |
Where it is deployed | europe-west1, next to everything else |
--runtime |
Language version | Pinned explicitly, not implicitly |
--source |
Where the code comes from | . locally; also gs:// or a repository |
--entry-point |
Name of the Python function | It must match exactly |
--trigger-http |
Triggered over HTTP | As opposed to event triggers |
--no-allow-unauthenticated |
Requires IAM authentication | Section 4 |
--service-account |
Execution identity | Never the default account |
--memory |
Memory per instance | It also determines the CPU allocated |
--timeout |
Maximum time per invocation | Short: if it takes longer, something is wrong |
--max-instances |
Scaling ceiling | It protects what is behind it |
--max-instances deserves an explanation, because its absence causes real incidents. With no ceiling, a traffic spike — or an infinite loop, or an attack — can start hundreds of instances that open hundreds of connections to alpinashop-pedidos and bring the database down. Infinite scaling is not a virtue if what is behind it does not scale the same way. Setting a limit turns a total outage into a partial degradation, which is infinitely preferable.
- Authentication: why they should almost never be public
It is tempting to deploy with --allow-unauthenticated because "it is easier to test". That is exactly how data gets leaked.
A public function has a URL that is guessable from a known pattern, it is exposed to the whole internet, anybody can invoke it as many times as they like — and you pay for every invocation — and it usually has IAM permissions over databases and buckets in your project. It is a privileged endpoint with no door.
With --no-allow-unauthenticated, invoking the function requires the roles/run.invoker role — Cloud Run's, consistent with section 2:
# Only the catalogue website can call this function
gcloud functions add-invoker-policy-binding comprobar-stock \
--region=europe-west1 \
--member="serviceAccount:[email protected]" \
--project=alpinashop-prodAnd from the catalogue, the call is authenticated with an OIDC token:
import google.auth.transport.requests
import google.oauth2.id_token
STOCK_URL = "https://europe-west1-alpinashop-prod.cloudfunctions.net/comprobar-stock"
def remote_stock_lookup(sku: str) -> dict:
# The "audience" must be exactly the URL of the function
auth_request = google.auth.transport.requests.Request()
token = google.oauth2.id_token.fetch_id_token(auth_request, STOCK_URL)
response = requests.get(
STOCK_URL,
params={"sku": sku},
headers={"Authorization": f"Bearer {token}"},
timeout=5,
)
response.raise_for_status()
return response.json()The library obtains the token from the environment's identity — the service account attached to the GKE pod — with no key at all. It is the same mechanism as the push subscriptions with OIDC from 04-04.
The three cases in which a function can be public, and what is needed in each:
| Case | What else you need |
|---|---|
| Third-party webhook (payment gateway) | Verify the provider's signature in the request body |
| A genuinely public endpoint (web form) | Cloud Armor in front (03-05), request rate limiting and strict validation |
| Development and testing | That it be in alpinashop-dev and touch nothing real |
The webhook case deserves a note: a public function that receives payment notifications must validate the HMAC signature the provider includes in the header. Without that validation, anybody can send a POST saying "order 1234 is paid". It is a failure that shows up with depressing frequency.
- Event-driven functions: CloudEvents and Eventarc
An event-driven function does not receive an HTTP request: it receives a CloudEvent, an industry-standard format — not a Google invention — with common metadata (id, source, type, time, subject) and a payload specific to the event type.
In Python it is declared with @functions_framework.cloud_event:
import functions_framework
@functions_framework.cloud_event
def process(event):
print("id:", event["id"]) # unique identifier of the event
print("type:", event["type"]) # google.cloud.pubsub.topic.v1.messagePublished
print("source:", event["source"]) # the resource that generated it
data = event.data # the payload, depending on the typeEventarc is GCP's universal event router. It receives events from more than 130 sources, normalises them to CloudEvents and delivers them to the destination. The most useful triggers:
| Source | Event type | It fires when | Typical use at AlpinaShop |
|---|---|---|---|
| Pub/Sub | ...pubsub.topic.v1.messagePublished |
Something is published to a topic | imagenes-subidas (section 6) |
| Cloud Storage | ...storage.object.v1.finalized |
An object is created or overwritten | Direct alternative to the topic |
| Cloud Storage | ...storage.object.v1.deleted |
An object is deleted | Cleaning up orphaned references |
| Firestore | ...firestore.document.v1.written |
A document is written | Reacting to a cart |
| Firestore | ...document.v1.created / .deleted |
Creation or removal | Functional auditing |
| Audit logs | google.cloud.audit.log.v1.written |
Any logged operation | Alerting on firewall changes |
| Cloud Scheduler | Via Pub/Sub | At the scheduled time | Light periodic tasks |
| BigQuery | Via audit logs | A job finishes | Chaining analytical processes |
The audit log trigger is the least known and one of the most powerful: it lets you react to any GCP API operation. For example, running a function every time somebody modifies a firewall rule in alpinashop-prod:
gcloud functions deploy alertar-cambio-firewall \
--gen2 --region=europe-west1 --runtime=python312 \
--entry-point=alertar --source=. \
--trigger-event-filters="type=google.cloud.audit.log.v1.written" \
--trigger-event-filters="serviceName=compute.googleapis.com" \
--trigger-event-filters="methodName=v1.compute.firewalls.insert" \
--service-account=sa-alertas-seguridad@alpinashop-prod.iam.gserviceaccount.com \
--project=alpinashop-prodAn important note about Cloud Storage: you can trigger directly from the bucket (storage.object.v1.finalized) or by publishing to a Pub/Sub topic and triggering from there. AlpinaShop chose the second option in 05-05, and the reason is solid: with Pub/Sub in the middle, several consumers can react to the same event. Today it only processes images; tomorrow, a second subscriber can update the search index without touching anything that already exists. With the direct trigger, every new consumer forces you to reconfigure the bucket.
- The pending case:
procesar-imagen-producto from start to finish
procesar-imagen-producto from start to finishHere the loose end from 05-05 is tied off. Let us recall the flow designed back then:
flowchart LR
A[Dani uploads a photo to<br/>gs://alpinashop-catalogo] -->|notification| B[Pub/Sub topic<br/>imagenes-subidas]
B -->|CloudEvent via Eventarc| C[Function<br/>procesar-imagen-producto]
C --> D[Vision API:<br/>labels, colours, SafeSearch]
C --> E[400px thumbnail<br/>to Cloud Storage]
D --> F[BigQuery<br/>imagenes_vision]
C -.->|error after 5 attempts| G[imagenes-subidas-dlq]
Step 1: notify the topic from the bucket. A single command, which is probably already done since 05-05:
gcloud storage buckets notifications create gs://alpinashop-catalogo \
--topic=imagenes-subidas \
--event-types=OBJECT_FINALIZE \
--object-prefix=productos/ \
--project=alpinashop-prodThe --object-prefix=productos/ avoids processing objects that are not product photos — logos, temporary files — and it is the cheapest way to filter: the event is not even generated.
Step 2: the function. It is commented block by block because each one solves a specific problem:
# main.py
import base64
import json
import os
from datetime import datetime, timezone
import functions_framework
from google.cloud import bigquery, storage, vision
from PIL import Image
import io
# The clients are created ONCE, outside the function.
# They are reused for as long as the instance lives: it saves hundreds of ms per invocation.
vision_client = vision.ImageAnnotatorClient()
storage_client = storage.Client()
bq_client = bigquery.Client()
PROYECTO_DATOS = os.environ["PROYECTO_DATOS"] # alpinashop-datos
TABLE = f"{PROYECTO_DATOS}.alpinashop_analitica.imagenes_vision"
BUCKET_MINIATURAS = os.environ["BUCKET_MINIATURAS"] # alpinashop-catalogo
ANCHO_MINIATURA = 400
@functions_framework.cloud_event
def procesar_imagen(event):
"""Processes a new image: Vision API + thumbnail + BigQuery."""
# 1. Decode the Pub/Sub message. The payload arrives base64-encoded.
message = base64.b64decode(event.data["message"]["data"]).decode("utf-8")
notice = json.loads(message)
bucket_name = notice["bucket"]
path = notice["name"]
generation = notice["generation"] # identifies the exact VERSION of the object
# 2. Guard: ignore what must not be processed.
# Without this, the thumbnail we write below would fire another event
# and we would have an infinite loop that also costs money.
if path.startswith("productos/miniaturas/"):
print(f"Thumbnail ignored: {path}")
return
if not path.lower().endswith((".jpg", ".jpeg", ".png", ".webp")):
print(f"Unsupported format ignored: {path}")
return
# 3. Idempotency: if this generation has already been processed, do not repeat it.
image_id = f"gs://{bucket_name}/{path}#{generation}"
if already_processed(image_id):
print(f"Already processed, skipping: {image_id}")
return
# 4. Vision API over the object in Cloud Storage, without downloading it.
image = vision.Image(source=vision.ImageSource(
gcs_image_uri=f"gs://{bucket_name}/{path}"))
response = vision_client.annotate_image({
"image": image,
"features": [
{"type_": vision.Feature.Type.LABEL_DETECTION, "max_results": 10},
{"type_": vision.Feature.Type.IMAGE_PROPERTIES},
{"type_": vision.Feature.Type.SAFE_SEARCH_DETECTION},
],
})
if response.error.message:
# API error: raise an exception so that Pub/Sub retries
raise RuntimeError(f"Vision API: {response.error.message}")
# 5. Generate the thumbnail and upload it to the prefix that step 2 ignores
blob = storage_client.bucket(bucket_name).blob(path)
original = Image.open(io.BytesIO(blob.download_as_bytes()))
original.thumbnail((ANCHO_MINIATURA, ANCHO_MINIATURA))
output = io.BytesIO()
original.convert("RGB").save(output, format="JPEG", quality=82)
thumb_path = f"productos/miniaturas/{os.path.basename(path)}"
storage_client.bucket(BUCKET_MINIATURAS).blob(thumb_path).upload_from_string(
output.getvalue(), content_type="image/jpeg")
# 6. Write to BigQuery, with the identifier that guarantees idempotency
row = {
"id_imagen": image_id,
"ruta_gcs": f"gs://{bucket_name}/{path}",
"ruta_miniatura": f"gs://{BUCKET_MINIATURAS}/{thumb_path}",
"etiquetas": [
{"descripcion": l.description, "puntuacion": round(l.score, 4)}
for l in response.label_annotations
],
"color_dominante": dominant_colour(response),
"safesearch_adulto": response.safe_search_annotation.adult.name,
"procesada_en": datetime.now(timezone.utc).isoformat(),
}
errors = bq_client.insert_rows_json(TABLE, [row], row_ids=[image_id])
if errors:
raise RuntimeError(f"BigQuery: {errors}")
print(f"Processed successfully: {image_id}")The three details that make this work in production, and that separate a tutorial example from real code:
The clients outside the function. Creating an ImageAnnotatorClient means resolving credentials and establishing connections: hundreds of milliseconds. By creating it at module level, it is created once per instance and reused across every invocation that instance serves. In a function with concurrency 20 and a thousand events, the difference is enormous.
The guard against the infinite loop (step 2). It is the classic mistake with functions triggered by Cloud Storage and it deserves a slow explanation: the function writes the thumbnail into the same bucket, which generates an OBJECT_FINALIZE event, which fires the function, which generates another thumbnail, which fires the function… Each iteration costs invocations, Vision API calls — which are billed — and rows in BigQuery. A loop like that discovered on a Monday morning may have burned through an entire budget over the weekend. The three possible defences are the prefix in the notification filter, the guard in the code and writing to a different bucket; use at least two.
The row_ids in insert_rows_json. BigQuery deduplicates by that identifier within a time window. It is a second net beneath the idempotency check in step 3.
Step 3: deploy it with its identity and its permissions.
# Its own service account with minimum permissions
gcloud iam service-accounts create sa-procesar-imagen \
--display-name="Function procesar-imagen-producto" --project=alpinashop-prod
[email protected]
gcloud projects add-iam-policy-binding alpinashop-prod \
--member="serviceAccount:${SA}" --role=roles/storage.objectAdmin
gcloud projects add-iam-policy-binding alpinashop-datos \
--member="serviceAccount:${SA}" --role=roles/bigquery.dataEditor
gcloud projects add-iam-policy-binding alpinashop-prod \
--member="serviceAccount:${SA}" --role=roles/eventarc.eventReceiver
gcloud functions deploy procesar-imagen-producto \
--gen2 --region=europe-west1 --runtime=python312 \
--source=. --entry-point=procesar_imagen \
--trigger-topic=imagenes-subidas \
--service-account="${SA}" \
--set-env-vars=PROYECTO_DATOS=alpinashop-datos,BUCKET_MINIATURAS=alpinashop-catalogo \
--memory=1Gi \
--timeout=120s \
--max-instances=50 \
--retry \
--project=alpinashop-prod--memory=1Gi is not a whim: opening a 12-megapixel product photo with Pillow consumes a fair amount of memory, and in the 2nd generation the allocated CPU grows with the memory, so it also runs faster. --retry enables retries, which is the subject of the next section.
- Idempotency, retries and dead letter
This is the part that separates a function that works from a function you can trust. And it starts from a fact already established in 04-04:
Pub/Sub guarantees at least once delivery. Your function will receive duplicate messages. It is not a remote possibility: it is a statistical certainty.
Duplicates arrive for perfectly normal reasons: the function processed correctly but took longer than the acknowledgement deadline; there was a network failure while acknowledging; the instance restarted after processing and before acknowledging; or the event was redelivered because a previous invocation raised an exception.
Idempotent means that processing the same event N times produces the same result as processing it once. Let us analyse our function's three actions:
| Action | Idempotent by nature? | What would happen with a duplicate |
|---|---|---|
| Calling the Vision API | Yes in result, no in cost | You pay twice for the same thing |
| Writing the thumbnail | Yes: same path, it is overwritten | No harm, just wasted compute |
| Inserting into BigQuery | No | Duplicate row: the statistics lie |
The third is the dangerous one, and that is why the design includes three layers of defence:
Layer 1, the stable identifier. gs://bucket/path#generation identifies a specific version of an object. Notice that the path alone would not be enough: if Dani uploads a corrected photo with the same name, it is a different object that does need reprocessing, and the generation distinguishes it.
Layer 2, the up-front check with a lightweight control table:
def already_processed(image_id: str) -> bool:
"""Checks in a control table whether this id has already been processed."""
query = f"""
SELECT 1 FROM `{PROYECTO_DATOS}.alpinashop_analitica.imagenes_procesadas`
WHERE id_imagen = @id
AND procesada_en > TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 7 DAY)
LIMIT 1
"""
job = bq_client.query(query, job_config=bigquery.QueryJobConfig(
query_parameters=[bigquery.ScalarQueryParameter("id", "STRING", image_id)]))
return next(job.result(), None) is not NoneThe 7-day window is a deliberate compromise: Pub/Sub duplicates arrive within seconds or minutes, not days, and limiting the query lets you take advantage of the table's partitioning instead of scanning the full history on every invocation. It is the same cost logic as in 04-01.
Layer 3, BigQuery's row_ids, which deduplicates automatically even if the previous two fail because of a race condition.
Retries. With --retry, if the function raises an exception, Pub/Sub receives no acknowledgement and redelivers the message. And here there is a design decision that is constantly got wrong:
| Type of error | Example | Retry? | What to do |
|---|---|---|---|
| Transient | Vision API returning 503, network timeout | Yes | Raise an exception |
| Permanent | Corrupt image, unsupported format | No | Log it and return normally |
| Configuration | An IAM permission is missing | It does not help | Log as a serious error and alert |
A permanent error that gets retried is a function failing forever on the same message, consuming quota and filling the logs. The rule: only raise an exception if trying again has some chance of working.
try:
original = Image.open(io.BytesIO(blob.download_as_bytes()))
except UnidentifiedImageError:
# PERMANENT error: retrying a thousand times will not fix a corrupt file
print(f"PERMANENT ERROR: unreadable image {path}")
record_permanent_failure(image_id, "unreadable_image")
return # normal return → Pub/Sub acknowledges and does not retryDead letter. Even with the distinction above, some message will fail indefinitely. The dead letter topic stops it blocking the queue:
gcloud pubsub topics create imagenes-subidas-dlq --project=alpinashop-prod
gcloud pubsub subscriptions update eventarc-europe-west1-procesar-imagen-sub \
--dead-letter-topic=projects/alpinashop-prod/topics/imagenes-subidas-dlq \
--max-delivery-attempts=5 \
--project=alpinashop-prodAfter five attempts, the message goes to the DLQ instead of being retried forever. And a DLQ nobody looks at is worse than having no DLQ, because it creates a false sense of control: you have to put an alert on its number of unacknowledged messages — exactly the kind of alert configured in 06-04.
- The real life cycle: cold starts and concurrency
A function instance goes through three phases, and understanding which one is expensive explains almost all the observed behaviour:
flowchart LR
A[No instances<br/>cost 0] -->|an event arrives| B[COLD START<br/>create instance +<br/>load code +<br/>run the module]
B --> C[Invocation<br/>WARM START]
C -->|more events| C
C -->|no events<br/>for a few minutes| A
The cold start is the time from the event arriving to your code starting to run: provisioning the instance, loading the runtime, importing the dependencies and running the module-level code. It ranges from a few hundred milliseconds to several seconds, and it depends above all on how heavy your imports are.
The five levers for mitigating it, ordered by real effectiveness:
| Lever | How | Effect | Cost |
|---|---|---|---|
--min-instances |
Keep N instances always alive | Removes cold starts for the baseline traffic | You pay for the idle instance |
| Fewer dependencies | Import only what is needed | Reduces start-up a lot | None |
| Deferred imports | import inside the function |
Only if it is not always used | Complicates the code |
| High concurrency | --concurrency=80 |
Fewer instances, fewer starts | Requires thread-safe code |
| More CPU | --cpu=2 |
Faster start-up | More expensive per second |
# Critical function on the customer path: no cold starts
gcloud functions deploy comprobar-stock --gen2 --region=europe-west1 \
--min-instances=2 --max-instances=50 --concurrency=80 \
--project=alpinashop-prodWhen cold starts matter and when they do not, which is the decision you really have to make:
| Situation | Does it matter? | Recommendation |
|---|---|---|
| A function on the path of a customer request | A lot | --min-instances ≥ 1 |
| Payment gateway webhook | Yes: there are timeouts | --min-instances=1 |
procesar-imagen-producto |
No | --min-instances=0: two extra seconds bother nobody |
| Scheduled nightly task | No | 0 |
--min-instances costs money even with no traffic, and that is the price of partially giving up "scale to zero". Setting it by default on every function cancels out one of the model's economic advantages. Set it where a customer perceives the latency.
About concurrency, an important warning: with --concurrency=80, eighty requests run simultaneously in the same Python process. Any mutable global variable becomes shared:
# DANGEROUS with concurrency > 1
results = [] # shared between simultaneous invocations!
@functions_framework.http
def process(request):
results.append(request.args["id"]) # race condition
return str(len(results)) # returns anything at all
# CORRECT: global objects only for reusable, immutable clients
bq_client = bigquery.Client() # safe: it is designed for concurrent use
@functions_framework.http
def process(request):
local_results = [] # state inside the invocation
...The rule: at module level, only clients and constants. All mutable state, inside the function.
- Configuration, secrets and identity
The three things every production function needs to have properly sorted.
Environment variables for non-sensitive configuration:
gcloud functions deploy procesar-imagen-producto --gen2 \
--set-env-vars=PROYECTO_DATOS=alpinashop-datos,ENTORNO=prod,ANCHO_MINIATURA=400 \
--region=europe-west1 --project=alpinashop-prodSecrets from Secret Manager, never as an environment variable carrying the value:
gcloud functions deploy procesar-pago --gen2 \
--set-secrets='API_KEY_PASARELA=api-key-pasarela-pago:latest' \
--region=europe-west1 --project=alpinashop-prodWith that syntax, the function reads os.environ["API_KEY_PASARELA"] as normal, but the value is not in the deployment configuration: it is injected at run time from Secret Manager (03-06). Practical consequences: the value does not show up in the console, it is not left in the deployment history, it can be rotated without redeploying if you use :latest, and access is recorded in the audit logs. They can also be mounted as a file with --set-secrets='/etc/claves/api=secreto:latest', which is preferable for large values such as certificates.
And the detail that gets forgotten: the function's service account needs roles/secretmanager.secretAccessor on that specific secret, granted in the secret's policy.
Identity. By default, a function uses the project's Compute Engine service account, which usually has the Editor role — the same anti-pattern as in 06-01. Always --service-account with an account of its own per function. A function that only writes to BigQuery must not be able to delete buckets.
- Connecting to the VPC to reach Cloud SQL
By default, a function lives outside your VPC: it can go out to the internet but it cannot reach resources with a private IP. If alpinashop-pedidos only has a private IP — as it should, according to 03-01 — a function cannot get there.
The bridge is a Serverless VPC Access connector:
gcloud compute networks vpc-access connectors create conector-alpinashop \
--region=europe-west1 \
--network=alpinashop-vpc \
--range=10.8.0.0/28 \
--min-instances=2 --max-instances=4 \
--machine-type=e2-micro \
--project=alpinashop-prod
gcloud functions deploy sincronizar-stock --gen2 \
--vpc-connector=projects/alpinashop-prod/locations/europe-west1/connectors/conector-alpinashop \
--egress-settings=private-ranges-only \
--region=europe-west1 --project=alpinashop-prodDetails that matter:
- The
/28range must be free and must not overlap withsn-web-euw1orsn-datos-euw1. It is a strict requirement and a common source of errors. - The connector is billed per instance and hour, whether it is in use or not. With
--min-instances=2you pay for two small machines permanently. It partially breaks "scale to zero", so it is shared between all the functions that need it. --egress-settings=private-ranges-onlysends only traffic to private ranges through the VPC; the rest goes straight out to the internet. Withall-traffic, everything goes through the VPC and out via the Cloud NATalpinashop-nat-euw1from 03-01, which gives a fixed egress IP — useful if a third party demands IP allowlisting.
| Need | Solution | Added cost |
|---|---|---|
| Cloud SQL with a public IP | Cloud SQL connector over TLS | None |
| Cloud SQL with a private IP | VPC connector | Connector instances |
| Fixed egress IP | Connector + all-traffic + Cloud NAT |
Connector + NAT |
| Google APIs only | Nothing: it already works | None |
Tip: do not add a connector "just in case". If the function only talks to Google APIs — Vision, BigQuery, Storage, like procesar-imagen-producto — it does not need one, and adding it is cost and complexity for nothing.
- Limits and cost
The 2nd generation limits worth keeping in mind — always check the current values in the official documentation:
| Limit | Approximate value | What to do if it falls short |
|---|---|---|
| Maximum time (HTTP) | 60 minutes | Long work → Cloud Run Jobs or Dataflow |
| Maximum time (events) | 9 minutes | Split the work, chain it via Pub/Sub |
| Memory | Up to 32 GB | Heavy processing → Cloud Run or Batch |
| CPU | Up to 8 vCPU | Same |
| Size of the deployed code | Tens of MB compressed | Models and data in Cloud Storage |
| Size of the HTTP request | 32 MB | Upload straight to Cloud Storage + event |
| Simultaneous instances | Thousands (quota can be raised) | Request a quota increase |
| Environment variables | A few KB in total | Configuration in Firestore or Secret Manager |
Cost has four components, and there is a generous monthly free tier:
| Component | You pay for | Comment |
|---|---|---|
| Invocations | Each call | Cents per million |
| GB-second | Memory × time | The main lever |
| GHz-second | CPU × time | Tied to the memory |
| Network egress | GB out to the internet | Zero between services in the same region |
An indicative calculation for procesar-imagen-producto. Let us assume 20,000 images a month, 1 GiB of memory, 3 seconds per image:
- Invocations: 20,000 → covered many times over by the free tier.
- GB-second: 20,000 × 3 s × 1 GiB = 60,000 GB-s, mostly within the free tier.
- GHz-second: proportional, same order of magnitude.
- Real compute cost: practically zero.
And now the figure that really matters: the Vision API for those 20,000 images, with three features each, costs an order of magnitude more than the entire execution of the function. The economic lesson is that in event-driven architectures compute is almost never the main cost: the costs are the APIs you call and the data you move. Optimising the function's memory while making redundant calls to a paid API is optimising what does not matter.
That said, three ways for the bill to genuinely explode:
- Infinite loops (section 6). The most expensive by a wide margin.
--min-instanceson functions that do not need it. It is a fixed 24×7 cost and it cancels out the model's advantage.- A long timeout with hung errors. A function with
--timeout=540sthat sits waiting for a downed service pays nine minutes per failed invocation. Short timeouts, matched to reality.
- Local testing and deployment from Cloud Build
Locally, the Functions Framework brings up a server:
pip install functions-framework
functions-framework --target=procesar_imagen --signature-type=cloudevent --port=8080And a Pub/Sub event is simulated with the real structure, base64 included:
DATA=$(printf '{"bucket":"alpinashop-catalogo","name":"productos/piolet-01.jpg","generation":"1712345678"}' | base64 -w0)
curl -X POST http://localhost:8080 \
-H "Content-Type: application/json" \
-H "ce-id: 1234" \
-H "ce-source: //pubsub.googleapis.com/projects/alpinashop-prod/topics/imagenes-subidas" \
-H "ce-type: google.cloud.pubsub.topic.v1.messagePublished" \
-H "ce-specversion: 1.0" \
-d "{\"message\":{\"data\":\"${DATA}\"}}"Better still, unit tests that do not need anything brought up:
# test_main.py
import base64, json
from unittest.mock import patch, MagicMock
from cloudevents.http import CloudEvent
import main
def make_event(bucket, name, generation="1"):
data = base64.b64encode(json.dumps(
{"bucket": bucket, "name": name, "generation": generation}).encode())
return CloudEvent(
{"type": "google.cloud.pubsub.topic.v1.messagePublished",
"source": "//pubsub.googleapis.com/", "id": "1"},
{"message": {"data": data.decode()}})
def test_ignores_thumbnails():
"""The infinite loop guard is the MOST important test of all."""
with patch.object(main, "vision_client") as vision_mock:
main.procesar_imagen(make_event("alpinashop-catalogo",
"productos/miniaturas/x.jpg"))
vision_mock.annotate_image.assert_not_called()
def test_skips_if_already_processed():
with patch.object(main, "already_processed", return_value=True), \
patch.object(main, "vision_client") as vision_mock:
main.procesar_imagen(make_event("alpinashop-catalogo", "productos/a.jpg"))
vision_mock.annotate_image.assert_not_called()The first test deserves a comment: it verifies that the Vision API is not called for a thumbnail. It is the test that protects against the most expensive possible failure of this architecture, and it costs six lines.
Deployment from Cloud Build, fitting in with 06-01 and 06-02:
# cloudbuild.yaml of the functions repository
steps:
- name: 'python:3.12-slim'
id: 'pruebas'
entrypoint: 'bash'
args:
- '-c'
- |
pip install -r requirements.txt -r requirements-dev.txt -t /workspace/lib
PYTHONPATH=/workspace/lib python -m pytest -q
- name: 'gcr.io/google.com/cloudsdktool/cloud-sdk:slim'
id: 'desplegar'
waitFor: ['pruebas']
args:
- 'gcloud'
- 'functions'
- 'deploy'
- 'procesar-imagen-producto'
- '--gen2'
- '--region=europe-west1'
- '--runtime=python312'
- '--source=.'
- '--entry-point=procesar_imagen'
- '--trigger-topic=imagenes-subidas'
- '--service-account=sa-procesar-imagen@alpinashop-prod.iam.gserviceaccount.com'
- '--set-env-vars=PROYECTO_DATOS=alpinashop-datos,BUCKET_MINIATURAS=alpinashop-catalogo'
- '--memory=1Gi'
- '--max-instances=50'
- '--retry'
- '--project=alpinashop-prod'
options:
logging: CLOUD_LOGGING_ONLYThe Cloud Build service account needs roles/cloudfunctions.developer and roles/iam.serviceAccountUser over sa-procesar-imagen — this second permission is always forgotten and produces a permissions error that never mentions the word "function".
- When a function and when a service
We come back to the question left open in section 2: if a 2nd generation function is a Cloud Run, why choose a function?
| Criterion | Cloud Functions (2nd gen) | Cloud Run |
|---|---|---|
| What you deploy | Source code | A container |
| Who builds the image | Google, with buildpacks | You, with your Dockerfile |
| Control of the environment | The runtime Google offers | Total |
| HTTP routes | One function, one entry point | Any complete web application |
| Event triggers | Built into the deployment | Eventarc configured separately |
| Learning curve | Very low | Medium |
| Migrating elsewhere | Depends on the platform | A container runs anywhere |
| System dependencies | Those of the runtime | Whatever you install |
Choose Cloud Functions when:
- The work is one single thing triggered by an event:
procesar-imagen-productois the perfect example. - You do not need system dependencies beyond the usual.
- You want the shortest path between "I have an idea" and "it is working".
- The team does not want to maintain Dockerfiles for every small piece.
Choose Cloud Run when:
- It is an application with several routes, like the Flask catalogue.
- You need to control the image: a specific version of a system library, a binary, a packaged model.
- You already have a container — AlpinaShop has had one since 02-05.
- You want real portability between GKE, Cloud Run and anywhere else.
For AlpinaShop, the allocation looks like this, consistent with decision DA-001 to move the catalogue to Cloud Run (developed in 07-02):
| Piece | Choice | Reason |
|---|---|---|
| Flask catalogue website | Cloud Run (07-02) | A complete application, its own container, DA-001 |
procesar-imagen-producto |
Cloud Function | One event, one action |
| Payment gateway webhook | Cloud Function | A single endpoint with signature verification |
| Nightly reports | Workflows + Cloud Run Job (04-06) | Long work, does not fit into 9 minutes |
The warning: the distributed monolith of functions
There is an anti-pattern that shows up when people like functions too much, and it is worth recognising before falling into it:
flowchart LR
A[funcion-validar] --> B[funcion-calcular-precio]
B --> C[funcion-comprobar-stock]
C --> D[funcion-reservar]
D --> E[funcion-cobrar]
E --> F[funcion-notificar]
Six chained functions to process an order. It looks modular. In reality it is a monolith chopped up with network latency between its lines of code, and it inherits all the drawbacks of both worlds:
- Accumulated latency: six possible cold starts instead of one.
- Very difficult debugging: following a request means correlating six sets of logs. It is exactly the problem that the distributed tracing in 06-06 solves, but which it is better not to have.
- No transactions: if
funcion-cobrarworks andfuncion-notificarfails, the system is left in an inconsistent state and has to be compensated by hand. - Changes that cross six deployments: modifying the flow means coordinating six functions.
- Multiplied cost: six invocations and six times the overhead.
The warning sign is simple: if your functions call each other synchronously, they should probably be a single service. And if you really do need a multi-step stateful flow, the right tool is not chaining functions: it is Workflows, the orchestration AlpinaShop already chose in 04-06, which manages state, retries and errors explicitly.
Functions shine when they are leaves of the tree: they react to an event, do their job and finish. When they start becoming intermediate nodes coordinating others, they have been chosen badly.
Common Mistakes and Tips
The Cloud Storage infinite loop. A function triggered by a bucket that writes to that same bucket. It is the most expensive mistake in this lesson. Defend yourself with at least two of the three barriers: prefix in the notification filter, guard in the code and a different bucket for the outputs.
Assuming events arrive exactly once. They arrive at least once. Every event-driven function needs a stable identifier and an idempotency check. Without that, your data will have duplicates and you will not know why.
Retrying permanent errors. A corrupt image does not fix itself on the fifth attempt. Raise an exception only if retrying can work; for everything else, log it and return normally.
Creating the clients inside the function. Hundreds of milliseconds wasted on every invocation. Clients go at module level. But only clients and constants: any mutable global state is a race condition waiting for you to raise the concurrency.
Deploying with --allow-unauthenticated "just to test". That "just to test" sticks around. Use --no-allow-unauthenticated and grant run.invoker to whoever should be calling it.
Leaving the default service account. It is the Compute Engine one, with Editor. An account of its own per function, with the exact permissions.
Not setting --max-instances. A spike of events can open hundreds of connections to Cloud SQL and bring the database down. Unlimited scaling is not a virtue if what is behind it does not scale the same way.
Putting secrets into environment variables. Use --set-secrets with Secret Manager: it is not left in the configuration, it does not show up in the console and it is rotated without redeploying.
Adding a VPC connector "just in case". It costs money permanently. Only if you need to reach private IPs.
Using the 1st generation because you copied an old tutorial. --gen2 always: concurrency, more memory, more time and far more event sources.
And the tip that sums up section 13: if your functions call one another, stop and rethink. It is probably a service, or a Workflow.
Exercises
Exercise 1: design an idempotent function for orders
AlpinaShop wants a function triggered by the pedidos-nuevos topic that sends a confirmation email to the customer and adds a row to the billing table in BigQuery. Sending an email is not idempotent: the customer is annoyed if they receive three. Design the function explaining the idempotency mechanism, which errors you would retry and which you would not, and what deployment configuration you would use. Point out the exact place where a race condition could produce a duplicate email and how you would mitigate it.
Exercise 2: decide between a function, a service and a workflow
For each of these four AlpinaShop cases, choose Cloud Function, Cloud Run or Workflows, and justify the decision with criteria from this lesson: (a) generating the PDF invoice for an order, about 2 seconds per invoice, triggered after payment; (b) the internal administration panel, about 30 HTTP routes, used by 5 people during office hours; (c) the nightly process that recalculates recommendations, about 40 minutes, with 6 dependent sequential steps; (d) resizing images from customer reviews, with spikes of 500 images within a few minutes after a campaign.
Exercise 3: diagnose an unexpected bill
One Monday, Marta sees that the weekend's bill has been 40 times the usual. The data: procesar-imagen-producto recorded 1.2 million invocations in 48 hours against the usual 600; the imagenes_vision table has 1.2 million new rows, many of them with the same ruta_gcs; the alpinashop-catalogo bucket has grown by 300 GB; and the imagenes-subidas-dlq DLQ is empty. On Friday a "minor" change was deployed: also saving a black and white version of each photo. Diagnose the cause, explain why the empty DLQ is a clue rather than a reassurance, and detail the immediate containment and prevention measures.
Solutions
Solution 1
The core problem: the function has two effects with opposite properties. Inserting into BigQuery is controllable with row_ids; sending an email is irreversible. Once it is sent, there is no undo.
Idempotency mechanism with a marker before sending. The key is to record the intention before the irreversible action, with an atomic conditional write. Firestore fits better than BigQuery here because it offers transactions and low latency:
from google.cloud import firestore
db = firestore.Client()
bq_client = bigquery.Client()
@functions_framework.cloud_event
def confirmar_pedido(event):
message = json.loads(base64.b64decode(event.data["message"]["data"]))
id_pedido = message["id_pedido"]
event_id = event["id"] # unique identifier of the Pub/Sub message
ref = db.collection("emails_sent").document(id_pedido)
# 1. Atomic reservation: create() fails if the document already exists.
# It is a server-side atomic operation, not a "read and write".
try:
ref.create({
"status": "sending",
"event_id": event_id,
"started_at": firestore.SERVER_TIMESTAMP,
})
except google.api_core.exceptions.AlreadyExists:
doc = ref.get().to_dict()
if doc["status"] == "sent":
print(f"Email already sent for {id_pedido}, skipping")
return # acknowledge without retrying
# Status "sending": another invocation is on it, or died halfway
if age_of(doc["started_at"]) < 300:
print(f"Another invocation is processing {id_pedido}")
return
print(f"WARNING: orphaned reservation on {id_pedido}, retrying")
# 2. Irreversible action
try:
send_confirmation_email(message)
except TransientEmailError as e:
ref.delete() # release the reservation so it can be retried
raise # exception → Pub/Sub retries
except PermanentEmailError as e:
ref.update({"status": "failed", "error": str(e)})
flag_for_review(id_pedido, e)
return # do NOT retry
ref.update({"status": "sent", "sent_at": firestore.SERVER_TIMESTAMP})
# 3. Action idempotent by design: it can be repeated without harm
bq_client.insert_rows_json(BILLING_TABLE, [row_from(message)],
row_ids=[id_pedido])Error classification:
| Error | Retry? | Handling |
|---|---|---|
| Mail server timeout | Yes | Release the reservation + exception |
| 5xx from the mail provider | Yes | The same |
| Invalid email address | No | Mark as failed, notify customer support |
| A required field is missing from the message | No | Log and return; it is the sender's error |
| Permission denied in BigQuery | It does not help | Serious error + alert: it is a configuration failure |
Deployment:
gcloud functions deploy confirmar-pedido --gen2 --region=europe-west1 \
--runtime=python312 --entry-point=confirmar_pedido \
--trigger-topic=pedidos-nuevos \
--service-account=sa-confirmar-pedido@alpinashop-prod.iam.gserviceaccount.com \
--set-secrets='API_KEY_CORREO=api-key-correo:latest' \
--memory=512Mi --timeout=60s --max-instances=30 --min-instances=1 --retry \
--project=alpinashop-prod--min-instances=1 is justified here: a customer who has just paid is waiting for their confirmation, and two seconds of cold start at that moment are noticeable.
The race condition and its mitigation. The gap is between the create() and the send_confirmation_email(). If the instance dies right there, the reservation stays in the sending state forever and the customer never receives the email, because redeliveries will see the reservation and back off.
The code mitigates it with the 300-second timer: a sending reservation older than that is considered orphaned and is retried. The trade-off is explicit and has to be accepted: if the instance did not die but simply took a long time, two emails will be sent.
And here is the underlying lesson: with an irreversible external action there is no such thing as an exactly-once guarantee. You can only choose which side to fail on:
| Strategy | Risk | When to choose it |
|---|---|---|
| Mark before sending | It may not be sent | When duplicating is worse (charges) |
| Mark after sending | It may be sent twice | When not sending is worse (notifications) |
| Mark before + timer | Both, with low probability | A reasonable compromise |
For a confirmation email, an occasional duplicate is annoying but harmless, whereas not sending it generates a call to customer support. That is why the timer is the right choice here. For a card charge, the answer would be the opposite, and the appropriate solution would go through an idempotency key from the payment provider itself.
Solution 2
(a) Invoice PDF: Cloud Function. It is the canonical case: one event, one bounded action, 2 seconds of work, no exotic dependencies. Triggered by pedidos-nuevos or by a dedicated "payment confirmed" topic, with --memory=512Mi and a moderate --max-instances. Nuance: if generating the PDF needed corporate fonts or LaTeX, the system dependency would push towards Cloud Run with a container of its own.
(b) Administration panel: Cloud Run. Thirty HTTP routes are an application, not a function. A Cloud Function has one entry point, and putting a router inside it would be using the tool backwards. Besides, the usage pattern — five people during office hours — makes scaling to zero at night and at weekends ideal, and that is precisely Cloud Run. --min-instances=0, and if the cold start bothers people first thing in the morning, --min-instances=1 during working hours only. It is consistent with DA-001.
(c) 40-minute nightly recalculation: Workflows. Two reasons, each sufficient on its own. First, 40 minutes exceeds the 9-minute limit of an event-driven function. Second, and more important: six dependent sequential steps are orchestration, and chaining them with functions would be the distributed monolith from section 13. Workflows — already chosen in 04-06 — manages state, per-step retries and errors explicitly, and each step invokes whatever is appropriate: a Cloud Run Job, a BigQuery job or a Vertex AI pipeline.
(d) Resizing review images: Cloud Function. It is identical to procesar-imagen-producto: one event, one action, no state. The spikes of 500 images are exactly where the model shines — it scales on its own and goes back to zero afterwards. Configuration: --memory=1Gi, --max-instances=100 so that the spike is absorbed quickly, --min-instances=0 because nobody is waiting in real time, and --retry with idempotency by generation.
| Case | Choice | Decisive criterion |
|---|---|---|
| (a) Invoice PDF | Cloud Function | One event, one action |
| (b) Administration panel | Cloud Run | An application with many routes |
| (c) Nightly recalculation | Workflows | Exceeds 9 min + it is orchestration |
| (d) Review images | Cloud Function | Event-driven, stateless, spiky |
The criterion that unifies all four: count how many different things the piece does and how long it takes. One thing and little time → function. Many routes → service. Many coordinated steps → workflow. And when you hesitate between a function and a service for something that is already containerised, Cloud Run almost always wins on portability.
Solution 3
Diagnosis: an infinite loop, exactly the one from section 6.
Friday's "minor" change added writing a black and white version into the same bucket, and in all likelihood into a prefix that is not productos/miniaturas/, which is the only thing the guard in the code checks. Reconstructing the sequence:
productos/piolet.jpgis uploaded → event → the function processes it.- It writes
productos/miniaturas/piolet.jpg→ event → ignored by the guard. Correct. - It writes
productos/bn/piolet.jpg→ event → the guard does NOT cover it → it is processed. - Processing
productos/bn/piolet.jpgwritesproductos/bn/bn/piolet.jpg→ event → it is processed…
Each level generates the next one. The growth is exponential until something stops it. All four pieces of evidence fit without exception: 1.2 million invocations (recursion); repeated rows with the same ruta_gcs in imagenes_vision (each level writes a row, and idempotency does not protect you because each derived object is a different path); 300 GB of growth (the generated objects); and the runaway bill, dominated by the Vision API, not by the function.
Why the empty DLQ is a clue rather than a reassurance. Instinct says "the DLQ is empty, nothing has failed". It is exactly the other way round: the empty DLQ confirms that everything worked correctly. The function had no errors at all; it did perfectly what it was asked to do, one million two hundred thousand times. That is the danger of event-driven systems: the expensive failure does not produce errors, it produces successes. No alert based on error rate would have caught this. What would have caught it is an alert on the volume of invocations — a subject for 06-04 — and its absence is the real finding of the incident.
Immediate containment, in this order:
# 1. STOP IT NOW: max-instances to 0 halts processing without deleting anything
gcloud functions deploy procesar-imagen-producto --gen2 \
--region=europe-west1 --max-instances=0 --project=alpinashop-prod
# 2. Drain the queue of pending events, which may be enormous
gcloud pubsub subscriptions seek eventarc-europe-west1-procesar-imagen-sub \
--time=$(date -u +%Y-%m-%dT%H:%M:%SZ) --project=alpinashop-prod
# 3. Measure the scope before deleting anything
gcloud storage du -s gs://alpinashop-catalogo/productos/bn/ --project=alpinashop-prodSetting --max-instances=0 instead of deleting the function is deliberate: it stops the bleeding immediately, it preserves the whole configuration for the diagnosis and it is reversible with one command.
Clean-up: delete the derived objects recursively (productos/bn/bn/... and onwards) keeping the first level if it turns out to be useful; remove from imagenes_vision the rows whose ruta_gcs contains /bn/; and check whether any downstream process — the search index, productos_color from 05-05 — consumed that data and needs to be redone.
Prevention, in five measures:
| Measure | What it prevents | Cost |
|---|---|---|
Write the outputs to a different bucket (alpinashop-derivadas) |
The recursion, at the root | None |
Notification filter with the prefix productos/originales/ |
The event being generated at all | None |
| A guard by allowlist, not by blocklist | The next forgotten variant | None |
| Alert on invocations per hour for the function | Catching it in 15 minutes, not 48 hours | None |
| Budget alert on the project (01-04) | Any cost leak recurring | None |
The third measure is the most transferable design lesson. The original guard was a blocklist: "if it starts with productos/miniaturas/, ignore it". A blocklist fails whenever a new case appears, and one will. The right thing is an allowlist:
ORIGINALS_PREFIX = "productos/originales/"
if not path.startswith(ORIGINALS_PREFIX):
print(f"Ignored, not an original: {path}")
returnWith that version, Friday's change would have been harmless: the function only processes what is explicitly permitted. In an event-driven architecture, forbidding the known is fragile; permitting only the known is robust.
And the final reflection on the incident: there was no technical failure. There was a one-line change, reviewed by nobody, in a system where writing to a bucket means triggering code. Three things would have prevented it and all three are in this module: the code review from 06-02 — somebody would have asked where the new file gets written —, the unit test from section 12 — which verifies that Vision is not called for derived objects —, and the observability from 06-04 — a volume alert that warns you within minutes. The 40× bill is the price of having none of the three.
Conclusion
Module 5's loose end is tied off. procesar-imagen-producto exists, it fires from imagenes-subidas, it calls the Vision API, it generates the thumbnail, it writes to alpinashop_analitica and it wakes nobody up in the middle of the night.
You know what FaaS is and what its deployment unit is — the function —, with the four properties that define it: it is triggered by events, it scales to zero, it scales up on its own, and it is ephemeral and stateless, with the practical consequence that everything that has to survive lives outside.
You understand what a 2nd generation Cloud Function is today: a Cloud Run service with a container built by Google and triggers managed by Eventarc. And you know what that gives you — configurable concurrency instead of one request per instance, up to 60 minutes over HTTP, up to 32 GB, traffic splitting and more than 130 event sources — with the clear rule of always using --gen2.
You know how to write an HTTP function with the Functions Framework and deploy it with the options that matter, --max-instances included, because unlimited scaling is not a virtue if Cloud SQL does not scale the same way. And you know that almost no function should be public: --no-allow-unauthenticated, run.invoker to whoever needs it and an OIDC token in the caller, with the three legitimate exceptions and what each of them demands.
You know event-driven functions, CloudEvents and the table of what triggers what — Pub/Sub, Cloud Storage, Firestore and the audit logs — and why AlpinaShop goes through a topic instead of triggering straight from the bucket: so that tomorrow a second consumer can react without touching anything.
You have the complete case solved, with the three details that make it viable in production: clients created at module level, the guard against the infinite loop and BigQuery's row_ids. And you have the discipline that separates a function that works from one you can trust: idempotency with a stable identifier — path plus generation —, a time-bounded up-front check and deduplication at the destination; the distinction between transient errors that get retried and permanent ones that do not; and a dead letter with an alert, because a DLQ nobody looks at is worse than not having one.
You know what a cold start is, the five levers for mitigating it and — more importantly — when it matters and when it does not: --min-instances where a customer perceives it, zero in batch processing. You know the danger of mutable global variables with high concurrency, with the rule that only clients and constants go at module level. You know how to inject configuration with environment variables and secrets with --set-secrets from Secret Manager, give each function its own service account, and connect to the VPC with a connector when — and only when — you have to reach a private IP.
You have the limits and the cost with its indicative calculation, and the economic lesson that goes beyond the example: in an event-driven architecture, compute is almost never the main cost; the costs are the APIs you call and the data you move. You know how to test locally, write the six-line unit test that protects against the most expensive possible failure, and deploy from Cloud Build.
And you have the criterion from section 13: a function for one thing triggered by an event, a service for an application with many routes, a workflow for several coordinated steps, with the warning against the distributed monolith of functions and its warning sign — if your functions call each other synchronously, rethink it.
Now look at the state of AlpinaShop. There is a shop on GKE, a function processing images, data pipelines, models training themselves, a CI/CD pipeline and a network infrastructure holding all of that up. Lots of pieces. Many more than fitted three modules ago.
And there is still only one way to know whether they work: a customer writing an email to say the website is slow.
That is the problem in 06-04. It is time to stop looking at the console and start measuring: metrics, dashboards and alerts with Cloud Monitoring, so as to find out about problems before the customers do.
Google Cloud Platform (GCP) Course
Module 1: Introduction to Google Cloud Platform
- What is Google Cloud Platform?
- Setting Up Your GCP Account
- A Tour of the GCP Console
- Projects, Resource Hierarchy and Billing
- Regions, Zones and the Shared Responsibility Model
- Cloud Shell and the gcloud CLI
Module 2: Core GCP Services
- Compute Engine: Virtual Machines on Google Cloud
- Cloud Storage: Object Storage
- Cloud SQL: Managed Relational Databases
- App Engine: Platform as a Service
- Google Kubernetes Engine (GKE)
- NoSQL Databases: Firestore, Bigtable and Spanner
- How to Choose the Right Compute Service
Module 3: Networking and Security
- VPC Networks
- Cloud Load Balancing
- Cloud CDN
- Identity and Access Management (IAM)
- Cloud Armor
- Secrets and Encryption: Secret Manager and Cloud KMS
- Cloud DNS, TLS Certificates and Publishing Services Securely
Module 4: Data and Analytics
- BigQuery: The Analytical Data Warehouse
- Cloud Dataflow: Batch and Streaming Data Processing
- Cloud Dataproc: Managed Spark and Hadoop
- Cloud Pub/Sub: Asynchronous Messaging
- Cloud Data Fusion: Code-Free Data Integration
- Orchestrating Pipelines with Cloud Composer and Workflows
- Data Governance and Dashboards with Dataplex and Looker Studio
Module 5: Machine Learning and AI
- Vertex AI: The Machine Learning Platform on GCP
- AutoML: Custom Models Without Writing Code
- TensorFlow on GCP: Training and Serving Models
- Natural Language API
- Vision API
- Generative AI on Vertex AI: Gemini Models and Embeddings
- MLOps: From Model to Product with Vertex AI Pipelines
Module 6: DevOps and Monitoring
- Cloud Build: Continuous Integration on GCP
- Cloud Source Repositories and Source Code Management
- Cloud Functions: Serverless Functions
- Cloud Monitoring (formerly Stackdriver): Metrics, Dashboards and Alerts
- Cloud Deployment Manager and Native Infrastructure as Code
- Cloud Logging and Cloud Trace: Logs, Traces and Diagnostics
- Terraform on GCP: Infrastructure as Code in Practice
Module 7: Advanced GCP Topics
- Hybrid and Multicloud with Anthos
- Serverless Computing with Cloud Run
- Advanced Networking: Shared VPC, Peering and Hybrid Connectivity
- Security Best Practices
- Cost Management and Optimization
- Reliability: SLOs, High Availability and Disaster Recovery
- Governance at Scale: Organization, Policies and Auditing
