You have the blueprint. Now it is time to build.
And this is where most personal projects go wrong, for a very specific reason: building does not fail because of technical difficulty, it fails because of order. Whoever starts with the application ends up with a pretty app hanging off hand-crafted infrastructure that cannot be reproduced. Whoever does security backwards — opening everything up "so that it works" and promising to lock it down later — never locks it down. Whoever does not control spending discovers in week three that the entire credit has been eaten.
This lesson is a working manual. It does not explain what Cloud Run is or what Terraform is: you already saw that in 07-02 and in 06-07. It explains what order the pieces go in, why that order and not another, and how you check after each step that what you have done is right.
It is organised into nine phases, from 0 to 8. Each phase has its objective, its steps, its commands and — most importantly when you work alone — its verifiable "done" criterion: a command you run and an output you expect. If the command does not give what is expected, the phase is not finished, however much it "looks like" it works.
By the end of the lesson you will have your system up and running, reproducible from scratch, with automated delivery, secure and observed. And you will know exactly where you stand at any moment, which when you work alone is half the battle.
Contents
- The build order and why it is that one
- Phase 0 — Preparation: projects, APIs, budget and repository
- Phase 1 — The foundation with Terraform: state, provider, modules and network
- Phase 2 — Identity: service accounts, minimum roles and keyless federation
- Phase 3 — Data: private database, buckets and migrations as code
- Phase 4 — The application: container, configuration, secrets, health and logs
- Phase 5 — Automated delivery: tests, build, deployment and promotion
- Phase 6 — The data and AI layer
- Phase 7 — Exposure: domain, TLS and protection
- Phase 8 — Observability: logs, dashboard, alerts and SLO
- Cross-cutting practices throughout the implementation
- What to do when you get stuck: the five-step method
- Progress log and "done" criteria
- The realistic warning: the first time, everything takes twice as long
- The build order and why it is that one
flowchart TD
F0["Phase 0 — Preparation<br/>projects · APIs · budget · repo"]
F1["Phase 1 — Foundation<br/>Terraform state · network · firewall"]
F2["Phase 2 — Identity<br/>service accounts · roles · WIF"]
F3["Phase 3 — Data<br/>private DB · buckets · migrations"]
F4["Phase 4 — Application<br/>container · secrets · health · logs"]
F5["Phase 5 — Delivery<br/>CI/CD · tests · promotion"]
F6["Phase 6 — Data and AI<br/>events · BigQuery · dashboard · model"]
F7["Phase 7 — Exposure<br/>domain · TLS · protection"]
F8["Phase 8 — Observability<br/>dashboard · alerts · SLO"]
F0 --> F1 --> F2 --> F3 --> F4 --> F5
F4 --> F6
F5 --> F7
F6 --> F8
F7 --> F8
style F0 fill:#e8f0fe
style F4 fill:#fef7e0
style F8 fill:#e6f4ea
The justification for each dependency, which is what makes the order rational rather than habitual:
| Phase | Depends on | Why exactly |
|---|---|---|
| 0. Preparation | Nothing | The projectId is irreversible and APIs take minutes to propagate. And the budget has to exist before you can spend |
| 1. Foundation | 0 | Terraform needs a project with the Resource Manager API enabled and a bucket for the state. The network cannot be reconfigured without destroying what is inside it |
| 2. Identity | 1 | Service accounts belong to a project, but permissions on specific resources (bucket, secret, topic) require those resources to exist… or for you to create them in the same apply. It comes before data because the DB is created with its user and its secret already |
| 3. Data | 1, 2 | Cloud SQL with a private IP requires the network and the private services access peering to already exist. And its password goes to Secret Manager, which needs the identity that will read it |
| 4. Application | 3 | The app does not start without a DB or secrets. Deploying it earlier forces you to deploy it twice |
| 5. Delivery | 4 | You cannot automate a deployment you have never done by hand. First you do it once and understand it; then you automate it |
| 6. Data and AI | 4 | The events are emitted by the application. No application, no events to process |
| 7. Exposure | 5 | The domain points at a stable service. If the app is still changing shape every day, DNS propagation and the certificate are noise |
| 8. Observability | 6, 7 | You observe the complete system. A dashboard built on half a system has to be redone |
The general rule behind all of this: you build from the irreversible to the reversible, and from what depends on nothing to what depends on everything. The projectId cannot be changed; a Monitoring dashboard is rebuilt in ten minutes. That is why the first is in phase 0 and the second in phase 8.
The legitimate exception: if at any point you are stuck for more than an hour in a phase, jump to the next one that does not depend on it and come back afterwards. Phase 6 (data and AI) and phase 7 (exposure) are independent of each other; phase 8 depends on both. Document the jump in the log.
- Phase 0 — Preparation
Objective: get the ground ready for Terraform to work, with spending under control from the very first minute.
Estimated time: 1-2 hours.
2.1 Create the projects with the names decided in 08-02
# Project variables — adjust them to what you decided in the design
export PROY_BASE="refugio"
export PROY_DEV="${PROY_BASE}-dev"
export PROY_PROD="${PROY_BASE}-prod"
export PROY_DATOS="${PROY_BASE}-datos"
export REGION="europe-west1"
export BILLING_ACCOUNT="0X0X0X-0X0X0X-0X0X0X" # gcloud billing accounts list
# Create the projects
for P in "${PROY_DEV}" "${PROY_PROD}" "${PROY_DATOS}"; do
gcloud projects create "${P}" --name="${P}"
gcloud billing projects link "${P}" --billing-account="${BILLING_ACCOUNT}"
done
# Check
gcloud projects list --filter="projectId:${PROY_BASE}-*" \
--format="table(projectId, name, lifecycleState)"⚠️ Before pressing Enter, read the names out loud. It is the last chance. A
projectIdcannot be changed, cannot be reused even after deleting it, and will appear in every screenshot of your presentation.
If gcloud projects create fails with already exists, it is not that you have it: it is that somebody in the world has it, because the namespace is global. Add a short, distinctive suffix, not a -2.
2.2 Enable the APIs
APIs take from seconds to minutes to propagate. Enable them all in one go now and you save yourself ten interruptions later:
APIS=(
cloudresourcemanager.googleapis.com # Terraform needs this one first
serviceusage.googleapis.com
iam.googleapis.com
iamcredentials.googleapis.com
compute.googleapis.com # network, addresses, firewall
servicenetworking.googleapis.com # peering for private Cloud SQL
vpcaccess.googleapis.com # serverless connector
run.googleapis.com
artifactregistry.googleapis.com
cloudbuild.googleapis.com
sqladmin.googleapis.com
secretmanager.googleapis.com
storage.googleapis.com
pubsub.googleapis.com
bigquery.googleapis.com
cloudfunctions.googleapis.com
eventarc.googleapis.com
cloudscheduler.googleapis.com
monitoring.googleapis.com
logging.googleapis.com
cloudtrace.googleapis.com
language.googleapis.com # replace with your AI API
)
for P in "${PROY_DEV}" "${PROY_PROD}"; do
gcloud services enable "${APIS[@]}" --project="${P}"
done
gcloud services enable bigquery.googleapis.com storage.googleapis.com \
--project="${PROY_DATOS}""Done" criterion:
2.3 The budget, before anything else
You already created it in 08-01. If not, do it now, before creating the first billable resource:
gcloud billing budgets create \
--billing-account="${BILLING_ACCOUNT}" \
--display-name="Final project ${PROY_BASE}" \
--budget-amount=12EUR \
--threshold-rule=percent=0.5 \
--threshold-rule=percent=0.9 \
--threshold-rule=percent=1.0 \
--filter-projects="projects/${PROY_DEV}","projects/${PROY_PROD}","projects/${PROY_DATOS}"
gcloud billing budgets list --billing-account="${BILLING_ACCOUNT}"And enable the billing export to BigQuery from the console (Billing → Billing export). It takes up to 24 hours to start populating data, so the sooner it is enabled, the sooner you will have history for the cost section of your presentation.
2.4 The repository and its structure
If you did exercise 3 of 08-01, you already have it. Extended for the implementation:
mi-proyecto/
├── README.md
├── .gitignore
├── cloudbuild.yaml # dev pipeline
├── cloudbuild-prod.yaml # promotion to prod
├── app/
│ ├── Dockerfile
│ ├── requirements.txt
│ ├── src/
│ │ ├── main.py
│ │ ├── db.py
│ │ ├── eventos.py
│ │ └── observabilidad.py
│ └── tests/
│ ├── test_unitarios.py
│ └── test_integracion.py
├── infra/
│ ├── modules/
│ │ ├── red/
│ │ ├── identidad/
│ │ ├── datos/
│ │ ├── servicio/
│ │ └── observabilidad/
│ ├── envs/
│ │ ├── dev/{main.tf,variables.tf,terraform.tfvars,backend.tf}
│ │ └── prod/{main.tf,variables.tf,terraform.tfvars,backend.tf}
│ └── bootstrap/ # creates the state bucket. Applied once
├── data/
│ ├── seed/generar.py
│ ├── migraciones/
│ │ ├── 001_esquema_inicial.sql
│ │ └── 002_indices.sql
│ └── sql/analitica.sql
├── ml/
│ └── analizar_opiniones.py
└── docs/
├── arquitectura.md
├── runbook.md
├── diario.md
└── adr/The README that actually gets read. Write it now, not at the end, and with this structure:
# RefugioReserva Mountain refuge booking platform. Final project of the GCP course. **All data is fictional.** 🔗 **Demo:** https://refugioreserva.example 📊 **Dashboard:** [Looker Studio](...) 📐 **Architecture:** [docs/arquitectura.md](docs/arquitectura.md) ## What it does A hiker checks the availability of places in 12 refuges and books. The warden sees the day's bookings. The federation checks indicators. ## Architecture in one line Cloud Run (Python/FastAPI) → private Cloud SQL PostgreSQL · photos in Cloud Storage · events over Pub/Sub → BigQuery → Looker Studio · sentiment of reviews with the Natural Language API. All in Terraform, deployed by Cloud Build with no keys (WIF). ## How to bring it up from scratch
cd infra/bootstrap && terraform init && terraform apply cd ../envs/dev && terraform init && terraform apply make seed
## Cost **~€10/month.** See [docs/adr/ADR-006-sin-balanceador.md](docs/adr/) for the most relevant cost decision. ## Status and known technical debt See the "Technical debt" section below. Yes, there is some, and it is prioritised.
Phase 0 "done" criterion:
| Check | Command | Expected result |
|---|---|---|
| Projects created and billable | gcloud billing projects describe ${PROY_PROD} |
billingEnabled: true |
| APIs enabled | gcloud services list --enabled --project=${PROY_PROD} |
≥25 lines |
| Budget created | gcloud billing budgets list --billing-account=${BILLING_ACCOUNT} |
1 budget |
Repository with structure and README |
git log --oneline |
≥2 commits |
- Phase 1 — The foundation with Terraform
Objective: that from here on everything is created with code.
Estimated time: 4-8 hours the first time.
3.1 The state bucket (bootstrap)
There is a chicken-and-egg problem: Terraform stores its state in a bucket, but the bucket also has to be created. The standard solution is a small bootstrap module with local state that is applied just once:
# infra/bootstrap/main.tf
terraform {
required_version = ">= 1.9"
required_providers {
google = { source = "hashicorp/google", version = "~> 6.0" }
}
# No backend: local state, applied once and the .tfstate is uploaded encrypted
# or you simply accept that this module gets re-created by hand if it is lost.
}
provider "google" {
project = var.proyecto_estado
region = var.region
}
resource "google_storage_bucket" "estado" {
name = "${var.prefijo}-terraform-estado"
location = var.region
force_destroy = false
uniform_bucket_level_access = true
public_access_prevention = "enforced"
versioning { enabled = true } # ← essential: lets you recover a corrupted state
lifecycle_rule {
condition { num_newer_versions = 20 }
action { type = "Delete" }
}
labels = {
proyecto = var.prefijo
componente = "iac"
gestionado-por = "terraform"
}
}The three options that are not negotiable:
versioning.enabled = true: if the state gets corrupted (and it happens), the previous version saves your project.public_access_prevention = "enforced": the Terraform state contains resource identifiers and, sometimes, sensitive values. Never public.force_destroy = false: stops an accidentaldestroytaking the state of everything else with it.
cd infra/bootstrap
terraform init && terraform apply
gcloud storage buckets describe gs://refugio-terraform-estado \
--format="value(versioning.enabled,iamConfiguration.publicAccessPrevention)"
# Expected: True enforced3.2 The remote backend and the pinned provider
# infra/envs/prod/backend.tf
terraform {
required_version = ">= 1.9"
backend "gcs" {
bucket = "refugio-terraform-estado"
prefix = "envs/prod" # dev uses "envs/dev": separate states
}
required_providers {
google = {
source = "hashicorp/google"
version = "~> 6.0" # pinned. NEVER without a version
}
random = { source = "hashicorp/random", version = "~> 3.6" }
}
}Why a different prefix per environment and not two buckets: one bucket, two prefixes, two completely independent states. It is simpler to manage and there is no risk of a dev apply touching prod, because they are different state files.
Why pin the provider version: without version, Terraform picks the latest one every time you run init. On some random Tuesday version 7.0 comes out with breaking changes and your plan proposes destroying half your infrastructure. With ~> 6.0 you stay on the 6 branch until you decide to move up yourself, consciously and with a commit that says so.
3.3 Your own modules
The practical rule: create a module when you are going to use the same thing in two environments. With dev and prod, that is nearly everything.
# infra/modules/red/main.tf
variable "proyecto" { type = string }
variable "prefijo" { type = string }
variable "region" { type = string }
variable "cidr_app" { type = string }
variable "cidr_conector" { type = string }
variable "cidr_privado" { type = string }
variable "etiquetas" { type = map(string) }
resource "google_compute_network" "vpc" {
project = var.proyecto
name = "${var.prefijo}-vpc"
auto_create_subnetworks = false # ← NEVER the default network
routing_mode = "REGIONAL"
}
resource "google_compute_subnetwork" "app" {
project = var.proyecto
name = "${var.prefijo}-app-${substr(var.region, 0, 8)}"
network = google_compute_network.vpc.id
region = var.region
ip_cidr_range = var.cidr_app
private_ip_google_access = true # egress to Google APIs without a public IP
}
# Serverless access connector: requires an exact /28
resource "google_vpc_access_connector" "conector" {
project = var.proyecto
name = "${var.prefijo}-conn"
region = var.region
ip_cidr_range = var.cidr_conector
network = google_compute_network.vpc.name
min_instances = 2
max_instances = 3 # cost cap
}
# --- Private services access (needed for Cloud SQL with a private IP) ---
resource "google_compute_global_address" "rango_privado" {
project = var.proyecto
name = "${var.prefijo}-rango-privado"
purpose = "VPC_PEERING"
address_type = "INTERNAL"
prefix_length = 20
address = split("/", var.cidr_privado)[0]
network = google_compute_network.vpc.id
}
resource "google_service_networking_connection" "peering" {
network = google_compute_network.vpc.id
service = "servicenetworking.googleapis.com"
reserved_peering_ranges = [google_compute_global_address.rango_privado.name]
}
# --- Firewall: explicit deny by default ---
resource "google_compute_firewall" "denegar_entrada" {
project = var.proyecto
name = "${var.prefijo}-denegar-entrada"
network = google_compute_network.vpc.name
direction = "INGRESS"
priority = 65534
deny { protocol = "all" }
source_ranges = ["0.0.0.0/0"]
log_config { metadata = "INCLUDE_ALL_METADATA" }
}
output "red_id" { value = google_compute_network.vpc.id }
output "red_nombre" { value = google_compute_network.vpc.name }
output "conector_id" { value = google_vpc_access_connector.conector.id }
output "peering_listo" { value = google_service_networking_connection.peering.id }The two resources people forget are in there: google_compute_global_address with purpose = "VPC_PEERING" and google_service_networking_connection. Without them, Cloud SQL with a private IP fails with an error that mentions peering nowhere at all.
And the output "peering_listo" is not decorative: it lets the data module declare a depends_on against it so that Terraform does not try to create the database before the peering exists.
3.4 Variables per environment
# infra/envs/prod/terraform.tfvars
proyecto = "refugio-prod"
entorno = "prod"
prefijo = "refugio"
region = "europe-west1"
cidr_app = "10.20.0.0/24"
cidr_conector = "10.20.8.0/28"
cidr_privado = "10.20.16.0/20"
bd_tier = "db-f1-micro"
bd_alta_disp = false
bd_backup_retencion = 7
run_min_instancias = 0
run_max_instancias = 5# infra/envs/dev/terraform.tfvars — same keys, smaller values
proyecto = "refugio-dev"
entorno = "dev"
prefijo = "refugio"
region = "europe-west1"
cidr_app = "10.10.0.0/24"
cidr_conector = "10.10.8.0/28"
cidr_privado = "10.10.16.0/20"
bd_tier = "db-f1-micro"
bd_alta_disp = false
bd_backup_retencion = 1
run_min_instancias = 0
run_max_instancias = 2The rule: both files have exactly the same keys. If dev has a variable prod does not have, the environments have diverged and promotion stops being reliable.
3.5 The first apply
cd infra/envs/dev
terraform init
terraform fmt -recursive ../.. # consistent formatting, free
terraform validate # syntax and references
terraform plan -out=plan.tfplan # READ THE WHOLE OUTPUT
terraform apply plan.tfplanRead the whole plan. Always. It is cross-cutting practice number one from section 11, and the first time is when you learn the most: the plan shows you exactly which resources each block you have written implies.
Phase 1 "done" criterion:
| Check | Command | Expected |
|---|---|---|
| Remote state | gcloud storage ls gs://refugio-terraform-estado/envs/dev/ |
default.tfstate |
| Network created, not the default one | gcloud compute networks list --project=${PROY_DEV} |
refugio-vpc, no default |
| Connector active | gcloud compute networks vpc-access connectors list --region=${REGION} --project=${PROY_DEV} |
state READY |
| Peering established | gcloud services vpc-peerings list --network=refugio-vpc --project=${PROY_DEV} |
1 peering |
| Clean plan | terraform plan |
No changes. |
That last row is the one that matters: a plan that says No changes means the code and reality match. It is the criterion checked in 08-04 and it is worth 3 points of the rubric.
- Phase 2 — Identity
Objective: every workload with its own identity and its minimum permissions, and CI/CD working without a single downloaded key.
Estimated time: 3-5 hours (2 of them the first time you configure WIF).
4.1 Service accounts and minimum roles
# infra/modules/identidad/main.tf
resource "google_service_account" "web" {
project = var.proyecto
account_id = "sa-${var.prefijo}-web"
display_name = "Web application account"
description = "Cloud Run: reads the DB, reads 2 secrets, writes photos, publishes events"
}
# --- PROJECT-level roles: only those that cannot be narrowed further ---
resource "google_project_iam_member" "web_proyecto" {
for_each = toset([
"roles/cloudsql.client",
"roles/logging.logWriter",
"roles/cloudtrace.agent",
"roles/monitoring.metricWriter",
])
project = var.proyecto
role = each.value
member = "serviceAccount:${google_service_account.web.email}"
}
# --- Roles scoped TO THE RESOURCE: this is how it is done properly ---
resource "google_secret_manager_secret_iam_member" "web_secretos" {
for_each = toset(var.secretos_de_la_web) # ["refugio-db-password", "refugio-session-key"]
project = var.proyecto
secret_id = each.value
role = "roles/secretmanager.secretAccessor"
member = "serviceAccount:${google_service_account.web.email}"
}
resource "google_storage_bucket_iam_member" "web_fotos" {
bucket = var.bucket_fotos
role = "roles/storage.objectAdmin"
member = "serviceAccount:${google_service_account.web.email}"
}
resource "google_pubsub_topic_iam_member" "web_publica" {
project = var.proyecto
topic = var.topic_eventos
role = "roles/pubsub.publisher"
member = "serviceAccount:${google_service_account.web.email}"
}The difference that separates a good project from a mediocre one is in the names of those resources. google_project_iam_member with secretAccessor gives access to all the secrets in the project, present and future. google_secret_manager_secret_iam_member gives it to that secret. It is the same amount of code and an enormous difference in attack surface.
4.2 Workload Identity Federation: keyless CI/CD
This is the point that will set you apart most. The idea, in two sentences: instead of downloading a JSON key and storing it in GitHub, you establish a trust relationship between GCP and your CI's identity provider. The CI presents a signed token saying "I am the main branch of the repo usuario/refugioreserva", and GCP exchanges it for temporary credentials.
Zero keys. Zero rotation. Zero risk of leakage.
# infra/modules/identidad/wif.tf
resource "google_iam_workload_identity_pool" "github" {
project = var.proyecto
workload_identity_pool_id = "gh-pool"
display_name = "GitHub Actions"
}
resource "google_iam_workload_identity_pool_provider" "github" {
project = var.proyecto
workload_identity_pool_id = google_iam_workload_identity_pool.github.workload_identity_pool_id
workload_identity_pool_provider_id = "gh-provider"
attribute_mapping = {
"google.subject" = "assertion.sub"
"attribute.repository" = "assertion.repository"
"attribute.ref" = "assertion.ref"
}
# MANDATORY CONDITION: without this, ANY GitHub repository
# in the world could authenticate against your project.
attribute_condition = "assertion.repository == '${var.github_repo}'"
oidc { issuer_uri = "https://token.actions.githubusercontent.com" }
}
resource "google_service_account" "deploy" {
project = var.proyecto
account_id = "sa-${var.prefijo}-deploy"
display_name = "Deployment from CI"
}
# Only the repository's main branch can impersonate the deployment account
resource "google_service_account_iam_member" "deploy_wif" {
service_account_id = google_service_account.deploy.name
role = "roles/iam.workloadIdentityUser"
member = "principalSet://iam.googleapis.com/${google_iam_workload_identity_pool.github.name}/attribute.repository/${var.github_repo}"
}
resource "google_project_iam_member" "deploy_roles" {
for_each = toset([
"roles/run.developer",
"roles/artifactregistry.writer",
])
project = var.proyecto
role = each.value
member = "serviceAccount:${google_service_account.deploy.email}"
}
# THE PERMISSION EVERYBODY ALWAYS FORGETS:
# to deploy a service that runs as sa-web, deploy must be able to "act as" it.
resource "google_service_account_iam_member" "deploy_actua_como_web" {
service_account_id = google_service_account.web.name
role = "roles/iam.serviceAccountUser"
member = "serviceAccount:${google_service_account.deploy.email}"
}⚠️ The
attribute_conditionis not optional. Without it, the provider accepts tokens from any GitHub repository on the planet, and anybody who knows your pool's name can deploy into your project. It is a real, documented security flaw that appears in many tutorials by omission.
And in the GitHub Actions workflow:
# .github/workflows/deploy.yml
permissions:
contents: read
id-token: write # essential for GitHub to issue the OIDC token
jobs:
desplegar:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: google-github-actions/auth@v2
with:
workload_identity_provider: projects/123456789/locations/global/workloadIdentityPools/gh-pool/providers/gh-provider
service_account: [email protected]
# From here on, gcloud is authenticated. With no secret in GitHub at all.Phase 2 "done" criterion:
# 1. No service account with a primitive role
gcloud projects get-iam-policy "${PROY_PROD}" --format=json | \
jq -r '.bindings[] | select(.role | test("roles/(owner|editor)")) |
.members[] | select(startswith("serviceAccount:"))'
# Expected: empty
# 2. Zero user-managed service account keys
for SA in $(gcloud iam service-accounts list --project="${PROY_PROD}" --format="value(email)"); do
echo -n "$SA: "
gcloud iam service-accounts keys list --iam-account="$SA" \
--managed-by=user --format="value(name)" | wc -l
done
# Expected: 0 for all of themSave the output of those two commands: they are direct evidence for 6 points of block D of the rubric.
- Phase 3 — Data
Objective: a managed, private database, with its password in Secret Manager, and the schema as versioned code.
Estimated time: 3-5 hours.
5.1 The database, with no public IP
# infra/modules/datos/sql.tf
resource "random_password" "bd" {
length = 32
special = true
}
resource "google_secret_manager_secret" "bd_password" {
project = var.proyecto
secret_id = "${var.prefijo}-db-password"
replication { auto {} }
}
resource "google_secret_manager_secret_version" "bd_password" {
secret = google_secret_manager_secret.bd_password.id
secret_data = random_password.bd.result
}
resource "google_sql_database_instance" "principal" {
project = var.proyecto
name = "${var.prefijo}-db"
region = var.region
database_version = "POSTGRES_16"
# The peering must exist first. A legitimate case for an explicit depends_on.
depends_on = [var.peering_listo]
settings {
tier = var.bd_tier
availability_type = var.bd_alta_disp ? "REGIONAL" : "ZONAL"
disk_size = 10
disk_type = "PD_HDD" # cheaper; enough at this volume
disk_autoresize = true
ip_configuration {
ipv4_enabled = false # ← NO PUBLIC IP
private_network = var.red_id
ssl_mode = "ENCRYPTED_ONLY"
}
backup_configuration {
enabled = true
start_time = "03:00"
point_in_time_recovery_enabled = var.entorno == "prod"
backup_retention_settings { retained_backups = var.bd_backup_retencion }
}
maintenance_window { day = 7, hour = 4 } # early Sunday morning
database_flags {
name = "cloudsql.iam_authentication"
value = "on"
}
user_labels = var.etiquetas
}
# Protection against an accidental destroy in production
deletion_protection = var.entorno == "prod"
}
resource "google_sql_database" "app" {
project = var.proyecto
instance = google_sql_database_instance.principal.name
name = "reservas"
}
resource "google_sql_user" "app" {
project = var.proyecto
instance = google_sql_database_instance.principal.name
name = "app"
password = random_password.bd.result
}Five details that count:
ipv4_enabled = falseis the line worth 2 points of block D and, more importantly, the one that stops your database being scannable from the internet.- The password is generated with
random_passwordand you never see it. It goes straight to Secret Manager. Nobody types it, nobody copies it, nobody uploads it by mistake. depends_onon the peering is one of the few cases where an explicitdepends_onis correct: the dependency is real but Terraform cannot infer it from the graph.deletion_protectionconditional on the environment: dev gets destroyed on Fridays, prod does not get destroyed by accident.disk_type = "PD_HDD": for a portfolio project, SSD adds nothing and costs more.
5.2 Buckets with a lifecycle
resource "random_id" "sufijo" { byte_length = 2 }
resource "google_storage_bucket" "fotos" {
project = var.proyecto
name = "${var.prefijo}-fotos-${random_id.sufijo.hex}" # global name
location = var.region
uniform_bucket_level_access = true
public_access_prevention = "enforced"
versioning { enabled = true }
lifecycle_rule {
condition { age = 90, matches_storage_class = ["STANDARD"] }
action { type = "SetStorageClass", storage_class = "NEARLINE" }
}
lifecycle_rule {
condition { age = 365, matches_storage_class = ["NEARLINE"] }
action { type = "SetStorageClass", storage_class = "COLDLINE" }
}
lifecycle_rule {
condition { num_newer_versions = 3 } # do not pile up old versions
action { type = "Delete" }
}
cors {
origin = ["https://${var.dominio}"]
method = ["GET", "HEAD"]
response_header = ["Content-Type"]
max_age_seconds = 3600
}
labels = merge(var.etiquetas, { componente = "web" })
}public_access_prevention = "enforced" prevents anyone, at bucket level, from making it public either by mistake or on purpose. It is one line and it eliminates the most common cloud incident at the root.
5.3 Schema migrations as code
The schema is not created by hand in a SQL console. It is versioned:
data/migraciones/
├── 001_esquema_inicial.sql
├── 002_indice_disponibilidad.sql
└── 003_columnas_sentimiento.sql-- data/migraciones/001_esquema_inicial.sql
-- Idempotent: it can be run twice without breaking anything.
BEGIN;
CREATE TABLE IF NOT EXISTS refugio (
id SERIAL PRIMARY KEY,
nombre TEXT NOT NULL UNIQUE,
altitud_m INTEGER NOT NULL CHECK (altitud_m BETWEEN 500 AND 3500),
capacidad INTEGER NOT NULL CHECK (capacidad > 0),
activo BOOLEAN NOT NULL DEFAULT TRUE,
creado_en TIMESTAMPTZ NOT NULL DEFAULT NOW()
);
CREATE TABLE IF NOT EXISTS reserva (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
refugio_id INTEGER NOT NULL REFERENCES refugio(id),
fecha DATE NOT NULL,
plazas INTEGER NOT NULL CHECK (plazas BETWEEN 1 AND 12),
nombre_titular TEXT NOT NULL,
email_titular TEXT NOT NULL,
estado TEXT NOT NULL DEFAULT 'confirmada'
CHECK (estado IN ('confirmada','cancelada')),
creada_en TIMESTAMPTZ NOT NULL DEFAULT NOW()
);
-- Record of applied migrations
CREATE TABLE IF NOT EXISTS _migraciones (
version TEXT PRIMARY KEY,
aplicada_en TIMESTAMPTZ NOT NULL DEFAULT NOW()
);
INSERT INTO _migraciones (version) VALUES ('001')
ON CONFLICT (version) DO NOTHING;
COMMIT;To run them against a DB with no public IP, use the auth proxy:
# Download cloud-sql-proxy if you do not have it
./cloud-sql-proxy --port 5432 "${PROY_DEV}:${REGION}:refugio-db" &
PGPASSWORD=$(gcloud secrets versions access latest --secret=refugio-db-password --project="${PROY_DEV}")
export PGPASSWORD
for f in data/migraciones/*.sql; do
echo "→ $f"
psql -h 127.0.0.1 -U app -d reservas -v ON_ERROR_STOP=1 -f "$f"
done
unset PGPASSWORDNote the unset PGPASSWORD and the fact that the password is read from Secret Manager on the spot: it is never written into a file, nor into the shell history if you use HISTCONTROL=ignorespace and a leading space.
5.4 Seeding fictional data
# data/seed/generar.py — generates 100% invented data
import random, uuid
from datetime import date, timedelta
random.seed(42) # reproducible: the same data on every run
REFUGIOS = [
("Refugio de Cotiella", 2100, 40), ("Refugio Peña Blanca", 1850, 28),
("Refugio del Ibón Verde", 2340, 22), ("Refugio de Valdellosa", 1620, 55),
]
NOMBRES = ["Ana", "Luis", "Marta", "Jorge", "Carmen", "Diego", "Elena", "Pablo"]
APELLIDOS = ["Soler", "Ibáñez", "Marín", "Vidal", "Rey", "Castaño", "Lorca"]
def titular_ficticio(i):
nombre = f"{random.choice(NOMBRES)} {random.choice(APELLIDOS)}"
# example.com is reserved by RFC 2606 for exactly this
correo = f"usuario{i:04d}@example.com"
return nombre, correo
def factor_estacional(d: date) -> float:
"""July-August x4, weekends x2.5, the rest normal."""
f = 4.0 if d.month in (7, 8) else (2.0 if d.month in (6, 9) else 1.0)
if d.weekday() >= 5:
f *= 2.5
return f
def generar_reservas(n=2000):
inicio = date.today() - timedelta(days=540)
filas = []
i = 0
while len(filas) < n:
d = inicio + timedelta(days=random.randint(0, 540))
if random.random() > factor_estacional(d) / 10:
continue
i += 1
nombre, correo = titular_ficticio(i)
filas.append((
str(uuid.uuid4()), random.randint(1, len(REFUGIOS)), d.isoformat(),
random.randint(1, 6), nombre, correo,
"cancelada" if random.random() < 0.08 else "confirmada",
))
return filasrandom.seed(42) makes the generator reproducible: if you destroy and recreate the environment, you get exactly the same data. That keeps the dashboard screenshots valid and makes the 08-04 tests deterministic.
Phase 3 "done" criterion:
| Check | Command | Expected |
|---|---|---|
| DB with no public IP | gcloud sql instances describe refugio-db --format="value(settings.ipConfiguration.ipv4Enabled)" |
False |
| Secret created with a version | gcloud secrets versions list refugio-db-password |
≥1 ENABLED version |
| Bucket not public | gcloud storage buckets describe gs://... --format="value(iamConfiguration.publicAccessPrevention)" |
enforced |
| Migrations applied | psql -c "SELECT version FROM _migraciones ORDER BY version" |
All of them |
| Data seeded | psql -c "SELECT count(*) FROM reserva" |
~2,000 |
- Phase 4 — The application
Objective: a container that honours the Cloud Run contract, configured per environment, with injected secrets, health probes and correlatable logs. And deployed by hand once, as a proof of life.
Estimated time: 8-12 hours.
6.1 The Cloud Run contract
Four rules. Breaking any of them makes the deployment fail with a rather uninformative error:
| Rule | What it means | Error if you break it |
|---|---|---|
Listen on $PORT |
The PORT environment variable, not a fixed port |
The container does not pass the startup check |
Listen on 0.0.0.0 |
Not on 127.0.0.1 |
Same: it looks started but does not respond |
| Start in <4 min | No migrations or slow loading at startup | Deployment timeout |
| No state on disk | The filesystem is ephemeral and per instance | Data that disappears with no explanation |
# app/Dockerfile — multi-stage, no root, and with the bare minimum inside
FROM python:3.12-slim AS build
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir --user -r requirements.txt
FROM python:3.12-slim
RUN useradd --create-home --uid 1001 app
WORKDIR /app
COPY --from=build /root/.local /home/app/.local
COPY --chown=app:app src/ ./src/
USER app
ENV PATH=/home/app/.local/bin:$PATH \
PYTHONUNBUFFERED=1
# $PORT is injected by Cloud Run. 8080 is just the local default.
ENV PORT=8080
CMD exec uvicorn src.main:app --host 0.0.0.0 --port ${PORT}USER app is worth 1 security point and costs two lines. PYTHONUNBUFFERED=1 makes logs come out immediately instead of sitting in the buffer; without it, the logs of a container that crashes are lost.
6.2 Configuration and secrets
Non-sensitive configuration → service environment variables. Secrets → Secret Manager, mounted by reference:
resource "google_cloud_run_v2_service" "web" {
project = var.proyecto
name = "${var.prefijo}-web"
location = var.region
ingress = "INGRESS_TRAFFIC_ALL"
template {
service_account = var.sa_web_email
scaling {
min_instance_count = var.run_min_instancias # 0: scales to zero
max_instance_count = var.run_max_instancias # spending cap
}
vpc_access {
connector = var.conector_id
egress = "PRIVATE_RANGES_ONLY" # only the DB goes through the VPC
}
containers {
image = var.imagen
resources {
limits = { cpu = "1", memory = "512Mi" }
cpu_idle = true # do not pay for CPU between requests
}
# --- Configuration: in the clear, not sensitive ---
env { name = "ENTORNO" value = var.entorno }
env { name = "REGION" value = var.region }
env { name = "BD_HOST" value = var.bd_ip_privada }
env { name = "BD_NOMBRE" value = "reservas" }
env { name = "TOPIC_EVENTOS" value = var.topic_eventos }
env { name = "BUCKET_FOTOS" value = var.bucket_fotos }
# --- Secrets: by reference, never by value ---
env {
name = "BD_PASSWORD"
value_source {
secret_key_ref {
secret = var.secreto_bd_password
version = "latest"
}
}
}
startup_probe {
http_get { path = "/salud/arranque" }
initial_delay_seconds = 3
period_seconds = 3
failure_threshold = 10
}
liveness_probe {
http_get { path = "/salud/vivo" }
period_seconds = 30
}
}
}
traffic {
type = "TRAFFIC_TARGET_ALLOCATION_TYPE_LATEST"
percent = 100
}
}cpu_idle = true is the option that stops Cloud Run charging you for CPU between requests. In a portfolio project with sporadic traffic, it is the difference between cents and euros.
6.3 The two probes and why they are different
# app/src/main.py (excerpt)
from fastapi import FastAPI, Response
from . import db
app = FastAPI()
@app.get("/salud/arranque")
async def arranque(response: Response):
"""Startup probe: am I ready to receive traffic?
Checks critical dependencies. If it fails, Cloud Run sends no traffic."""
try:
await db.ping()
return {"estado": "listo"}
except Exception as e:
response.status_code = 503
return {"estado": "no listo", "motivo": str(e)[:200]}
@app.get("/salud/vivo")
async def vivo():
"""Liveness probe: is the process responding?
Does NOT check dependencies: if the DB goes down, we do not want
Cloud Run restarting the instance in a loop — it would fix nothing."""
return {"estado": "vivo"}The distinction is important and it comes up in interviews: startup checks dependencies, liveness does not. If the liveness probe checked the database, a DB outage would trigger continuous restarts of every instance, turning a problem into a storm.
6.4 Structured logs with trace_id
From 06-06, in its minimal, effective form:
# app/src/observabilidad.py
import json, os, sys, contextvars
_trace = contextvars.ContextVar("trace", default=None)
PROYECTO = os.environ.get("GOOGLE_CLOUD_PROJECT", "")
def fijar_trace(cabecera: str | None):
"""Cloud Run sends X-Cloud-Trace-Context: TRACE_ID/SPAN_ID;o=1"""
if cabecera:
_trace.set(cabecera.split("/")[0])
def log(severidad: str, mensaje: str, **campos):
entrada = {
"severity": severidad, # names Cloud Logging understands
"message": mensaje,
**campos,
}
t = _trace.get()
if t and PROYECTO:
# This exact key is what links the log to the trace in the console
entrada["logging.googleapis.com/trace"] = f"projects/{PROYECTO}/traces/{t}"
print(json.dumps(entrada, ensure_ascii=False), file=sys.stdout, flush=True)# In the application middleware
@app.middleware("http")
async def correlacion(request, call_next):
fijar_trace(request.headers.get("X-Cloud-Trace-Context"))
respuesta = await call_next(request)
log("INFO", "peticion",
ruta=request.url.path, metodo=request.method,
codigo=respuesta.status_code)
return respuestaThe logging.googleapis.com/trace key with that exact name is what lets you click on a slow trace in the console and see every log for that specific request. It is one line of code and it transforms debugging.
⚠️ Never log personal data. No emails, no names, no form contents, no tokens. Log identifiers (
reserva_id), not people. Logs get replicated, exported and retained; a piece of personal data in a log is a piece of personal data you have lost sight of.
6.5 The first deployment, by hand
It is done manually and just once. The reason is both pedagogical and practical: understanding each step before automating it, and having a proof of life to compare against when the pipeline fails.
export IMAGEN="${REGION}-docker.pkg.dev/${PROY_DEV}/refugio-imagenes/web:manual-1"
gcloud builds submit app/ --tag="${IMAGEN}" --project="${PROY_DEV}"
gcloud run deploy refugio-web \
--image="${IMAGEN}" \
--region="${REGION}" \
--project="${PROY_DEV}" \
--service-account="sa-refugio-web@${PROY_DEV}.iam.gserviceaccount.com" \
--vpc-connector="refugio-conn" \
--vpc-egress=private-ranges-only \
--set-env-vars="ENTORNO=dev,BD_HOST=10.10.16.3,BD_NOMBRE=reservas" \
--set-secrets="BD_PASSWORD=refugio-db-password:latest" \
--min-instances=0 --max-instances=2 \
--no-allow-unauthenticated
URL=$(gcloud run services describe refugio-web --region="${REGION}" \
--project="${PROY_DEV}" --format="value(status.url)")
curl -s -H "Authorization: Bearer $(gcloud auth print-identity-token)" "${URL}/salud/arranque"Once it works, erase it from your mental history and do it from Terraform. The manual deployment was the proof of life; the permanent state is governed by code.
Phase 4 "done" criterion:
| Check | Command | Expected |
|---|---|---|
| Service deployed | gcloud run services describe refugio-web --region=$REGION --format="value(status.conditions[0].status)" |
True |
| Startup probe OK | curl .../salud/arranque |
{"estado":"listo"} |
| Reads from the DB | curl .../api/refugios |
A list with data |
| Structured logs | gcloud logging read 'resource.type="cloud_run_revision"' --limit=1 --format=json |
JSON with jsonPayload and trace |
| Does not run as root | docker run --rm $IMAGEN id -u |
1001 |
- Phase 5 — Automated delivery
Objective: that a git push to main tests, builds, publishes and deploys to development without you touching anything; and that promotion to production is the same image with an approval.
Estimated time: 4-6 hours.
7.1 The development pipeline
# cloudbuild.yaml
substitutions:
_REGION: europe-west1
_SERVICIO: refugio-web
_REPO: refugio-imagenes
steps:
# 1. Tests BEFORE building. If they fail, there is no image.
- id: pruebas
name: python:3.12-slim
entrypoint: bash
args:
- -c
- |
pip install --no-cache-dir -r app/requirements.txt -r app/requirements-dev.txt
cd app && python -m pytest tests/ -v --tb=short
# 2. Build using the previous image as cache
- id: construir
name: gcr.io/cloud-builders/docker
args:
- build
- --cache-from=${_REGION}-docker.pkg.dev/$PROJECT_ID/${_REPO}/web:latest
- -t=${_REGION}-docker.pkg.dev/$PROJECT_ID/${_REPO}/web:$SHORT_SHA
- -t=${_REGION}-docker.pkg.dev/$PROJECT_ID/${_REPO}/web:latest
- app/
waitFor: [pruebas]
# 3. Publish
- id: publicar
name: gcr.io/cloud-builders/docker
args: [push, --all-tags, "${_REGION}-docker.pkg.dev/$PROJECT_ID/${_REPO}/web"]
# 4. Deploy to DEVELOPMENT
- id: desplegar
name: gcr.io/google.com/cloudsdktool/cloud-sdk:slim
entrypoint: gcloud
args:
- run
- deploy
- ${_SERVICIO}
- --image=${_REGION}-docker.pkg.dev/$PROJECT_ID/${_REPO}/web:$SHORT_SHA
- --region=${_REGION}
- --revision-suffix=$SHORT_SHA
# 5. Smoke: if the freshly deployed revision does not respond, the build fails
- id: humo
name: gcr.io/google.com/cloudsdktool/cloud-sdk:slim
entrypoint: bash
args:
- -c
- |
URL=$(gcloud run services describe ${_SERVICIO} --region=${_REGION} --format='value(status.url)')
TOKEN=$(gcloud auth print-identity-token)
CODIGO=$(curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $$TOKEN" "$$URL/salud/arranque")
echo "Code: $$CODIGO"
test "$$CODIGO" = "200"
options:
logging: CLOUD_LOGGING_ONLY
machineType: E2_HIGHCPU_8
timeout: 900sThe order is the important part: tests → build → publish → deploy → smoke. The tests go before the build, because building an image from code that does not pass the tests is time and money thrown away. And the smoke test goes after the deployment, because it is the only thing that distinguishes "the deployment finished" from "the deployment worked".
7.2 Promotion to production
# cloudbuild-prod.yaml — does NOT build. It promotes the already-tested image.
substitutions:
_IMAGEN_SHA: "" # passed explicitly: the one already working in dev
steps:
- id: verificar-imagen-existe
name: gcr.io/google.com/cloudsdktool/cloud-sdk:slim
entrypoint: bash
args:
- -c
- |
test -n "${_IMAGEN_SHA}" || { echo "Missing _IMAGEN_SHA"; exit 1; }
gcloud artifacts docker images describe \
europe-west1-docker.pkg.dev/refugio-dev/refugio-imagenes/web:${_IMAGEN_SHA}
# CANARY deployment: 10% of traffic to the new revision
- id: canario
name: gcr.io/google.com/cloudsdktool/cloud-sdk:slim
entrypoint: gcloud
args:
- run
- deploy
- refugio-web
- --image=europe-west1-docker.pkg.dev/refugio-dev/refugio-imagenes/web:${_IMAGEN_SHA}
- --region=europe-west1
- --project=refugio-prod
- --no-traffic # deployed without receiving traffic
- --revision-suffix=${_IMAGEN_SHA}
- id: repartir-10
name: gcr.io/google.com/cloudsdktool/cloud-sdk:slim
entrypoint: gcloud
args:
- run
- services
- update-traffic
- refugio-web
- --region=europe-west1
- --project=refugio-prod
- --to-revisions=refugio-web-${_IMAGEN_SHA}=10The principle to internalise: the same image, without rebuilding. Rebuilding for production means you are deploying something you have never tested: the pip install may resolve a different version, the base image may have changed. The image that passed the tests in dev is the one that goes to production, byte for byte.
Manual approval is configured on the Cloud Build trigger (--require-approval) or, in GitHub Actions, with a protected environment.
Phase 5 "done" criterion:
| Check | How | Expected |
|---|---|---|
| A push deploys on its own | git push and check gcloud builds list --limit=1 |
SUCCESS in <10 min |
| A broken test stops the deployment | Break a test on purpose, push | Build in FAILURE, Cloud Run revision unchanged |
| No keys in CI | Review the repository secrets | No GCP credential |
| Promotion without rebuilding | Compare the digest in dev and prod | Identical |
The second row is the real test. A pipeline that has never failed is untested: break a test on purpose, check that the deployment stops, and save the screenshot. It is worth 3 points and it demonstrates that the safety net exists.
- Phase 6 — The data and AI layer
Objective: that an action in the application ends up visible in the dashboard, and that there is an AI component that works.
Estimated time: 6-9 hours.
8.1 Ingestion: the application publishes events
# app/src/eventos.py
import json, os
from google.cloud import pubsub_v1
_publisher = pubsub_v1.PublisherClient()
_TOPIC = _publisher.topic_path(os.environ["GOOGLE_CLOUD_PROJECT"],
os.environ["TOPIC_EVENTOS"])
def publicar_reserva(evento: dict) -> None:
"""Publishes the business event. NEVER includes contact details."""
carga = {
"evento_id": evento["id"],
"tipo_evento": evento["tipo"], # creada | cancelada
"reserva_id": evento["reserva_id"],
"refugio_id": evento["refugio_id"],
"refugio_nombre": evento["refugio_nombre"],
"fecha_estancia": evento["fecha"],
"plazas": evento["plazas"],
"ocurrido_en": evento["ts"],
}
futuro = _publisher.publish(_TOPIC, json.dumps(carga).encode("utf-8"))
futuro.result(timeout=10)The comment is not decorative: the event carries no nombre_titular and no email_titular. It is the data minimisation from 08-02 applied in the code, and it is the difference between having personal data in one place or in four.
8.2 Transformation: the function that writes into BigQuery
# ml/../funcion/main.py
import base64, json, os
from google.cloud import bigquery
_bq = bigquery.Client()
_TABLA = os.environ["TABLA_EVENTOS"]
def procesar(evento, contexto):
"""Pub/Sub subscriber. Inserts the event into BigQuery."""
datos = json.loads(base64.b64decode(evento["data"]).decode("utf-8"))
errores = _bq.insert_rows_json(_TABLA, [datos], row_ids=[datos["evento_id"]])
if errores:
# Raising an exception makes Pub/Sub retry; after N attempts,
# the message goes to the dead-letter queue.
raise RuntimeError(f"Error inserting into BigQuery: {errores}")row_ids with the event identifier activates BigQuery deduplication: if Pub/Sub delivers the same message twice — and it will, because its guarantee is "at least once" — the row is not duplicated. It is one line that prevents a dashboard with inflated figures.
And configure the dead-letter queue in Terraform, with its topic and its subscription, or failing messages will retry forever.
8.3 The analytical tables
-- data/sql/analitica.sql
CREATE TABLE IF NOT EXISTS `refugio-datos.refugio_analitica.reservas_eventos` (
evento_id STRING NOT NULL,
ocurrido_en TIMESTAMP NOT NULL,
tipo_evento STRING NOT NULL,
reserva_id STRING NOT NULL,
refugio_id INT64 NOT NULL,
refugio_nombre STRING,
fecha_estancia DATE NOT NULL,
plazas INT64 NOT NULL
)
PARTITION BY DATE(ocurrido_en)
CLUSTER BY refugio_id
OPTIONS (partition_expiration_days = 1095, require_partition_filter = TRUE);
-- Aggregated view: this is what the dashboard consumes.
-- Having the dashboard query a view rather than the base table reduces scanning.
CREATE OR REPLACE VIEW `refugio-datos.refugio_analitica.v_ocupacion_diaria` AS
SELECT
fecha_estancia,
refugio_id,
ANY_VALUE(refugio_nombre) AS refugio,
SUM(IF(tipo_evento = 'creada', plazas, 0)) AS plazas_reservadas,
SUM(IF(tipo_evento = 'cancelada', plazas, 0)) AS plazas_canceladas,
SUM(IF(tipo_evento = 'creada', plazas, -plazas)) AS plazas_netas
FROM `refugio-datos.refugio_analitica.reservas_eventos`
WHERE DATE(ocurrido_en) >= DATE_SUB(CURRENT_DATE(), INTERVAL 730 DAY)
GROUP BY fecha_estancia, refugio_id;The WHERE DATE(ocurrido_en) >= ... in the view is not optional: with require_partition_filter = TRUE, a view without a partition filter would fail.
8.4 The AI component
The recommendation, repeated because it matters: start with a pre-trained API.
# ml/analizar_opiniones.py
from google.cloud import language_v2
_cliente = language_v2.LanguageServiceClient()
def analizar(texto: str) -> dict:
"""Sentiment of a review. FICTIONAL text, no personal data."""
documento = language_v2.Document(
content=texto,
type_=language_v2.Document.Type.PLAIN_TEXT,
language_code="es",
)
r = _cliente.analyze_sentiment(request={"document": documento})
return {
"sentimiento": round(r.document_sentiment.score, 3), # -1..1
"magnitud": round(r.document_sentiment.magnitude, 3),
}And the architecture decision that saves money: the analysis is done once, when the review is created, and the result is stored in the database. The API is not called every time somebody opens the dashboard. It is the same reasoning that took AlpinaShop to batch recommendations in DA-002: if the result does not change, it is not recomputed.
Phase 6 "done" criterion:
# End-to-end test: create a booking and see it in BigQuery
curl -s -X POST "${URL}/api/reservas" -H 'Content-Type: application/json' \
-d '{"refugio_id":1,"fecha":"2026-12-20","plazas":2,
"nombre_titular":"Fictional Test","email_titular":"[email protected]"}'
sleep 30
bq query --use_legacy_sql=false --project_id="${PROY_DATOS}" \
'SELECT evento_id, tipo_evento, refugio_id, plazas
FROM `refugio-datos.refugio_analitica.reservas_eventos`
WHERE DATE(ocurrido_en) = CURRENT_DATE()
ORDER BY ocurrido_en DESC LIMIT 5'If that row appears, you have closed milestone M3 from 08-01: the complete flow works end to end. Take a screenshot: you will need it in the presentation.
- Phase 7 — Exposure
Objective: that the system responds on a domain of your own, with valid HTTPS, and with whatever protection you decided in the design.
Estimated time: 3-5 hours, plus DNS propagation and certificate issuance time (from 15 minutes to several hours: do not leave it to the last day).
9.1 The two possible routes
| Cloud Run custom domain | Global HTTPS load balancer | |
|---|---|---|
| Cost | €0 | ~€18/month for the forwarding rule |
| TLS | Managed and automatic | Managed and automatic |
| Cloud CDN | ❌ | ✅ |
| Cloud Armor (WAF) | ❌ | ✅ |
| Several backends (Run + bucket) | ❌ | ✅ |
| Complexity | 2 resources | 7-8 resources |
If your cost limit is tight, the first option meets RNF-5 and costs nothing. Document why in an ADR, as RefugioReserva did.
9.2 The cheap route: custom domain
resource "google_cloud_run_domain_mapping" "web" {
project = var.proyecto
location = var.region
name = var.dominio # "refugioreserva.example"
metadata { namespace = var.proyecto }
spec { route_name = google_cloud_run_v2_service.web.name }
}
output "registros_dns" {
description = "Records to create at the registrar"
value = google_cloud_run_domain_mapping.web.status[0].resource_records
}You create the records that output returns at your registrar (or in Cloud DNS) and Google issues the certificate on its own.
9.3 The complete route: load balancer, CDN and WAF
Even if you end up destroying it on cost grounds, write the module and deploy it at least once: the knowledge stays, you have screenshots and you can talk about it in the presentation.
resource "google_compute_region_network_endpoint_group" "neg" {
project = var.proyecto
name = "${var.prefijo}-neg"
region = var.region
network_endpoint_type = "SERVERLESS"
cloud_run { service = google_cloud_run_v2_service.web.name }
}
resource "google_compute_backend_service" "bs" {
project = var.proyecto
name = "${var.prefijo}-bs"
protocol = "HTTPS"
load_balancing_scheme = "EXTERNAL_MANAGED"
enable_cdn = true
security_policy = google_compute_security_policy.waf.id
backend { group = google_compute_region_network_endpoint_group.neg.id }
cdn_policy {
cache_mode = "CACHE_ALL_STATIC"
default_ttl = 3600
client_ttl = 3600
negative_caching = true
}
log_config { enable = true, sample_rate = 1.0 }
}
resource "google_compute_security_policy" "waf" {
project = var.proyecto
name = "${var.prefijo}-waf"
# Rate limiting: it protects your wallet as much as your application
rule {
action = "throttle"
priority = 1000
match {
versioned_expr = "SRC_IPS_V1"
config { src_ip_ranges = ["*"] }
}
rate_limit_options {
conform_action = "allow"
exceed_action = "deny(429)"
enforce_on_key = "IP"
rate_limit_threshold { count = 100, interval_sec = 60 }
}
}
rule {
action = "allow"
priority = 2147483647
match {
versioned_expr = "SRC_IPS_V1"
config { src_ip_ranges = ["*"] }
}
description = "Default rule"
}
}The rate limiting rule deserves a note: in a project with a €12 budget, a bot making 100,000 requests can cost you the whole month's budget. The limit of 100 requests per minute per IP protects the application as much as the bill. If you do not use a load balancer, the equivalent is --max-instances, which is a hard spending cap.
Phase 7 "done" criterion:
curl -sI "https://${DOMINIO}" | head -1 # HTTP/2 200
curl -sI "https://${DOMINIO}" | grep -i strict-transport # HSTS present
echo | openssl s_client -connect "${DOMINIO}:443" -servername "${DOMINIO}" 2>/dev/null \
| openssl x509 -noout -dates -issuer # valid certificate
curl -sI "http://${DOMINIO}" | head -1 # 301 to HTTPS
- Phase 8 — Observability
Objective: finding out something is wrong before somebody tells you.
Estimated time: 4-6 hours.
10.1 The four signals dashboard
resource "google_monitoring_dashboard" "principal" {
project = var.proyecto
dashboard_json = jsonencode({
displayName = "RefugioReserva — overview"
gridLayout = { columns = 2, widgets = [
{
title = "Traffic (requests/s)"
xyChart = { dataSets = [{ timeSeriesQuery = { timeSeriesFilter = {
filter = "metric.type=\"run.googleapis.com/request_count\" resource.type=\"cloud_run_revision\""
aggregation = { alignmentPeriod = "60s", perSeriesAligner = "ALIGN_RATE" }
}}}]}
},
{
title = "Errors (5xx/s)"
xyChart = { dataSets = [{ timeSeriesQuery = { timeSeriesFilter = {
filter = "metric.type=\"run.googleapis.com/request_count\" metric.label.response_code_class=\"5xx\""
aggregation = { alignmentPeriod = "60s", perSeriesAligner = "ALIGN_RATE" }
}}}]}
},
{
title = "p95 latency (ms)"
xyChart = { dataSets = [{ timeSeriesQuery = { timeSeriesFilter = {
filter = "metric.type=\"run.googleapis.com/request_latencies\""
aggregation = { alignmentPeriod = "60s", perSeriesAligner = "ALIGN_DELTA",
crossSeriesReducer = "REDUCE_PERCENTILE_95" }
}}}]}
},
{
title = "Active instances (saturation)"
xyChart = { dataSets = [{ timeSeriesQuery = { timeSeriesFilter = {
filter = "metric.type=\"run.googleapis.com/container/instance_count\""
aggregation = { alignmentPeriod = "60s", perSeriesAligner = "ALIGN_MEAN" }
}}}]}
}
]}
})
}10.2 The alert that actually notifies
resource "google_monitoring_notification_channel" "correo" {
project = var.proyecto
display_name = "Owner's email"
type = "email"
labels = { email_address = var.correo_alertas }
}
resource "google_monitoring_alert_policy" "errores_5xx" {
project = var.proyecto
display_name = "5xx error rate > 5%"
combiner = "OR"
conditions {
display_name = "Elevated 5xx for 5 minutes"
condition_threshold {
filter = join(" ", [
"metric.type=\"run.googleapis.com/request_count\"",
"resource.type=\"cloud_run_revision\"",
"metric.label.response_code_class=\"5xx\"",
])
comparison = "COMPARISON_GT"
threshold_value = 0.05
duration = "300s"
aggregations {
alignment_period = "60s"
per_series_aligner = "ALIGN_RATE"
}
}
}
notification_channels = [google_monitoring_notification_channel.correo.id]
documentation {
content = <<-EOT
## Elevated 5xx errors
**First steps** (see `docs/runbook.md`):
1. `gcloud logging read 'severity>=ERROR' --limit=20 --freshness=15m`
2. Does it coincide with a deployment? `gcloud run revisions list --limit=5`
3. If it does: roll back with
`gcloud run services update-traffic refugio-web --to-revisions=<previous>=100`
4. Is the database reachable? `gcloud sql instances describe refugio-db`
EOT
mime_type = "text/markdown"
}
}The documentation block is what turns an alert into something useful. An alert that only says "something is wrong" wakes you up; one that says what to look at and how to roll back lets you fix it. Always write the first steps in there.
10.3 The SLO as a resource
resource "google_monitoring_slo" "disponibilidad" {
project = var.proyecto
service = google_monitoring_service.web.service_id
slo_id = "disponibilidad-api"
display_name = "99.5% of requests with no 5xx error (30 days)"
goal = 0.995
rolling_period_days = 30
request_based_sli {
good_total_ratio {
total_service_filter = "metric.type=\"run.googleapis.com/request_count\" resource.type=\"cloud_run_revision\""
bad_service_filter = "metric.type=\"run.googleapis.com/request_count\" resource.type=\"cloud_run_revision\" metric.label.response_code_class=\"5xx\""
}
}
}10.4 The uptime check
resource "google_monitoring_uptime_check_config" "web" {
project = var.proyecto
display_name = "RefugioReserva available"
timeout = "10s"
period = "300s"
http_check {
path = "/salud/arranque"
port = 443
use_ssl = true
validate_ssl = true
}
monitored_resource {
type = "uptime_url"
labels = { host = var.dominio, project_id = var.proyecto }
}
selected_regions = ["EUROPE", "USA"]
}Phase 8 "done" criterion — and here there is one you do not meet by looking, but by provoking:
| Check | How | Expected |
|---|---|---|
| Dashboard exists | gcloud monitoring dashboards list |
1 dashboard |
| Alert exists | gcloud alpha monitoring policies list |
≥1 policy |
| The alert really notifies | Provoke it (see below) | Email received |
| SLO calculating | Console → SLO | Error budget with a value |
| Uptime check | gcloud monitoring uptime list-configs |
1, in a correct state |
# Provoke the alert on purpose: deploy a revision that returns 500
# on a test endpoint, generate traffic, and wait for the email.
for i in $(seq 1 200); do curl -s -o /dev/null "${URL}/api/error-de-prueba"; done
# Wait 5-10 minutes. If the email does not arrive, the alert is useless.An alert that has never fired is not an alert: it is an intention. Provoking it is the only way to know that the notification channel works, that the threshold is reachable and that the email does not end up in spam. Save the screenshot of the email received: it is worth 2 points.
- Cross-cutting practices throughout the implementation
These six things do not belong to any one phase: they belong to all of them.
11.1 Small, frequent commits
One commit per comprehensible unit of work. Add network module with serverless connector is a commit; Various progress is not.
git add infra/modules/red/
git commit -m "Add network module: VPC, subnet, /28 connector and service peering"Concrete benefits: you can revert one thing without reverting five, the history tells the story of the project in the presentation, and git bisect is actually useful when something breaks.
11.2 terraform plan always reviewed
Never apply without having read the plan. And pay special attention to three phrases:
| In the plan | Meaning | Reaction |
|---|---|---|
will be created |
New resource | Normal |
will be updated in-place |
Change without recreating | Normal |
must be replaced |
It is destroyed and created again | STOP and understand why |
will be destroyed |
It disappears | Check that it is intentional |
must be replaced on a database means losing the data. On a bucket, losing the objects. It is almost always caused by changing a ForceNew attribute (the name, the region, a CIDR). If it appears and you were not expecting it, cancel.
terraform plan -out=plan.tfplan
terraform show -json plan.tfplan | \
jq -r '.resource_changes[] | select(.change.actions | index("delete")) | .address'
# If this returns something you were not expecting, do not apply.11.3 Do not touch anything by hand outside Terraform
The ideal rule. And reality: you are going to do it, in a hurry, on a Tuesday night.
What to do then, in order:
- Note it down immediately in
docs/diario.md. The sin is not the click; it is forgetting it. - Detect the drift:
terraform planwill tell you that reality does not match the code. - Decide: either you bring the change into the code (the usual choice) or you revert the manual change.
- If you created a resource by hand, import it instead of recreating it:
# Terraform 1.5+: import block, versionable in the code
cat >> infra/envs/dev/imports.tf <<'EOF'
import {
to = google_storage_bucket.temporal
id = "refugio-dev/refugio-temporal-a1b2"
}
EOF
terraform plan # generates the missing configurationThe free detector: the gestionado-por = terraform label from 08-02. Any resource without it was created by hand.
11.4 Documenting on the fly
Three lines per session in docs/diario.md and an ADR every time you decide something with alternatives. It costs nothing and it is the difference between a presentation with substance and one invented the night before.
11.5 Checking spending every few days
# Aliases worth having to hand
gcloud billing accounts list
# And, once the BigQuery export has been populating for a few days:
bq query --use_legacy_sql=false \
'SELECT service.description AS service, ROUND(SUM(cost),2) AS cost_eur
FROM `mi-proyecto.facturacion.gcp_billing_export_v1_XXXX`
WHERE DATE(usage_start_time) >= DATE_SUB(CURRENT_DATE(), INTERVAL 7 DAY)
GROUP BY service ORDER BY cost_eur DESC'Twice a week, thirty seconds. And destroy dev when you are not using it:
terraform -chdir=infra/envs/dev destroy -auto-approve # Friday
terraform -chdir=infra/envs/dev apply -auto-approve # SaturdayThis has a double benefit worth underlining: you save money and you are running the hardest test of RNF-1 every week. If one Friday Saturday's apply does not rebuild the environment, you have just discovered a flaw in your IaC at the cheapest possible moment.
11.6 Keeping the environments aligned
Every time you change something in prod, check that dev has the equivalent. Drift between environments turns promotion into a lottery.
- What to do when you get stuck: the five-step method
| Step | Question | Command |
|---|---|---|
| 1 | What does the error say, in full? | gcloud logging read 'severity>=ERROR' --limit=20 --freshness=1h --format=json |
| 2 | Is the API enabled? | gcloud services list --enabled | grep <api> |
| 3 | Is it a permission? | gcloud policy-troubleshoot iam <resource> --principal-email=<sa> --permission=<permission> |
| 4 | Is it the network? | gcloud network-management connectivity-tests create ... |
| 5 | Is the resource what I think it is? | gcloud <service> describe <resource> --format=yaml |
The ten errors you are going to hit, with their cause
| Symptom | Almost certain cause | Fix |
|---|---|---|
PERMISSION_DENIED when deploying |
Missing iam.serviceAccountUser on the runtime SA |
Grant it on that specific SA |
API not enabled |
Exactly that, or propagation still under way | gcloud services enable ... and wait 2 min |
| Private Cloud SQL will not be created | Missing private services access peering | Create global_address + service_networking_connection |
| Cloud Run cannot connect to the DB | Missing connector, or wrong vpc-egress |
--vpc-connector + --vpc-egress=private-ranges-only |
| Container does not start | It does not listen on $PORT or on 0.0.0.0 |
Fix the CMD |
| Deployment times out | Startup > 4 min (migrations at startup) | Take the migrations out of startup |
| Bucket "already exists" | Global name, somebody else has it | Add a suffix with random_id |
| VPC connector will not be created | The range is not an exact /28 |
Adjust the CIDR |
| Alert that never arrives | Unverified channel, or email in spam | Verify the channel and test by provoking it |
Unexpected must be replaced |
You changed a ForceNew attribute |
Check the Terraform registry before applying |
And the two-hour rule: if you have spent two hours on the same error, switch phase. Note the exact error and what you have tried in the log. Coming back the next day solves more blockages than persisting.
- Progress log and "done" criteria
Keep this table in docs/diario.md and update it as you finish each phase. When you work alone, it is the only thing that objectively tells you where you stand:
| Phase | Deliverable | Verifiable "done" criterion | Verification command | Status |
|---|---|---|---|---|
| 0 | Projects and budget | 3 billable projects, ≥25 APIs, 1 budget | gcloud billing projects describe |
⬜ |
| 0 | Repository | Structure + README + ≥2 commits |
git log --oneline |
⬜ |
| 1 | Remote state | Bucket with versioning and no public access | gcloud storage buckets describe |
⬜ |
| 1 | Network | Own VPC, connector READY, peering active |
gcloud compute networks list |
⬜ |
| 1 | Clean plan | terraform plan → No changes |
terraform plan |
⬜ |
| 2 | Identities | 0 primitive roles on SAs, 0 user keys | Phase 2 script | ⬜ |
| 2 | WIF | Authenticated build with no secrets in the CI | gcloud builds list |
⬜ |
| 3 | Database | ipv4Enabled = False, backups on |
gcloud sql instances describe |
⬜ |
| 3 | Secrets | Password in Secret Manager, not in the repo | git log -p | grep -i password → empty |
⬜ |
| 3 | Fictional data | ~2,000 rows, migrations recorded | psql -c "SELECT count(*)..." |
⬜ |
| 4 | Application | Service Ready, /salud/arranque = 200 |
curl |
⬜ |
| 4 | Logs | Entries with jsonPayload and a trace field |
gcloud logging read |
⬜ |
| 5 | CI/CD | A push deploys in <10 min | gcloud builds list --limit=1 |
⬜ |
| 5 | Test that protects | Broken test → deployment stopped | Break it on purpose | ⬜ |
| 6 | Complete flow | Booking → event → BigQuery in <2 min | bq query |
⬜ |
| 6 | Dashboard | 4 visualisations with data | Report URL | ⬜ |
| 6 | AI | Reviews with a sentiment score | psql -c "SELECT ... WHERE sentimiento IS NOT NULL" |
⬜ |
| 7 | Domain and TLS | HTTP/2 200 and a valid certificate |
curl -sI + openssl |
⬜ |
| 8 | Dashboard | 4 signals visible | gcloud monitoring dashboards list |
⬜ |
| 8 | Tested alert | Email received after provoking it | Screenshot of the email | ⬜ |
| 8 | SLO | Error budget with a numeric value | Monitoring console | ⬜ |
The three rows in bold are the ones people tick without checking. Do not: they are precisely the ones that prove the system really works.
- The realistic warning: the first time, everything takes twice as long
It deserves its own section because it is the main cause of abandonment, and because it is not a problem with you.
| Task | First time | Second time |
|---|---|---|
| Configuring WIF | 2-3 h | 15 min |
| Cloud SQL with a private IP | 2 h | 20 min |
| First Cloud Run deployment that works | 3-4 h | 30 min |
| Complete Cloud Build pipeline | 4 h | 45 min |
| Monitoring dashboard with 4 charts | 2 h | 30 min |
| Reusable Terraform module | 3 h | 45 min |
What happens the first time and not the second: reading documentation, understanding the mental model, getting a field name wrong, waiting for propagation, discovering a permission was missing, undoing and redoing.
Three practical consequences:
- Plan for double. If you think phase 5 is 4 hours, set aside 8.
- Do not measure yourself against a tutorial. The twenty-minute video is edited and recorded by somebody who has done it fifty times.
- The "wasted" time is the learning. The two hours fighting with
iam.serviceAccountUserare exactly why next time you take fifteen minutes, and why in an interview you will be able to answer instantly.
Common Mistakes and Tips
Starting with the application. It is mistake 2 from 08-01 and it shows up here. If your twentieth commit has nothing in infra/, you started in the wrong place.
Creating "just one little thing" by hand. It is never one. Three weeks later you have fifteen orphaned resources and a terraform plan full of noise that you no longer read. Note it, import it or revert it, always.
Applying without reading the plan. The day must be replaced appears on your database and you do not see it, you will lose the data and the weekend.
Granting roles/editor "temporarily". That "temporarily" lasts until the presentation. Start restrictive and open up with policy-troubleshoot.
Leaving the domain until the last day. DNS propagation and certificate issuance have timings you do not control. Do it as soon as phase 5 is stable.
Loading the AI model on every request. Analyse once and store the result. Calling an AI API on every page load is expensive and slow with no advantage whatsoever.
Not testing the alert. It is the quickest check to do and the most forgotten. An untested alert has a 50 % chance of not working when you need it.
Tip: write the "done" criterion first, the code afterwards. Before starting a phase, write the command you will use to verify it. It forces you to define what finishing means and avoids "I think it is done".
Tip: use terraform plan as a learning tool. Every time you write a new resource, run plan and read what it implies. It is the best documentation that exists on what each block does.
Tip: save the important outputs in the repository. A docs/evidencias/ directory with the output of the identity audit script, the screenshot of the first end-to-end flow, the alert email. It is worth points in 08-05 and it is impossible to reconstruct afterwards.
Tip: git commit at the end of every session even if it does not work. A commit WIP: VPC connector, peering still failing is useful information. The history is a deliverable.
Exercises
Exercise 1 — Build the reproducible foundation (phases 0 to 2)
Complete phases 0, 1 and 2 in your project: projects created with the names from the design, APIs enabled, budget with alerts, repository with structure; state bucket with versioning and public access prevention, remote backend with a prefix per environment, provider pinned by version, at least two of your own modules (network and identity) and per-environment variables with the same keys; service accounts per workload with roles scoped to the resource, and Workload Identity Federation with a repository condition.
Deliverable: the output of the identity verification script (0 primitive roles, 0 keys) and a terraform plan that says No changes.
Exercise 2 — Data, application and end-to-end flow (phases 3, 4 and 6)
Deploy your managed database with no public IP, with its password generated and stored in Secret Manager, its migrations applied as code and its fictional data seeded reproducibly. Containerise your application honouring the Cloud Run contract — $PORT, 0.0.0.0, non-root user, fast startup — with configuration through environment variables, secrets by reference, two different health probes and structured logs with trace_id. Connect the data flow through to the dashboard and add the AI component.
Deliverable: the complete trace of an action in your application that ends up visible in the dashboard, with the verification commands for each hop.
Exercise 3 — Automate, expose and observe (phases 5, 7 and 8)
Set up the pipeline that tests, builds, publishes, deploys to development and runs a smoke test, plus the promotion to production of the same image with an approval. Break a test on purpose and demonstrate with a screenshot that the deployment stops. Expose the system on your domain with valid HTTPS and a redirect from HTTP. Create the four signals dashboard, the alert with its first-steps documentation, the uptime check and the SLO.
Provoke the alert on purpose and save the email received. Close the exercise with the progress log table from section 13 completely ticked and verified.
Solutions
Solution 1 — RefugioReserva's foundation
After completing the first three phases, verification gives this:
$ terraform -chdir=infra/envs/dev plan
No changes. Your infrastructure matches the configuration.
$ gcloud projects get-iam-policy refugio-dev --format=json | \
jq -r '.bindings[] | select(.role|test("roles/(owner|editor)")) |
"\(.role): \(.members[])"'
roles/owner: user:[email protected]
# No service account. Correct.
$ for SA in $(gcloud iam service-accounts list --project=refugio-dev --format="value(email)"); do
N=$(gcloud iam service-accounts keys list --iam-account="$SA" --managed-by=user \
--format="value(name)" | wc -l)
echo "$SA -> $N user keys"
done
[email protected] -> 0
[email protected] -> 0
[email protected] -> 0
[email protected] -> 0The three real problems that came up and what they cost:
| Problem | Symptom | Cause | Time lost |
|---|---|---|---|
| VPC connector would not be created | Invalid IP CIDR range |
I had put a /27; it requires a /28 |
25 min |
| WIF authenticated but would not deploy | PERMISSION_DENIED when deploying |
Missing iam.serviceAccountUser for sa-deploy on sa-web |
1 h 40 min |
| Photos bucket rejected | already exists |
Global name taken | 10 min → random_id |
The second is the module's classic. gcloud policy-troubleshoot pointed it out in two minutes; the problem was that it took me an hour and a half to remember to use it. Noted in the log so as not to repeat it.
Solution 2 — RefugioReserva's end-to-end flow
# 1. Verify the database
$ gcloud sql instances describe refugio-db --project=refugio-dev \
--format="value(settings.ipConfiguration.ipv4Enabled, state)"
False RUNNABLE
# 2. Verify the secret
$ gcloud secrets versions list refugio-db-password --project=refugio-dev \
--format="value(name,state)"
1 ENABLED
# 3. Data seeded
$ psql -h 127.0.0.1 -U app -d reservas -c \
"SELECT count(*) AS bookings, count(DISTINCT refugio_id) AS refuges FROM reserva"
bookings | refuges
----------+---------
2000 | 12
# 4. The application responds and reads from the DB
$ curl -s "${URL}/api/refugios" | jq '.[0]'
{"id":1,"nombre":"Refugio de Cotiella","altitud_m":2100,"capacidad":40}
# 5. Create a booking (FICTIONAL data)
$ curl -s -X POST "${URL}/api/reservas" -H 'Content-Type: application/json' \
-d '{"refugio_id":3,"fecha":"2026-12-20","plazas":2,
"nombre_titular":"Fictional Test","email_titular":"[email protected]"}' | jq
{"id":"7c2e...","estado":"confirmada","plazas_restantes":20}
# 6. The structured log, with its trace
$ gcloud logging read 'resource.type="cloud_run_revision" jsonPayload.ruta="/api/reservas"' \
--limit=1 --format="value(jsonPayload.message, jsonPayload.codigo, trace)"
peticion 201 projects/refugio-dev/traces/8a1f...
# 7. The event reached BigQuery (26 seconds later)
$ bq query --use_legacy_sql=false --project_id=refugio-datos \
'SELECT evento_id, tipo_evento, refugio_id, plazas, ocurrido_en
FROM `refugio-datos.refugio_analitica.reservas_eventos`
WHERE DATE(ocurrido_en) = CURRENT_DATE() ORDER BY ocurrido_en DESC LIMIT 1'
+-----------+-------------+------------+--------+---------------------+
| evento_id | tipo_evento | refugio_id | plazas | ocurrido_en |
+-----------+-------------+------------+--------+---------------------+
| 7c2e... | creada | 3 | 2 | 2026-10-14 18:42:11 |
+-----------+-------------+------------+--------+---------------------+
# 8. And the sentiment of the reviews
$ psql -c "SELECT count(*) FILTER (WHERE sentimiento IS NOT NULL) AS analysed,
round(avg(sentimiento)::numeric,3) AS average FROM opinion"
analysed | average
----------+---------
400 | 0.412Milestone M3 closed. Eight commands, eight verified hops, 26 seconds from the action to the analytical store. Screenshot saved in docs/evidencias/.
One detail that cost time and is worth pointing out: the events were appearing duplicated in BigQuery. Cause: Pub/Sub guarantees "at least once" and the function was being retried. Fix: row_ids=[evento_id] in insert_rows_json, which activates deduplication by identifier. Without that, the dashboard was showing 12 % more bookings than there really were, and the worst part is that I would not have noticed if I had not reconciled the total against the operational database. Lesson noted: always cross-check the analytical store against the operational one.
Solution 3 — RefugioReserva's delivery, exposure and observability
The proof that the pipeline protects. The capacity control test was broken on purpose:
# app/tests/test_unitarios.py — temporary change
def test_no_sobreventa():
disponible = calcular_disponibilidad(refugio_id=1, fecha="2026-08-15")
assert disponible == 999 # ← deliberately wrong value$ git commit -am "TEST: breaking the capacity test on purpose" && git push
$ gcloud builds list --limit=1 --format="table(id, status, createTime)"
ID STATUS CREATE_TIME
9f2a-... FAILURE 2026-10-21T19:14:22
$ gcloud run revisions list --service=refugio-web --region=europe-west1 --limit=2 \
--format="table(name, active, createTime)"
NAME ACTIVE CREATE_TIME
refugio-web-a3f81c True 2026-10-21T18:02:11 ← the previous one still servingThe deployment stopped at step 1. The previous revision carried on serving traffic. The safety net exists and is tested. Screenshot saved.
Exposure:
$ curl -sI https://dev.refugioreserva.example | head -3
HTTP/2 200
strict-transport-security: max-age=31536000; includeSubDomains
content-type: text/html; charset=utf-8
$ echo | openssl s_client -connect dev.refugioreserva.example:443 \
-servername dev.refugioreserva.example 2>/dev/null | \
openssl x509 -noout -dates -issuer
notBefore=Oct 20 09:14:00 2026 GMT
notAfter=Jan 18 09:13:59 2027 GMT
issuer=C = US, O = Google Trust Services, CN = WE1
$ curl -sI http://dev.refugioreserva.example | head -1
HTTP/1.1 301 Moved PermanentlyAfter ADR-006, Cloud Run's custom domain is used: managed certificate, zero cost, RNF-5 met. The load balancer module was left written and will be deployed for a week in 08-04 for the load tests.
The alert, genuinely provoked:
$ for i in $(seq 1 300); do curl -s -o /dev/null "${URL}/api/error-de-prueba"; done
$ # 6 minutes later:
$ gcloud alpha monitoring policies list --format="value(displayName,enabled)"
5xx error rate > 5% TrueEmail received at 20:41, six minutes after starting to generate errors. It contained the four first steps from the documentation block. Screenshot saved in docs/evidencias/alerta-recibida.png.
One detail that did not work first time: the first notification channel had been created but not verified, and the alerts were not arriving. There is no visible error: the policy shows as active and the incident opens, but the email never goes out. It was only discovered because the alert was provoked on purpose. That is exactly why provoking it is mandatory.
Final progress log:
| Phase | Criterion | Status | Evidence |
|---|---|---|---|
| 0 | 3 projects, 27 APIs, €12 budget | ✅ | docs/evidencias/fase0.txt |
| 1 | Clean plan, own VPC, peering |
✅ | docs/evidencias/plan-limpio.txt |
| 2 | 0 primitive roles, 0 keys, WIF | ✅ | docs/evidencias/identidades.txt |
| 3 | Private DB, secret, 2,000 rows | ✅ | docs/evidencias/datos.txt |
| 4 | Service Ready, logs with trace |
✅ | docs/evidencias/app.txt |
| 5 | Push deploys; a broken test stops it | ✅ | docs/evidencias/build-fallido.png |
| 6 | Booking → BigQuery in 26 s; 400 reviews analysed | ✅ | docs/evidencias/e2e.txt |
| 7 | HTTP/2 200, HSTS, valid certificate, 301 | ✅ | docs/evidencias/tls.txt |
| 8 | Dashboard, alert received, SLO, uptime | ✅ | docs/evidencias/alerta-recibida.png |
Actual time invested: 63 hours against the 50 estimated. The overrun was concentrated in three places: WIF (2.5 h against the 1 estimated), the Cloud SQL peering (2 h against 0.5) and debugging the duplicated events (3 h not foreseen). It matches the warning in section 14 almost exactly.
Conclusion
Your system exists, works and rebuilds itself from scratch.
You know why the build order is what it is: from the irreversible to the reversible, from what depends on nothing to what depends on everything. The projectId in phase 0 and the dashboard in phase 8, because one can never be changed and the other is rebuilt in ten minutes. And you know each specific dependency: why identity comes before data, why the application is not automated until you have deployed it by hand once, why observability is built on the complete system.
You have phase 0 with the projects created under their definitive names, the APIs enabled in one go, the budget existing before you could spend, and the repository with its structure and its README written at the start and not at the end.
You have the foundation with Terraform: the bootstrap that solves the chicken-and-egg problem, the state bucket with versioning and no public access, the backend with a prefix per environment, the provider pinned by version so that an update on some random Tuesday does not propose destroying half your infrastructure, your own modules reused in dev and prod, variables with exactly the same keys in both, and the two private services access peering resources everybody forgets.
You have identity done properly: one service account per workload, roles scoped to the resource and not to the project — google_secret_manager_secret_iam_member, not google_project_iam_member — Workload Identity Federation with its attribute_condition, without which it accepts tokens from any repository on the planet, and iam.serviceAccountUser accounted for so that the first automated deployment does not steal an afternoon from you.
You have the data with the database with no public IP, the password generated by Terraform and stored in Secret Manager without anybody ever seeing it, buckets with enforced public access prevention and a lifecycle, idempotent migrations versioned in the repository and a reproducible fictional data generator with a fixed seed.
You have the application honouring the Cloud Run contract — $PORT, 0.0.0.0, fast startup, no state — running as non-root, with configuration through variables and secrets by reference, with two probes that deliberately do different things, with cpu_idle so as not to pay between requests, and with structured logs whose logging.googleapis.com/trace key links every line to its trace.
You have automated delivery in the right order — tests before building, smoke after deploying — and the principle that makes it reliable: the same image is promoted to production, without rebuilding, because rebuilding is deploying something you have never tested. And you have verified it in the only way that counts: by breaking a test on purpose and checking that the deployment stops.
You have the data and AI layer with events that carry no personal data, deduplication via row_ids that prevents inflated figures, partitioned tables with a mandatory filter and an AI component based on a pre-trained API that analyses once and stores the result.
You have exposure with your own domain and valid TLS, knowing how to choose between the zero-cost route and the complete one with CDN and WAF, and with rate limiting understood as what it also is: budget protection.
And you have observability with the four signals dashboard, the alert with its first steps written in the documentation block itself, the uptime check, the SLO calculating its error budget, and the alert tested by provoking it, which is the only way to discover an unverified notification channel before you need it.
On top of all that you have the cross-cutting practices: small commits, the plan always read — with must be replaced as the alarm phrase — nothing created by hand and what to do when you do it anyway, documentation on the fly, spending checked twice a week with the Friday destroy that saves money and tests your IaC at the same time, and the environments kept aligned.
And you have the five-step method for when you get stuck, the ten most frequent errors with their cause, the two-hour rule, the progress log table with verifiable criteria and the most honest warning of all: the first time, everything takes twice as long, and that time is not wasted, it is exactly the learning.
In the next lesson, 08-04, the question is no longer "does it work?" but "how do I know?". You are going to build your project's testing pyramid, test the infrastructure by recreating it from scratch — the definitive proof that your IaC is real — run through a security checklist, do an honest and cheap load test, switch off a dependency on purpose to see what happens, time a restore, rehearse a rollback, and go through your pre-launch checklist before declaring the project live.
Google Cloud Platform (GCP) Course
Module 1: Introduction to Google Cloud Platform
- What is Google Cloud Platform?
- Setting Up Your GCP Account
- A Tour of the GCP Console
- Projects, Resource Hierarchy and Billing
- Regions, Zones and the Shared Responsibility Model
- Cloud Shell and the gcloud CLI
Module 2: Core GCP Services
- Compute Engine: Virtual Machines on Google Cloud
- Cloud Storage: Object Storage
- Cloud SQL: Managed Relational Databases
- App Engine: Platform as a Service
- Google Kubernetes Engine (GKE)
- NoSQL Databases: Firestore, Bigtable and Spanner
- How to Choose the Right Compute Service
Module 3: Networking and Security
- VPC Networks
- Cloud Load Balancing
- Cloud CDN
- Identity and Access Management (IAM)
- Cloud Armor
- Secrets and Encryption: Secret Manager and Cloud KMS
- Cloud DNS, TLS Certificates and Publishing Services Securely
Module 4: Data and Analytics
- BigQuery: The Analytical Data Warehouse
- Cloud Dataflow: Batch and Streaming Data Processing
- Cloud Dataproc: Managed Spark and Hadoop
- Cloud Pub/Sub: Asynchronous Messaging
- Cloud Data Fusion: Code-Free Data Integration
- Orchestrating Pipelines with Cloud Composer and Workflows
- Data Governance and Dashboards with Dataplex and Looker Studio
Module 5: Machine Learning and AI
- Vertex AI: The Machine Learning Platform on GCP
- AutoML: Custom Models Without Writing Code
- TensorFlow on GCP: Training and Serving Models
- Natural Language API
- Vision API
- Generative AI on Vertex AI: Gemini Models and Embeddings
- MLOps: From Model to Product with Vertex AI Pipelines
Module 6: DevOps and Monitoring
- Cloud Build: Continuous Integration on GCP
- Cloud Source Repositories and Source Code Management
- Cloud Functions: Serverless Functions
- Cloud Monitoring (formerly Stackdriver): Metrics, Dashboards and Alerts
- Cloud Deployment Manager and Native Infrastructure as Code
- Cloud Logging and Cloud Trace: Logs, Traces and Diagnostics
- Terraform on GCP: Infrastructure as Code in Practice
Module 7: Advanced GCP Topics
- Hybrid and Multicloud with Anthos
- Serverless Computing with Cloud Run
- Advanced Networking: Shared VPC, Peering and Hybrid Connectivity
- Security Best Practices
- Cost Management and Optimization
- Reliability: SLOs, High Availability and Disaster Recovery
- Governance at Scale: Organization, Policies and Auditing
