You have the blueprint. Now it is time to build.

And this is where most personal projects go wrong, for a very specific reason: building does not fail because of technical difficulty, it fails because of order. Whoever starts with the application ends up with a pretty app hanging off hand-crafted infrastructure that cannot be reproduced. Whoever does security backwards — opening everything up "so that it works" and promising to lock it down later — never locks it down. Whoever does not control spending discovers in week three that the entire credit has been eaten.

This lesson is a working manual. It does not explain what Cloud Run is or what Terraform is: you already saw that in 07-02 and in 06-07. It explains what order the pieces go in, why that order and not another, and how you check after each step that what you have done is right.

It is organised into nine phases, from 0 to 8. Each phase has its objective, its steps, its commands and — most importantly when you work alone — its verifiable "done" criterion: a command you run and an output you expect. If the command does not give what is expected, the phase is not finished, however much it "looks like" it works.

By the end of the lesson you will have your system up and running, reproducible from scratch, with automated delivery, secure and observed. And you will know exactly where you stand at any moment, which when you work alone is half the battle.

Contents

  1. The build order and why it is that one
  2. Phase 0 — Preparation: projects, APIs, budget and repository
  3. Phase 1 — The foundation with Terraform: state, provider, modules and network
  4. Phase 2 — Identity: service accounts, minimum roles and keyless federation
  5. Phase 3 — Data: private database, buckets and migrations as code
  6. Phase 4 — The application: container, configuration, secrets, health and logs
  7. Phase 5 — Automated delivery: tests, build, deployment and promotion
  8. Phase 6 — The data and AI layer
  9. Phase 7 — Exposure: domain, TLS and protection
  10. Phase 8 — Observability: logs, dashboard, alerts and SLO
  11. Cross-cutting practices throughout the implementation
  12. What to do when you get stuck: the five-step method
  13. Progress log and "done" criteria
  14. The realistic warning: the first time, everything takes twice as long

  1. The build order and why it is that one

flowchart TD
    F0["Phase 0 — Preparation<br/>projects · APIs · budget · repo"]
    F1["Phase 1 — Foundation<br/>Terraform state · network · firewall"]
    F2["Phase 2 — Identity<br/>service accounts · roles · WIF"]
    F3["Phase 3 — Data<br/>private DB · buckets · migrations"]
    F4["Phase 4 — Application<br/>container · secrets · health · logs"]
    F5["Phase 5 — Delivery<br/>CI/CD · tests · promotion"]
    F6["Phase 6 — Data and AI<br/>events · BigQuery · dashboard · model"]
    F7["Phase 7 — Exposure<br/>domain · TLS · protection"]
    F8["Phase 8 — Observability<br/>dashboard · alerts · SLO"]

    F0 --> F1 --> F2 --> F3 --> F4 --> F5
    F4 --> F6
    F5 --> F7
    F6 --> F8
    F7 --> F8

    style F0 fill:#e8f0fe
    style F4 fill:#fef7e0
    style F8 fill:#e6f4ea

The justification for each dependency, which is what makes the order rational rather than habitual:

Phase Depends on Why exactly
0. Preparation Nothing The projectId is irreversible and APIs take minutes to propagate. And the budget has to exist before you can spend
1. Foundation 0 Terraform needs a project with the Resource Manager API enabled and a bucket for the state. The network cannot be reconfigured without destroying what is inside it
2. Identity 1 Service accounts belong to a project, but permissions on specific resources (bucket, secret, topic) require those resources to exist… or for you to create them in the same apply. It comes before data because the DB is created with its user and its secret already
3. Data 1, 2 Cloud SQL with a private IP requires the network and the private services access peering to already exist. And its password goes to Secret Manager, which needs the identity that will read it
4. Application 3 The app does not start without a DB or secrets. Deploying it earlier forces you to deploy it twice
5. Delivery 4 You cannot automate a deployment you have never done by hand. First you do it once and understand it; then you automate it
6. Data and AI 4 The events are emitted by the application. No application, no events to process
7. Exposure 5 The domain points at a stable service. If the app is still changing shape every day, DNS propagation and the certificate are noise
8. Observability 6, 7 You observe the complete system. A dashboard built on half a system has to be redone

The general rule behind all of this: you build from the irreversible to the reversible, and from what depends on nothing to what depends on everything. The projectId cannot be changed; a Monitoring dashboard is rebuilt in ten minutes. That is why the first is in phase 0 and the second in phase 8.

The legitimate exception: if at any point you are stuck for more than an hour in a phase, jump to the next one that does not depend on it and come back afterwards. Phase 6 (data and AI) and phase 7 (exposure) are independent of each other; phase 8 depends on both. Document the jump in the log.

  1. Phase 0 — Preparation

Objective: get the ground ready for Terraform to work, with spending under control from the very first minute.

Estimated time: 1-2 hours.

2.1 Create the projects with the names decided in 08-02

# Project variables — adjust them to what you decided in the design
export PROY_BASE="refugio"
export PROY_DEV="${PROY_BASE}-dev"
export PROY_PROD="${PROY_BASE}-prod"
export PROY_DATOS="${PROY_BASE}-datos"
export REGION="europe-west1"
export BILLING_ACCOUNT="0X0X0X-0X0X0X-0X0X0X"   # gcloud billing accounts list

# Create the projects
for P in "${PROY_DEV}" "${PROY_PROD}" "${PROY_DATOS}"; do
  gcloud projects create "${P}" --name="${P}"
  gcloud billing projects link "${P}" --billing-account="${BILLING_ACCOUNT}"
done

# Check
gcloud projects list --filter="projectId:${PROY_BASE}-*" \
  --format="table(projectId, name, lifecycleState)"

⚠️ Before pressing Enter, read the names out loud. It is the last chance. A projectId cannot be changed, cannot be reused even after deleting it, and will appear in every screenshot of your presentation.

If gcloud projects create fails with already exists, it is not that you have it: it is that somebody in the world has it, because the namespace is global. Add a short, distinctive suffix, not a -2.

2.2 Enable the APIs

APIs take from seconds to minutes to propagate. Enable them all in one go now and you save yourself ten interruptions later:

APIS=(
  cloudresourcemanager.googleapis.com   # Terraform needs this one first
  serviceusage.googleapis.com
  iam.googleapis.com
  iamcredentials.googleapis.com
  compute.googleapis.com                # network, addresses, firewall
  servicenetworking.googleapis.com      # peering for private Cloud SQL
  vpcaccess.googleapis.com              # serverless connector
  run.googleapis.com
  artifactregistry.googleapis.com
  cloudbuild.googleapis.com
  sqladmin.googleapis.com
  secretmanager.googleapis.com
  storage.googleapis.com
  pubsub.googleapis.com
  bigquery.googleapis.com
  cloudfunctions.googleapis.com
  eventarc.googleapis.com
  cloudscheduler.googleapis.com
  monitoring.googleapis.com
  logging.googleapis.com
  cloudtrace.googleapis.com
  language.googleapis.com               # replace with your AI API
)

for P in "${PROY_DEV}" "${PROY_PROD}"; do
  gcloud services enable "${APIS[@]}" --project="${P}"
done

gcloud services enable bigquery.googleapis.com storage.googleapis.com \
  --project="${PROY_DATOS}"

"Done" criterion:

gcloud services list --enabled --project="${PROY_PROD}" | wc -l   # should be around 25+

2.3 The budget, before anything else

You already created it in 08-01. If not, do it now, before creating the first billable resource:

gcloud billing budgets create \
  --billing-account="${BILLING_ACCOUNT}" \
  --display-name="Final project ${PROY_BASE}" \
  --budget-amount=12EUR \
  --threshold-rule=percent=0.5 \
  --threshold-rule=percent=0.9 \
  --threshold-rule=percent=1.0 \
  --filter-projects="projects/${PROY_DEV}","projects/${PROY_PROD}","projects/${PROY_DATOS}"

gcloud billing budgets list --billing-account="${BILLING_ACCOUNT}"

And enable the billing export to BigQuery from the console (Billing → Billing export). It takes up to 24 hours to start populating data, so the sooner it is enabled, the sooner you will have history for the cost section of your presentation.

2.4 The repository and its structure

If you did exercise 3 of 08-01, you already have it. Extended for the implementation:

mi-proyecto/
├── README.md
├── .gitignore
├── cloudbuild.yaml                 # dev pipeline
├── cloudbuild-prod.yaml            # promotion to prod
├── app/
│   ├── Dockerfile
│   ├── requirements.txt
│   ├── src/
│   │   ├── main.py
│   │   ├── db.py
│   │   ├── eventos.py
│   │   └── observabilidad.py
│   └── tests/
│       ├── test_unitarios.py
│       └── test_integracion.py
├── infra/
│   ├── modules/
│   │   ├── red/
│   │   ├── identidad/
│   │   ├── datos/
│   │   ├── servicio/
│   │   └── observabilidad/
│   ├── envs/
│   │   ├── dev/{main.tf,variables.tf,terraform.tfvars,backend.tf}
│   │   └── prod/{main.tf,variables.tf,terraform.tfvars,backend.tf}
│   └── bootstrap/                  # creates the state bucket. Applied once
├── data/
│   ├── seed/generar.py
│   ├── migraciones/
│   │   ├── 001_esquema_inicial.sql
│   │   └── 002_indices.sql
│   └── sql/analitica.sql
├── ml/
│   └── analizar_opiniones.py
└── docs/
    ├── arquitectura.md
    ├── runbook.md
    ├── diario.md
    └── adr/

The README that actually gets read. Write it now, not at the end, and with this structure:

# RefugioReserva

Mountain refuge booking platform. Final project of the GCP course.
**All data is fictional.**

🔗 **Demo:** https://refugioreserva.example
📊 **Dashboard:** [Looker Studio](...)
📐 **Architecture:** [docs/arquitectura.md](docs/arquitectura.md)

## What it does
A hiker checks the availability of places in 12 refuges and books.
The warden sees the day's bookings. The federation checks indicators.

## Architecture in one line
Cloud Run (Python/FastAPI) → private Cloud SQL PostgreSQL · photos in Cloud
Storage · events over Pub/Sub → BigQuery → Looker Studio · sentiment of
reviews with the Natural Language API. All in Terraform, deployed by
Cloud Build with no keys (WIF).

## How to bring it up from scratch

cd infra/bootstrap && terraform init && terraform apply cd ../envs/dev && terraform init && terraform apply make seed

## Cost
**~€10/month.** See [docs/adr/ADR-006-sin-balanceador.md](docs/adr/) for
the most relevant cost decision.

## Status and known technical debt
See the "Technical debt" section below. Yes, there is some, and it is prioritised.

Phase 0 "done" criterion:

Check Command Expected result
Projects created and billable gcloud billing projects describe ${PROY_PROD} billingEnabled: true
APIs enabled gcloud services list --enabled --project=${PROY_PROD} ≥25 lines
Budget created gcloud billing budgets list --billing-account=${BILLING_ACCOUNT} 1 budget
Repository with structure and README git log --oneline ≥2 commits

  1. Phase 1 — The foundation with Terraform

Objective: that from here on everything is created with code.

Estimated time: 4-8 hours the first time.

3.1 The state bucket (bootstrap)

There is a chicken-and-egg problem: Terraform stores its state in a bucket, but the bucket also has to be created. The standard solution is a small bootstrap module with local state that is applied just once:

# infra/bootstrap/main.tf
terraform {
  required_version = ">= 1.9"
  required_providers {
    google = { source = "hashicorp/google", version = "~> 6.0" }
  }
  # No backend: local state, applied once and the .tfstate is uploaded encrypted
  # or you simply accept that this module gets re-created by hand if it is lost.
}

provider "google" {
  project = var.proyecto_estado
  region  = var.region
}

resource "google_storage_bucket" "estado" {
  name          = "${var.prefijo}-terraform-estado"
  location      = var.region
  force_destroy = false

  uniform_bucket_level_access = true
  public_access_prevention    = "enforced"

  versioning { enabled = true }        # ← essential: lets you recover a corrupted state

  lifecycle_rule {
    condition { num_newer_versions = 20 }
    action    { type = "Delete" }
  }

  labels = {
    proyecto       = var.prefijo
    componente     = "iac"
    gestionado-por = "terraform"
  }
}

The three options that are not negotiable:

  • versioning.enabled = true: if the state gets corrupted (and it happens), the previous version saves your project.
  • public_access_prevention = "enforced": the Terraform state contains resource identifiers and, sometimes, sensitive values. Never public.
  • force_destroy = false: stops an accidental destroy taking the state of everything else with it.
cd infra/bootstrap
terraform init && terraform apply
gcloud storage buckets describe gs://refugio-terraform-estado \
  --format="value(versioning.enabled,iamConfiguration.publicAccessPrevention)"
# Expected: True  enforced

3.2 The remote backend and the pinned provider

# infra/envs/prod/backend.tf
terraform {
  required_version = ">= 1.9"

  backend "gcs" {
    bucket = "refugio-terraform-estado"
    prefix = "envs/prod"                  # dev uses "envs/dev": separate states
  }

  required_providers {
    google = {
      source  = "hashicorp/google"
      version = "~> 6.0"                  # pinned. NEVER without a version
    }
    random = { source = "hashicorp/random", version = "~> 3.6" }
  }
}

Why a different prefix per environment and not two buckets: one bucket, two prefixes, two completely independent states. It is simpler to manage and there is no risk of a dev apply touching prod, because they are different state files.

Why pin the provider version: without version, Terraform picks the latest one every time you run init. On some random Tuesday version 7.0 comes out with breaking changes and your plan proposes destroying half your infrastructure. With ~> 6.0 you stay on the 6 branch until you decide to move up yourself, consciously and with a commit that says so.

3.3 Your own modules

The practical rule: create a module when you are going to use the same thing in two environments. With dev and prod, that is nearly everything.

# infra/modules/red/main.tf
variable "proyecto"    { type = string }
variable "prefijo"     { type = string }
variable "region"      { type = string }
variable "cidr_app"    { type = string }
variable "cidr_conector" { type = string }
variable "cidr_privado"  { type = string }
variable "etiquetas"   { type = map(string) }

resource "google_compute_network" "vpc" {
  project                 = var.proyecto
  name                    = "${var.prefijo}-vpc"
  auto_create_subnetworks = false        # ← NEVER the default network
  routing_mode            = "REGIONAL"
}

resource "google_compute_subnetwork" "app" {
  project                  = var.proyecto
  name                     = "${var.prefijo}-app-${substr(var.region, 0, 8)}"
  network                  = google_compute_network.vpc.id
  region                   = var.region
  ip_cidr_range            = var.cidr_app
  private_ip_google_access = true        # egress to Google APIs without a public IP
}

# Serverless access connector: requires an exact /28
resource "google_vpc_access_connector" "conector" {
  project       = var.proyecto
  name          = "${var.prefijo}-conn"
  region        = var.region
  ip_cidr_range = var.cidr_conector
  network       = google_compute_network.vpc.name
  min_instances = 2
  max_instances = 3                      # cost cap
}

# --- Private services access (needed for Cloud SQL with a private IP) ---
resource "google_compute_global_address" "rango_privado" {
  project       = var.proyecto
  name          = "${var.prefijo}-rango-privado"
  purpose       = "VPC_PEERING"
  address_type  = "INTERNAL"
  prefix_length = 20
  address       = split("/", var.cidr_privado)[0]
  network       = google_compute_network.vpc.id
}

resource "google_service_networking_connection" "peering" {
  network                 = google_compute_network.vpc.id
  service                 = "servicenetworking.googleapis.com"
  reserved_peering_ranges = [google_compute_global_address.rango_privado.name]
}

# --- Firewall: explicit deny by default ---
resource "google_compute_firewall" "denegar_entrada" {
  project   = var.proyecto
  name      = "${var.prefijo}-denegar-entrada"
  network   = google_compute_network.vpc.name
  direction = "INGRESS"
  priority  = 65534
  deny { protocol = "all" }
  source_ranges = ["0.0.0.0/0"]
  log_config { metadata = "INCLUDE_ALL_METADATA" }
}

output "red_id"        { value = google_compute_network.vpc.id }
output "red_nombre"    { value = google_compute_network.vpc.name }
output "conector_id"   { value = google_vpc_access_connector.conector.id }
output "peering_listo" { value = google_service_networking_connection.peering.id }

The two resources people forget are in there: google_compute_global_address with purpose = "VPC_PEERING" and google_service_networking_connection. Without them, Cloud SQL with a private IP fails with an error that mentions peering nowhere at all.

And the output "peering_listo" is not decorative: it lets the data module declare a depends_on against it so that Terraform does not try to create the database before the peering exists.

3.4 Variables per environment

# infra/envs/prod/terraform.tfvars
proyecto      = "refugio-prod"
entorno       = "prod"
prefijo       = "refugio"
region        = "europe-west1"
cidr_app      = "10.20.0.0/24"
cidr_conector = "10.20.8.0/28"
cidr_privado  = "10.20.16.0/20"

bd_tier            = "db-f1-micro"
bd_alta_disp       = false
bd_backup_retencion = 7

run_min_instancias = 0
run_max_instancias = 5
# infra/envs/dev/terraform.tfvars — same keys, smaller values
proyecto      = "refugio-dev"
entorno       = "dev"
prefijo       = "refugio"
region        = "europe-west1"
cidr_app      = "10.10.0.0/24"
cidr_conector = "10.10.8.0/28"
cidr_privado  = "10.10.16.0/20"

bd_tier            = "db-f1-micro"
bd_alta_disp       = false
bd_backup_retencion = 1
run_min_instancias = 0
run_max_instancias = 2

The rule: both files have exactly the same keys. If dev has a variable prod does not have, the environments have diverged and promotion stops being reliable.

3.5 The first apply

cd infra/envs/dev
terraform init
terraform fmt -recursive ../..     # consistent formatting, free
terraform validate                 # syntax and references
terraform plan -out=plan.tfplan    # READ THE WHOLE OUTPUT
terraform apply plan.tfplan

Read the whole plan. Always. It is cross-cutting practice number one from section 11, and the first time is when you learn the most: the plan shows you exactly which resources each block you have written implies.

Phase 1 "done" criterion:

Check Command Expected
Remote state gcloud storage ls gs://refugio-terraform-estado/envs/dev/ default.tfstate
Network created, not the default one gcloud compute networks list --project=${PROY_DEV} refugio-vpc, no default
Connector active gcloud compute networks vpc-access connectors list --region=${REGION} --project=${PROY_DEV} state READY
Peering established gcloud services vpc-peerings list --network=refugio-vpc --project=${PROY_DEV} 1 peering
Clean plan terraform plan No changes.

That last row is the one that matters: a plan that says No changes means the code and reality match. It is the criterion checked in 08-04 and it is worth 3 points of the rubric.

  1. Phase 2 — Identity

Objective: every workload with its own identity and its minimum permissions, and CI/CD working without a single downloaded key.

Estimated time: 3-5 hours (2 of them the first time you configure WIF).

4.1 Service accounts and minimum roles

# infra/modules/identidad/main.tf

resource "google_service_account" "web" {
  project      = var.proyecto
  account_id   = "sa-${var.prefijo}-web"
  display_name = "Web application account"
  description  = "Cloud Run: reads the DB, reads 2 secrets, writes photos, publishes events"
}

# --- PROJECT-level roles: only those that cannot be narrowed further ---
resource "google_project_iam_member" "web_proyecto" {
  for_each = toset([
    "roles/cloudsql.client",
    "roles/logging.logWriter",
    "roles/cloudtrace.agent",
    "roles/monitoring.metricWriter",
  ])
  project = var.proyecto
  role    = each.value
  member  = "serviceAccount:${google_service_account.web.email}"
}

# --- Roles scoped TO THE RESOURCE: this is how it is done properly ---
resource "google_secret_manager_secret_iam_member" "web_secretos" {
  for_each  = toset(var.secretos_de_la_web)   # ["refugio-db-password", "refugio-session-key"]
  project   = var.proyecto
  secret_id = each.value
  role      = "roles/secretmanager.secretAccessor"
  member    = "serviceAccount:${google_service_account.web.email}"
}

resource "google_storage_bucket_iam_member" "web_fotos" {
  bucket = var.bucket_fotos
  role   = "roles/storage.objectAdmin"
  member = "serviceAccount:${google_service_account.web.email}"
}

resource "google_pubsub_topic_iam_member" "web_publica" {
  project = var.proyecto
  topic   = var.topic_eventos
  role    = "roles/pubsub.publisher"
  member  = "serviceAccount:${google_service_account.web.email}"
}

The difference that separates a good project from a mediocre one is in the names of those resources. google_project_iam_member with secretAccessor gives access to all the secrets in the project, present and future. google_secret_manager_secret_iam_member gives it to that secret. It is the same amount of code and an enormous difference in attack surface.

4.2 Workload Identity Federation: keyless CI/CD

This is the point that will set you apart most. The idea, in two sentences: instead of downloading a JSON key and storing it in GitHub, you establish a trust relationship between GCP and your CI's identity provider. The CI presents a signed token saying "I am the main branch of the repo usuario/refugioreserva", and GCP exchanges it for temporary credentials.

Zero keys. Zero rotation. Zero risk of leakage.

# infra/modules/identidad/wif.tf

resource "google_iam_workload_identity_pool" "github" {
  project                   = var.proyecto
  workload_identity_pool_id = "gh-pool"
  display_name              = "GitHub Actions"
}

resource "google_iam_workload_identity_pool_provider" "github" {
  project                            = var.proyecto
  workload_identity_pool_id          = google_iam_workload_identity_pool.github.workload_identity_pool_id
  workload_identity_pool_provider_id = "gh-provider"

  attribute_mapping = {
    "google.subject"       = "assertion.sub"
    "attribute.repository" = "assertion.repository"
    "attribute.ref"        = "assertion.ref"
  }

  # MANDATORY CONDITION: without this, ANY GitHub repository
  # in the world could authenticate against your project.
  attribute_condition = "assertion.repository == '${var.github_repo}'"

  oidc { issuer_uri = "https://token.actions.githubusercontent.com" }
}

resource "google_service_account" "deploy" {
  project      = var.proyecto
  account_id   = "sa-${var.prefijo}-deploy"
  display_name = "Deployment from CI"
}

# Only the repository's main branch can impersonate the deployment account
resource "google_service_account_iam_member" "deploy_wif" {
  service_account_id = google_service_account.deploy.name
  role               = "roles/iam.workloadIdentityUser"
  member = "principalSet://iam.googleapis.com/${google_iam_workload_identity_pool.github.name}/attribute.repository/${var.github_repo}"
}

resource "google_project_iam_member" "deploy_roles" {
  for_each = toset([
    "roles/run.developer",
    "roles/artifactregistry.writer",
  ])
  project = var.proyecto
  role    = each.value
  member  = "serviceAccount:${google_service_account.deploy.email}"
}

# THE PERMISSION EVERYBODY ALWAYS FORGETS:
# to deploy a service that runs as sa-web, deploy must be able to "act as" it.
resource "google_service_account_iam_member" "deploy_actua_como_web" {
  service_account_id = google_service_account.web.name
  role               = "roles/iam.serviceAccountUser"
  member             = "serviceAccount:${google_service_account.deploy.email}"
}

⚠️ The attribute_condition is not optional. Without it, the provider accepts tokens from any GitHub repository on the planet, and anybody who knows your pool's name can deploy into your project. It is a real, documented security flaw that appears in many tutorials by omission.

And in the GitHub Actions workflow:

# .github/workflows/deploy.yml
permissions:
  contents: read
  id-token: write            # essential for GitHub to issue the OIDC token

jobs:
  desplegar:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: google-github-actions/auth@v2
        with:
          workload_identity_provider: projects/123456789/locations/global/workloadIdentityPools/gh-pool/providers/gh-provider
          service_account: [email protected]
      # From here on, gcloud is authenticated. With no secret in GitHub at all.

Phase 2 "done" criterion:

# 1. No service account with a primitive role
gcloud projects get-iam-policy "${PROY_PROD}" --format=json | \
  jq -r '.bindings[] | select(.role | test("roles/(owner|editor)")) |
         .members[] | select(startswith("serviceAccount:"))'
# Expected: empty

# 2. Zero user-managed service account keys
for SA in $(gcloud iam service-accounts list --project="${PROY_PROD}" --format="value(email)"); do
  echo -n "$SA: "
  gcloud iam service-accounts keys list --iam-account="$SA" \
    --managed-by=user --format="value(name)" | wc -l
done
# Expected: 0 for all of them

Save the output of those two commands: they are direct evidence for 6 points of block D of the rubric.

  1. Phase 3 — Data

Objective: a managed, private database, with its password in Secret Manager, and the schema as versioned code.

Estimated time: 3-5 hours.

5.1 The database, with no public IP

# infra/modules/datos/sql.tf

resource "random_password" "bd" {
  length  = 32
  special = true
}

resource "google_secret_manager_secret" "bd_password" {
  project   = var.proyecto
  secret_id = "${var.prefijo}-db-password"
  replication { auto {} }
}

resource "google_secret_manager_secret_version" "bd_password" {
  secret      = google_secret_manager_secret.bd_password.id
  secret_data = random_password.bd.result
}

resource "google_sql_database_instance" "principal" {
  project          = var.proyecto
  name             = "${var.prefijo}-db"
  region           = var.region
  database_version = "POSTGRES_16"

  # The peering must exist first. A legitimate case for an explicit depends_on.
  depends_on = [var.peering_listo]

  settings {
    tier              = var.bd_tier
    availability_type = var.bd_alta_disp ? "REGIONAL" : "ZONAL"
    disk_size         = 10
    disk_type         = "PD_HDD"      # cheaper; enough at this volume
    disk_autoresize   = true

    ip_configuration {
      ipv4_enabled    = false          # ← NO PUBLIC IP
      private_network = var.red_id
      ssl_mode        = "ENCRYPTED_ONLY"
    }

    backup_configuration {
      enabled                        = true
      start_time                     = "03:00"
      point_in_time_recovery_enabled = var.entorno == "prod"
      backup_retention_settings { retained_backups = var.bd_backup_retencion }
    }

    maintenance_window { day = 7, hour = 4 }   # early Sunday morning

    database_flags {
      name  = "cloudsql.iam_authentication"
      value = "on"
    }

    user_labels = var.etiquetas
  }

  # Protection against an accidental destroy in production
  deletion_protection = var.entorno == "prod"
}

resource "google_sql_database" "app" {
  project  = var.proyecto
  instance = google_sql_database_instance.principal.name
  name     = "reservas"
}

resource "google_sql_user" "app" {
  project  = var.proyecto
  instance = google_sql_database_instance.principal.name
  name     = "app"
  password = random_password.bd.result
}

Five details that count:

  • ipv4_enabled = false is the line worth 2 points of block D and, more importantly, the one that stops your database being scannable from the internet.
  • The password is generated with random_password and you never see it. It goes straight to Secret Manager. Nobody types it, nobody copies it, nobody uploads it by mistake.
  • depends_on on the peering is one of the few cases where an explicit depends_on is correct: the dependency is real but Terraform cannot infer it from the graph.
  • deletion_protection conditional on the environment: dev gets destroyed on Fridays, prod does not get destroyed by accident.
  • disk_type = "PD_HDD": for a portfolio project, SSD adds nothing and costs more.

5.2 Buckets with a lifecycle

resource "random_id" "sufijo" { byte_length = 2 }

resource "google_storage_bucket" "fotos" {
  project  = var.proyecto
  name     = "${var.prefijo}-fotos-${random_id.sufijo.hex}"   # global name
  location = var.region

  uniform_bucket_level_access = true
  public_access_prevention    = "enforced"

  versioning { enabled = true }

  lifecycle_rule {
    condition { age = 90, matches_storage_class = ["STANDARD"] }
    action    { type = "SetStorageClass", storage_class = "NEARLINE" }
  }
  lifecycle_rule {
    condition { age = 365, matches_storage_class = ["NEARLINE"] }
    action    { type = "SetStorageClass", storage_class = "COLDLINE" }
  }
  lifecycle_rule {
    condition { num_newer_versions = 3 }     # do not pile up old versions
    action    { type = "Delete" }
  }

  cors {
    origin          = ["https://${var.dominio}"]
    method          = ["GET", "HEAD"]
    response_header = ["Content-Type"]
    max_age_seconds = 3600
  }

  labels = merge(var.etiquetas, { componente = "web" })
}

public_access_prevention = "enforced" prevents anyone, at bucket level, from making it public either by mistake or on purpose. It is one line and it eliminates the most common cloud incident at the root.

5.3 Schema migrations as code

The schema is not created by hand in a SQL console. It is versioned:

data/migraciones/
├── 001_esquema_inicial.sql
├── 002_indice_disponibilidad.sql
└── 003_columnas_sentimiento.sql
-- data/migraciones/001_esquema_inicial.sql
-- Idempotent: it can be run twice without breaking anything.
BEGIN;

CREATE TABLE IF NOT EXISTS refugio (
    id        SERIAL PRIMARY KEY,
    nombre    TEXT NOT NULL UNIQUE,
    altitud_m INTEGER NOT NULL CHECK (altitud_m BETWEEN 500 AND 3500),
    capacidad INTEGER NOT NULL CHECK (capacidad > 0),
    activo    BOOLEAN NOT NULL DEFAULT TRUE,
    creado_en TIMESTAMPTZ NOT NULL DEFAULT NOW()
);

CREATE TABLE IF NOT EXISTS reserva (
    id             UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    refugio_id     INTEGER NOT NULL REFERENCES refugio(id),
    fecha          DATE NOT NULL,
    plazas         INTEGER NOT NULL CHECK (plazas BETWEEN 1 AND 12),
    nombre_titular TEXT NOT NULL,
    email_titular  TEXT NOT NULL,
    estado         TEXT NOT NULL DEFAULT 'confirmada'
                   CHECK (estado IN ('confirmada','cancelada')),
    creada_en      TIMESTAMPTZ NOT NULL DEFAULT NOW()
);

-- Record of applied migrations
CREATE TABLE IF NOT EXISTS _migraciones (
    version    TEXT PRIMARY KEY,
    aplicada_en TIMESTAMPTZ NOT NULL DEFAULT NOW()
);
INSERT INTO _migraciones (version) VALUES ('001')
  ON CONFLICT (version) DO NOTHING;

COMMIT;

To run them against a DB with no public IP, use the auth proxy:

# Download cloud-sql-proxy if you do not have it
./cloud-sql-proxy --port 5432 "${PROY_DEV}:${REGION}:refugio-db" &

PGPASSWORD=$(gcloud secrets versions access latest --secret=refugio-db-password --project="${PROY_DEV}")
export PGPASSWORD

for f in data/migraciones/*.sql; do
  echo "→ $f"
  psql -h 127.0.0.1 -U app -d reservas -v ON_ERROR_STOP=1 -f "$f"
done

unset PGPASSWORD

Note the unset PGPASSWORD and the fact that the password is read from Secret Manager on the spot: it is never written into a file, nor into the shell history if you use HISTCONTROL=ignorespace and a leading space.

5.4 Seeding fictional data

# data/seed/generar.py — generates 100% invented data
import random, uuid
from datetime import date, timedelta

random.seed(42)  # reproducible: the same data on every run

REFUGIOS = [
    ("Refugio de Cotiella", 2100, 40), ("Refugio Peña Blanca", 1850, 28),
    ("Refugio del Ibón Verde", 2340, 22), ("Refugio de Valdellosa", 1620, 55),
]
NOMBRES = ["Ana", "Luis", "Marta", "Jorge", "Carmen", "Diego", "Elena", "Pablo"]
APELLIDOS = ["Soler", "Ibáñez", "Marín", "Vidal", "Rey", "Castaño", "Lorca"]

def titular_ficticio(i):
    nombre = f"{random.choice(NOMBRES)} {random.choice(APELLIDOS)}"
    # example.com is reserved by RFC 2606 for exactly this
    correo = f"usuario{i:04d}@example.com"
    return nombre, correo

def factor_estacional(d: date) -> float:
    """July-August x4, weekends x2.5, the rest normal."""
    f = 4.0 if d.month in (7, 8) else (2.0 if d.month in (6, 9) else 1.0)
    if d.weekday() >= 5:
        f *= 2.5
    return f

def generar_reservas(n=2000):
    inicio = date.today() - timedelta(days=540)
    filas = []
    i = 0
    while len(filas) < n:
        d = inicio + timedelta(days=random.randint(0, 540))
        if random.random() > factor_estacional(d) / 10:
            continue
        i += 1
        nombre, correo = titular_ficticio(i)
        filas.append((
            str(uuid.uuid4()), random.randint(1, len(REFUGIOS)), d.isoformat(),
            random.randint(1, 6), nombre, correo,
            "cancelada" if random.random() < 0.08 else "confirmada",
        ))
    return filas

random.seed(42) makes the generator reproducible: if you destroy and recreate the environment, you get exactly the same data. That keeps the dashboard screenshots valid and makes the 08-04 tests deterministic.

Phase 3 "done" criterion:

Check Command Expected
DB with no public IP gcloud sql instances describe refugio-db --format="value(settings.ipConfiguration.ipv4Enabled)" False
Secret created with a version gcloud secrets versions list refugio-db-password ≥1 ENABLED version
Bucket not public gcloud storage buckets describe gs://... --format="value(iamConfiguration.publicAccessPrevention)" enforced
Migrations applied psql -c "SELECT version FROM _migraciones ORDER BY version" All of them
Data seeded psql -c "SELECT count(*) FROM reserva" ~2,000

  1. Phase 4 — The application

Objective: a container that honours the Cloud Run contract, configured per environment, with injected secrets, health probes and correlatable logs. And deployed by hand once, as a proof of life.

Estimated time: 8-12 hours.

6.1 The Cloud Run contract

Four rules. Breaking any of them makes the deployment fail with a rather uninformative error:

Rule What it means Error if you break it
Listen on $PORT The PORT environment variable, not a fixed port The container does not pass the startup check
Listen on 0.0.0.0 Not on 127.0.0.1 Same: it looks started but does not respond
Start in <4 min No migrations or slow loading at startup Deployment timeout
No state on disk The filesystem is ephemeral and per instance Data that disappears with no explanation
# app/Dockerfile — multi-stage, no root, and with the bare minimum inside
FROM python:3.12-slim AS build
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir --user -r requirements.txt

FROM python:3.12-slim
RUN useradd --create-home --uid 1001 app
WORKDIR /app
COPY --from=build /root/.local /home/app/.local
COPY --chown=app:app src/ ./src/
USER app
ENV PATH=/home/app/.local/bin:$PATH \
    PYTHONUNBUFFERED=1

# $PORT is injected by Cloud Run. 8080 is just the local default.
ENV PORT=8080
CMD exec uvicorn src.main:app --host 0.0.0.0 --port ${PORT}

USER app is worth 1 security point and costs two lines. PYTHONUNBUFFERED=1 makes logs come out immediately instead of sitting in the buffer; without it, the logs of a container that crashes are lost.

6.2 Configuration and secrets

Non-sensitive configuration → service environment variables. Secrets → Secret Manager, mounted by reference:

resource "google_cloud_run_v2_service" "web" {
  project  = var.proyecto
  name     = "${var.prefijo}-web"
  location = var.region
  ingress  = "INGRESS_TRAFFIC_ALL"

  template {
    service_account = var.sa_web_email

    scaling {
      min_instance_count = var.run_min_instancias   # 0: scales to zero
      max_instance_count = var.run_max_instancias   # spending cap
    }

    vpc_access {
      connector = var.conector_id
      egress    = "PRIVATE_RANGES_ONLY"   # only the DB goes through the VPC
    }

    containers {
      image = var.imagen

      resources {
        limits = { cpu = "1", memory = "512Mi" }
        cpu_idle = true                    # do not pay for CPU between requests
      }

      # --- Configuration: in the clear, not sensitive ---
      env { name = "ENTORNO"   value = var.entorno }
      env { name = "REGION"    value = var.region }
      env { name = "BD_HOST"   value = var.bd_ip_privada }
      env { name = "BD_NOMBRE" value = "reservas" }
      env { name = "TOPIC_EVENTOS" value = var.topic_eventos }
      env { name = "BUCKET_FOTOS"  value = var.bucket_fotos }

      # --- Secrets: by reference, never by value ---
      env {
        name = "BD_PASSWORD"
        value_source {
          secret_key_ref {
            secret  = var.secreto_bd_password
            version = "latest"
          }
        }
      }

      startup_probe {
        http_get { path = "/salud/arranque" }
        initial_delay_seconds = 3
        period_seconds        = 3
        failure_threshold     = 10
      }
      liveness_probe {
        http_get { path = "/salud/vivo" }
        period_seconds = 30
      }
    }
  }

  traffic {
    type    = "TRAFFIC_TARGET_ALLOCATION_TYPE_LATEST"
    percent = 100
  }
}

cpu_idle = true is the option that stops Cloud Run charging you for CPU between requests. In a portfolio project with sporadic traffic, it is the difference between cents and euros.

6.3 The two probes and why they are different

# app/src/main.py (excerpt)
from fastapi import FastAPI, Response
from . import db

app = FastAPI()

@app.get("/salud/arranque")
async def arranque(response: Response):
    """Startup probe: am I ready to receive traffic?
    Checks critical dependencies. If it fails, Cloud Run sends no traffic."""
    try:
        await db.ping()
        return {"estado": "listo"}
    except Exception as e:
        response.status_code = 503
        return {"estado": "no listo", "motivo": str(e)[:200]}

@app.get("/salud/vivo")
async def vivo():
    """Liveness probe: is the process responding?
    Does NOT check dependencies: if the DB goes down, we do not want
    Cloud Run restarting the instance in a loop — it would fix nothing."""
    return {"estado": "vivo"}

The distinction is important and it comes up in interviews: startup checks dependencies, liveness does not. If the liveness probe checked the database, a DB outage would trigger continuous restarts of every instance, turning a problem into a storm.

6.4 Structured logs with trace_id

From 06-06, in its minimal, effective form:

# app/src/observabilidad.py
import json, os, sys, contextvars

_trace = contextvars.ContextVar("trace", default=None)
PROYECTO = os.environ.get("GOOGLE_CLOUD_PROJECT", "")

def fijar_trace(cabecera: str | None):
    """Cloud Run sends X-Cloud-Trace-Context: TRACE_ID/SPAN_ID;o=1"""
    if cabecera:
        _trace.set(cabecera.split("/")[0])

def log(severidad: str, mensaje: str, **campos):
    entrada = {
        "severity": severidad,           # names Cloud Logging understands
        "message": mensaje,
        **campos,
    }
    t = _trace.get()
    if t and PROYECTO:
        # This exact key is what links the log to the trace in the console
        entrada["logging.googleapis.com/trace"] = f"projects/{PROYECTO}/traces/{t}"
    print(json.dumps(entrada, ensure_ascii=False), file=sys.stdout, flush=True)
# In the application middleware
@app.middleware("http")
async def correlacion(request, call_next):
    fijar_trace(request.headers.get("X-Cloud-Trace-Context"))
    respuesta = await call_next(request)
    log("INFO", "peticion",
        ruta=request.url.path, metodo=request.method,
        codigo=respuesta.status_code)
    return respuesta

The logging.googleapis.com/trace key with that exact name is what lets you click on a slow trace in the console and see every log for that specific request. It is one line of code and it transforms debugging.

⚠️ Never log personal data. No emails, no names, no form contents, no tokens. Log identifiers (reserva_id), not people. Logs get replicated, exported and retained; a piece of personal data in a log is a piece of personal data you have lost sight of.

6.5 The first deployment, by hand

It is done manually and just once. The reason is both pedagogical and practical: understanding each step before automating it, and having a proof of life to compare against when the pipeline fails.

export IMAGEN="${REGION}-docker.pkg.dev/${PROY_DEV}/refugio-imagenes/web:manual-1"

gcloud builds submit app/ --tag="${IMAGEN}" --project="${PROY_DEV}"

gcloud run deploy refugio-web \
  --image="${IMAGEN}" \
  --region="${REGION}" \
  --project="${PROY_DEV}" \
  --service-account="sa-refugio-web@${PROY_DEV}.iam.gserviceaccount.com" \
  --vpc-connector="refugio-conn" \
  --vpc-egress=private-ranges-only \
  --set-env-vars="ENTORNO=dev,BD_HOST=10.10.16.3,BD_NOMBRE=reservas" \
  --set-secrets="BD_PASSWORD=refugio-db-password:latest" \
  --min-instances=0 --max-instances=2 \
  --no-allow-unauthenticated

URL=$(gcloud run services describe refugio-web --region="${REGION}" \
      --project="${PROY_DEV}" --format="value(status.url)")
curl -s -H "Authorization: Bearer $(gcloud auth print-identity-token)" "${URL}/salud/arranque"

Once it works, erase it from your mental history and do it from Terraform. The manual deployment was the proof of life; the permanent state is governed by code.

Phase 4 "done" criterion:

Check Command Expected
Service deployed gcloud run services describe refugio-web --region=$REGION --format="value(status.conditions[0].status)" True
Startup probe OK curl .../salud/arranque {"estado":"listo"}
Reads from the DB curl .../api/refugios A list with data
Structured logs gcloud logging read 'resource.type="cloud_run_revision"' --limit=1 --format=json JSON with jsonPayload and trace
Does not run as root docker run --rm $IMAGEN id -u 1001

  1. Phase 5 — Automated delivery

Objective: that a git push to main tests, builds, publishes and deploys to development without you touching anything; and that promotion to production is the same image with an approval.

Estimated time: 4-6 hours.

7.1 The development pipeline

# cloudbuild.yaml
substitutions:
  _REGION: europe-west1
  _SERVICIO: refugio-web
  _REPO: refugio-imagenes

steps:
  # 1. Tests BEFORE building. If they fail, there is no image.
  - id: pruebas
    name: python:3.12-slim
    entrypoint: bash
    args:
      - -c
      - |
        pip install --no-cache-dir -r app/requirements.txt -r app/requirements-dev.txt
        cd app && python -m pytest tests/ -v --tb=short

  # 2. Build using the previous image as cache
  - id: construir
    name: gcr.io/cloud-builders/docker
    args:
      - build
      - --cache-from=${_REGION}-docker.pkg.dev/$PROJECT_ID/${_REPO}/web:latest
      - -t=${_REGION}-docker.pkg.dev/$PROJECT_ID/${_REPO}/web:$SHORT_SHA
      - -t=${_REGION}-docker.pkg.dev/$PROJECT_ID/${_REPO}/web:latest
      - app/
    waitFor: [pruebas]

  # 3. Publish
  - id: publicar
    name: gcr.io/cloud-builders/docker
    args: [push, --all-tags, "${_REGION}-docker.pkg.dev/$PROJECT_ID/${_REPO}/web"]

  # 4. Deploy to DEVELOPMENT
  - id: desplegar
    name: gcr.io/google.com/cloudsdktool/cloud-sdk:slim
    entrypoint: gcloud
    args:
      - run
      - deploy
      - ${_SERVICIO}
      - --image=${_REGION}-docker.pkg.dev/$PROJECT_ID/${_REPO}/web:$SHORT_SHA
      - --region=${_REGION}
      - --revision-suffix=$SHORT_SHA

  # 5. Smoke: if the freshly deployed revision does not respond, the build fails
  - id: humo
    name: gcr.io/google.com/cloudsdktool/cloud-sdk:slim
    entrypoint: bash
    args:
      - -c
      - |
        URL=$(gcloud run services describe ${_SERVICIO} --region=${_REGION} --format='value(status.url)')
        TOKEN=$(gcloud auth print-identity-token)
        CODIGO=$(curl -s -o /dev/null -w '%{http_code}' -H "Authorization: Bearer $$TOKEN" "$$URL/salud/arranque")
        echo "Code: $$CODIGO"
        test "$$CODIGO" = "200"

options:
  logging: CLOUD_LOGGING_ONLY
  machineType: E2_HIGHCPU_8
timeout: 900s

The order is the important part: tests → build → publish → deploy → smoke. The tests go before the build, because building an image from code that does not pass the tests is time and money thrown away. And the smoke test goes after the deployment, because it is the only thing that distinguishes "the deployment finished" from "the deployment worked".

7.2 Promotion to production

# cloudbuild-prod.yaml — does NOT build. It promotes the already-tested image.
substitutions:
  _IMAGEN_SHA: ""     # passed explicitly: the one already working in dev

steps:
  - id: verificar-imagen-existe
    name: gcr.io/google.com/cloudsdktool/cloud-sdk:slim
    entrypoint: bash
    args:
      - -c
      - |
        test -n "${_IMAGEN_SHA}" || { echo "Missing _IMAGEN_SHA"; exit 1; }
        gcloud artifacts docker images describe \
          europe-west1-docker.pkg.dev/refugio-dev/refugio-imagenes/web:${_IMAGEN_SHA}

  # CANARY deployment: 10% of traffic to the new revision
  - id: canario
    name: gcr.io/google.com/cloudsdktool/cloud-sdk:slim
    entrypoint: gcloud
    args:
      - run
      - deploy
      - refugio-web
      - --image=europe-west1-docker.pkg.dev/refugio-dev/refugio-imagenes/web:${_IMAGEN_SHA}
      - --region=europe-west1
      - --project=refugio-prod
      - --no-traffic                       # deployed without receiving traffic
      - --revision-suffix=${_IMAGEN_SHA}

  - id: repartir-10
    name: gcr.io/google.com/cloudsdktool/cloud-sdk:slim
    entrypoint: gcloud
    args:
      - run
      - services
      - update-traffic
      - refugio-web
      - --region=europe-west1
      - --project=refugio-prod
      - --to-revisions=refugio-web-${_IMAGEN_SHA}=10

The principle to internalise: the same image, without rebuilding. Rebuilding for production means you are deploying something you have never tested: the pip install may resolve a different version, the base image may have changed. The image that passed the tests in dev is the one that goes to production, byte for byte.

Manual approval is configured on the Cloud Build trigger (--require-approval) or, in GitHub Actions, with a protected environment.

Phase 5 "done" criterion:

Check How Expected
A push deploys on its own git push and check gcloud builds list --limit=1 SUCCESS in <10 min
A broken test stops the deployment Break a test on purpose, push Build in FAILURE, Cloud Run revision unchanged
No keys in CI Review the repository secrets No GCP credential
Promotion without rebuilding Compare the digest in dev and prod Identical

The second row is the real test. A pipeline that has never failed is untested: break a test on purpose, check that the deployment stops, and save the screenshot. It is worth 3 points and it demonstrates that the safety net exists.

  1. Phase 6 — The data and AI layer

Objective: that an action in the application ends up visible in the dashboard, and that there is an AI component that works.

Estimated time: 6-9 hours.

8.1 Ingestion: the application publishes events

# app/src/eventos.py
import json, os
from google.cloud import pubsub_v1

_publisher = pubsub_v1.PublisherClient()
_TOPIC = _publisher.topic_path(os.environ["GOOGLE_CLOUD_PROJECT"],
                               os.environ["TOPIC_EVENTOS"])

def publicar_reserva(evento: dict) -> None:
    """Publishes the business event. NEVER includes contact details."""
    carga = {
        "evento_id":      evento["id"],
        "tipo_evento":    evento["tipo"],       # creada | cancelada
        "reserva_id":     evento["reserva_id"],
        "refugio_id":     evento["refugio_id"],
        "refugio_nombre": evento["refugio_nombre"],
        "fecha_estancia": evento["fecha"],
        "plazas":         evento["plazas"],
        "ocurrido_en":    evento["ts"],
    }
    futuro = _publisher.publish(_TOPIC, json.dumps(carga).encode("utf-8"))
    futuro.result(timeout=10)

The comment is not decorative: the event carries no nombre_titular and no email_titular. It is the data minimisation from 08-02 applied in the code, and it is the difference between having personal data in one place or in four.

8.2 Transformation: the function that writes into BigQuery

# ml/../funcion/main.py
import base64, json, os
from google.cloud import bigquery

_bq = bigquery.Client()
_TABLA = os.environ["TABLA_EVENTOS"]

def procesar(evento, contexto):
    """Pub/Sub subscriber. Inserts the event into BigQuery."""
    datos = json.loads(base64.b64decode(evento["data"]).decode("utf-8"))

    errores = _bq.insert_rows_json(_TABLA, [datos], row_ids=[datos["evento_id"]])
    if errores:
        # Raising an exception makes Pub/Sub retry; after N attempts,
        # the message goes to the dead-letter queue.
        raise RuntimeError(f"Error inserting into BigQuery: {errores}")

row_ids with the event identifier activates BigQuery deduplication: if Pub/Sub delivers the same message twice — and it will, because its guarantee is "at least once" — the row is not duplicated. It is one line that prevents a dashboard with inflated figures.

And configure the dead-letter queue in Terraform, with its topic and its subscription, or failing messages will retry forever.

8.3 The analytical tables

-- data/sql/analitica.sql
CREATE TABLE IF NOT EXISTS `refugio-datos.refugio_analitica.reservas_eventos` (
  evento_id      STRING NOT NULL,
  ocurrido_en    TIMESTAMP NOT NULL,
  tipo_evento    STRING NOT NULL,
  reserva_id     STRING NOT NULL,
  refugio_id     INT64 NOT NULL,
  refugio_nombre STRING,
  fecha_estancia DATE NOT NULL,
  plazas         INT64 NOT NULL
)
PARTITION BY DATE(ocurrido_en)
CLUSTER BY refugio_id
OPTIONS (partition_expiration_days = 1095, require_partition_filter = TRUE);

-- Aggregated view: this is what the dashboard consumes.
-- Having the dashboard query a view rather than the base table reduces scanning.
CREATE OR REPLACE VIEW `refugio-datos.refugio_analitica.v_ocupacion_diaria` AS
SELECT
  fecha_estancia,
  refugio_id,
  ANY_VALUE(refugio_nombre) AS refugio,
  SUM(IF(tipo_evento = 'creada',    plazas, 0)) AS plazas_reservadas,
  SUM(IF(tipo_evento = 'cancelada', plazas, 0)) AS plazas_canceladas,
  SUM(IF(tipo_evento = 'creada', plazas, -plazas)) AS plazas_netas
FROM `refugio-datos.refugio_analitica.reservas_eventos`
WHERE DATE(ocurrido_en) >= DATE_SUB(CURRENT_DATE(), INTERVAL 730 DAY)
GROUP BY fecha_estancia, refugio_id;

The WHERE DATE(ocurrido_en) >= ... in the view is not optional: with require_partition_filter = TRUE, a view without a partition filter would fail.

8.4 The AI component

The recommendation, repeated because it matters: start with a pre-trained API.

# ml/analizar_opiniones.py
from google.cloud import language_v2

_cliente = language_v2.LanguageServiceClient()

def analizar(texto: str) -> dict:
    """Sentiment of a review. FICTIONAL text, no personal data."""
    documento = language_v2.Document(
        content=texto,
        type_=language_v2.Document.Type.PLAIN_TEXT,
        language_code="es",
    )
    r = _cliente.analyze_sentiment(request={"document": documento})
    return {
        "sentimiento": round(r.document_sentiment.score, 3),   # -1..1
        "magnitud":    round(r.document_sentiment.magnitude, 3),
    }

And the architecture decision that saves money: the analysis is done once, when the review is created, and the result is stored in the database. The API is not called every time somebody opens the dashboard. It is the same reasoning that took AlpinaShop to batch recommendations in DA-002: if the result does not change, it is not recomputed.

Phase 6 "done" criterion:

# End-to-end test: create a booking and see it in BigQuery
curl -s -X POST "${URL}/api/reservas" -H 'Content-Type: application/json' \
  -d '{"refugio_id":1,"fecha":"2026-12-20","plazas":2,
       "nombre_titular":"Fictional Test","email_titular":"[email protected]"}'

sleep 30

bq query --use_legacy_sql=false --project_id="${PROY_DATOS}" \
'SELECT evento_id, tipo_evento, refugio_id, plazas
 FROM `refugio-datos.refugio_analitica.reservas_eventos`
 WHERE DATE(ocurrido_en) = CURRENT_DATE()
 ORDER BY ocurrido_en DESC LIMIT 5'

If that row appears, you have closed milestone M3 from 08-01: the complete flow works end to end. Take a screenshot: you will need it in the presentation.

  1. Phase 7 — Exposure

Objective: that the system responds on a domain of your own, with valid HTTPS, and with whatever protection you decided in the design.

Estimated time: 3-5 hours, plus DNS propagation and certificate issuance time (from 15 minutes to several hours: do not leave it to the last day).

9.1 The two possible routes

Cloud Run custom domain Global HTTPS load balancer
Cost €0 ~€18/month for the forwarding rule
TLS Managed and automatic Managed and automatic
Cloud CDN ❌ ✅
Cloud Armor (WAF) ❌ ✅
Several backends (Run + bucket) ❌ ✅
Complexity 2 resources 7-8 resources

If your cost limit is tight, the first option meets RNF-5 and costs nothing. Document why in an ADR, as RefugioReserva did.

9.2 The cheap route: custom domain

resource "google_cloud_run_domain_mapping" "web" {
  project  = var.proyecto
  location = var.region
  name     = var.dominio                 # "refugioreserva.example"
  metadata { namespace = var.proyecto }
  spec     { route_name = google_cloud_run_v2_service.web.name }
}

output "registros_dns" {
  description = "Records to create at the registrar"
  value       = google_cloud_run_domain_mapping.web.status[0].resource_records
}

You create the records that output returns at your registrar (or in Cloud DNS) and Google issues the certificate on its own.

9.3 The complete route: load balancer, CDN and WAF

Even if you end up destroying it on cost grounds, write the module and deploy it at least once: the knowledge stays, you have screenshots and you can talk about it in the presentation.

resource "google_compute_region_network_endpoint_group" "neg" {
  project               = var.proyecto
  name                  = "${var.prefijo}-neg"
  region                = var.region
  network_endpoint_type = "SERVERLESS"
  cloud_run { service = google_cloud_run_v2_service.web.name }
}

resource "google_compute_backend_service" "bs" {
  project                 = var.proyecto
  name                    = "${var.prefijo}-bs"
  protocol                = "HTTPS"
  load_balancing_scheme   = "EXTERNAL_MANAGED"
  enable_cdn              = true
  security_policy         = google_compute_security_policy.waf.id

  backend { group = google_compute_region_network_endpoint_group.neg.id }

  cdn_policy {
    cache_mode        = "CACHE_ALL_STATIC"
    default_ttl       = 3600
    client_ttl        = 3600
    negative_caching  = true
  }

  log_config { enable = true, sample_rate = 1.0 }
}

resource "google_compute_security_policy" "waf" {
  project = var.proyecto
  name    = "${var.prefijo}-waf"

  # Rate limiting: it protects your wallet as much as your application
  rule {
    action   = "throttle"
    priority = 1000
    match {
      versioned_expr = "SRC_IPS_V1"
      config { src_ip_ranges = ["*"] }
    }
    rate_limit_options {
      conform_action = "allow"
      exceed_action  = "deny(429)"
      enforce_on_key = "IP"
      rate_limit_threshold { count = 100, interval_sec = 60 }
    }
  }

  rule {
    action   = "allow"
    priority = 2147483647
    match {
      versioned_expr = "SRC_IPS_V1"
      config { src_ip_ranges = ["*"] }
    }
    description = "Default rule"
  }
}

The rate limiting rule deserves a note: in a project with a €12 budget, a bot making 100,000 requests can cost you the whole month's budget. The limit of 100 requests per minute per IP protects the application as much as the bill. If you do not use a load balancer, the equivalent is --max-instances, which is a hard spending cap.

Phase 7 "done" criterion:

curl -sI "https://${DOMINIO}" | head -1                       # HTTP/2 200
curl -sI "https://${DOMINIO}" | grep -i strict-transport       # HSTS present
echo | openssl s_client -connect "${DOMINIO}:443" -servername "${DOMINIO}" 2>/dev/null \
  | openssl x509 -noout -dates -issuer                         # valid certificate
curl -sI "http://${DOMINIO}" | head -1                         # 301 to HTTPS

  1. Phase 8 — Observability

Objective: finding out something is wrong before somebody tells you.

Estimated time: 4-6 hours.

10.1 The four signals dashboard

resource "google_monitoring_dashboard" "principal" {
  project = var.proyecto
  dashboard_json = jsonencode({
    displayName = "RefugioReserva — overview"
    gridLayout  = { columns = 2, widgets = [
      {
        title = "Traffic (requests/s)"
        xyChart = { dataSets = [{ timeSeriesQuery = { timeSeriesFilter = {
          filter = "metric.type=\"run.googleapis.com/request_count\" resource.type=\"cloud_run_revision\""
          aggregation = { alignmentPeriod = "60s", perSeriesAligner = "ALIGN_RATE" }
        }}}]}
      },
      {
        title = "Errors (5xx/s)"
        xyChart = { dataSets = [{ timeSeriesQuery = { timeSeriesFilter = {
          filter = "metric.type=\"run.googleapis.com/request_count\" metric.label.response_code_class=\"5xx\""
          aggregation = { alignmentPeriod = "60s", perSeriesAligner = "ALIGN_RATE" }
        }}}]}
      },
      {
        title = "p95 latency (ms)"
        xyChart = { dataSets = [{ timeSeriesQuery = { timeSeriesFilter = {
          filter = "metric.type=\"run.googleapis.com/request_latencies\""
          aggregation = { alignmentPeriod = "60s", perSeriesAligner = "ALIGN_DELTA",
                          crossSeriesReducer = "REDUCE_PERCENTILE_95" }
        }}}]}
      },
      {
        title = "Active instances (saturation)"
        xyChart = { dataSets = [{ timeSeriesQuery = { timeSeriesFilter = {
          filter = "metric.type=\"run.googleapis.com/container/instance_count\""
          aggregation = { alignmentPeriod = "60s", perSeriesAligner = "ALIGN_MEAN" }
        }}}]}
      }
    ]}
  })
}

10.2 The alert that actually notifies

resource "google_monitoring_notification_channel" "correo" {
  project      = var.proyecto
  display_name = "Owner's email"
  type         = "email"
  labels       = { email_address = var.correo_alertas }
}

resource "google_monitoring_alert_policy" "errores_5xx" {
  project      = var.proyecto
  display_name = "5xx error rate > 5%"
  combiner     = "OR"

  conditions {
    display_name = "Elevated 5xx for 5 minutes"
    condition_threshold {
      filter = join(" ", [
        "metric.type=\"run.googleapis.com/request_count\"",
        "resource.type=\"cloud_run_revision\"",
        "metric.label.response_code_class=\"5xx\"",
      ])
      comparison      = "COMPARISON_GT"
      threshold_value = 0.05
      duration        = "300s"
      aggregations {
        alignment_period   = "60s"
        per_series_aligner = "ALIGN_RATE"
      }
    }
  }

  notification_channels = [google_monitoring_notification_channel.correo.id]

  documentation {
    content = <<-EOT
      ## Elevated 5xx errors

      **First steps** (see `docs/runbook.md`):
      1. `gcloud logging read 'severity>=ERROR' --limit=20 --freshness=15m`
      2. Does it coincide with a deployment? `gcloud run revisions list --limit=5`
      3. If it does: roll back with
         `gcloud run services update-traffic refugio-web --to-revisions=<previous>=100`
      4. Is the database reachable? `gcloud sql instances describe refugio-db`
    EOT
    mime_type = "text/markdown"
  }
}

The documentation block is what turns an alert into something useful. An alert that only says "something is wrong" wakes you up; one that says what to look at and how to roll back lets you fix it. Always write the first steps in there.

10.3 The SLO as a resource

resource "google_monitoring_slo" "disponibilidad" {
  project = var.proyecto
  service = google_monitoring_service.web.service_id
  slo_id  = "disponibilidad-api"
  display_name = "99.5% of requests with no 5xx error (30 days)"

  goal                = 0.995
  rolling_period_days = 30

  request_based_sli {
    good_total_ratio {
      total_service_filter = "metric.type=\"run.googleapis.com/request_count\" resource.type=\"cloud_run_revision\""
      bad_service_filter   = "metric.type=\"run.googleapis.com/request_count\" resource.type=\"cloud_run_revision\" metric.label.response_code_class=\"5xx\""
    }
  }
}

10.4 The uptime check

resource "google_monitoring_uptime_check_config" "web" {
  project      = var.proyecto
  display_name = "RefugioReserva available"
  timeout      = "10s"
  period       = "300s"

  http_check {
    path         = "/salud/arranque"
    port         = 443
    use_ssl      = true
    validate_ssl = true
  }

  monitored_resource {
    type   = "uptime_url"
    labels = { host = var.dominio, project_id = var.proyecto }
  }

  selected_regions = ["EUROPE", "USA"]
}

Phase 8 "done" criterion — and here there is one you do not meet by looking, but by provoking:

Check How Expected
Dashboard exists gcloud monitoring dashboards list 1 dashboard
Alert exists gcloud alpha monitoring policies list ≥1 policy
The alert really notifies Provoke it (see below) Email received
SLO calculating Console → SLO Error budget with a value
Uptime check gcloud monitoring uptime list-configs 1, in a correct state
# Provoke the alert on purpose: deploy a revision that returns 500
# on a test endpoint, generate traffic, and wait for the email.
for i in $(seq 1 200); do curl -s -o /dev/null "${URL}/api/error-de-prueba"; done
# Wait 5-10 minutes. If the email does not arrive, the alert is useless.

An alert that has never fired is not an alert: it is an intention. Provoking it is the only way to know that the notification channel works, that the threshold is reachable and that the email does not end up in spam. Save the screenshot of the email received: it is worth 2 points.

  1. Cross-cutting practices throughout the implementation

These six things do not belong to any one phase: they belong to all of them.

11.1 Small, frequent commits

One commit per comprehensible unit of work. Add network module with serverless connector is a commit; Various progress is not.

git add infra/modules/red/
git commit -m "Add network module: VPC, subnet, /28 connector and service peering"

Concrete benefits: you can revert one thing without reverting five, the history tells the story of the project in the presentation, and git bisect is actually useful when something breaks.

11.2 terraform plan always reviewed

Never apply without having read the plan. And pay special attention to three phrases:

In the plan Meaning Reaction
will be created New resource Normal
will be updated in-place Change without recreating Normal
must be replaced It is destroyed and created again STOP and understand why
will be destroyed It disappears Check that it is intentional

must be replaced on a database means losing the data. On a bucket, losing the objects. It is almost always caused by changing a ForceNew attribute (the name, the region, a CIDR). If it appears and you were not expecting it, cancel.

terraform plan -out=plan.tfplan
terraform show -json plan.tfplan | \
  jq -r '.resource_changes[] | select(.change.actions | index("delete")) | .address'
# If this returns something you were not expecting, do not apply.

11.3 Do not touch anything by hand outside Terraform

The ideal rule. And reality: you are going to do it, in a hurry, on a Tuesday night.

What to do then, in order:

  1. Note it down immediately in docs/diario.md. The sin is not the click; it is forgetting it.
  2. Detect the drift: terraform plan will tell you that reality does not match the code.
  3. Decide: either you bring the change into the code (the usual choice) or you revert the manual change.
  4. If you created a resource by hand, import it instead of recreating it:
# Terraform 1.5+: import block, versionable in the code
cat >> infra/envs/dev/imports.tf <<'EOF'
import {
  to = google_storage_bucket.temporal
  id = "refugio-dev/refugio-temporal-a1b2"
}
EOF
terraform plan   # generates the missing configuration

The free detector: the gestionado-por = terraform label from 08-02. Any resource without it was created by hand.

11.4 Documenting on the fly

Three lines per session in docs/diario.md and an ADR every time you decide something with alternatives. It costs nothing and it is the difference between a presentation with substance and one invented the night before.

11.5 Checking spending every few days

# Aliases worth having to hand
gcloud billing accounts list
# And, once the BigQuery export has been populating for a few days:
bq query --use_legacy_sql=false \
'SELECT service.description AS service, ROUND(SUM(cost),2) AS cost_eur
 FROM `mi-proyecto.facturacion.gcp_billing_export_v1_XXXX`
 WHERE DATE(usage_start_time) >= DATE_SUB(CURRENT_DATE(), INTERVAL 7 DAY)
 GROUP BY service ORDER BY cost_eur DESC'

Twice a week, thirty seconds. And destroy dev when you are not using it:

terraform -chdir=infra/envs/dev destroy -auto-approve   # Friday
terraform -chdir=infra/envs/dev apply   -auto-approve   # Saturday

This has a double benefit worth underlining: you save money and you are running the hardest test of RNF-1 every week. If one Friday Saturday's apply does not rebuild the environment, you have just discovered a flaw in your IaC at the cheapest possible moment.

11.6 Keeping the environments aligned

Every time you change something in prod, check that dev has the equivalent. Drift between environments turns promotion into a lottery.

  1. What to do when you get stuck: the five-step method

Step Question Command
1 What does the error say, in full? gcloud logging read 'severity>=ERROR' --limit=20 --freshness=1h --format=json
2 Is the API enabled? gcloud services list --enabled | grep <api>
3 Is it a permission? gcloud policy-troubleshoot iam <resource> --principal-email=<sa> --permission=<permission>
4 Is it the network? gcloud network-management connectivity-tests create ...
5 Is the resource what I think it is? gcloud <service> describe <resource> --format=yaml

The ten errors you are going to hit, with their cause

Symptom Almost certain cause Fix
PERMISSION_DENIED when deploying Missing iam.serviceAccountUser on the runtime SA Grant it on that specific SA
API not enabled Exactly that, or propagation still under way gcloud services enable ... and wait 2 min
Private Cloud SQL will not be created Missing private services access peering Create global_address + service_networking_connection
Cloud Run cannot connect to the DB Missing connector, or wrong vpc-egress --vpc-connector + --vpc-egress=private-ranges-only
Container does not start It does not listen on $PORT or on 0.0.0.0 Fix the CMD
Deployment times out Startup > 4 min (migrations at startup) Take the migrations out of startup
Bucket "already exists" Global name, somebody else has it Add a suffix with random_id
VPC connector will not be created The range is not an exact /28 Adjust the CIDR
Alert that never arrives Unverified channel, or email in spam Verify the channel and test by provoking it
Unexpected must be replaced You changed a ForceNew attribute Check the Terraform registry before applying

And the two-hour rule: if you have spent two hours on the same error, switch phase. Note the exact error and what you have tried in the log. Coming back the next day solves more blockages than persisting.

  1. Progress log and "done" criteria

Keep this table in docs/diario.md and update it as you finish each phase. When you work alone, it is the only thing that objectively tells you where you stand:

Phase Deliverable Verifiable "done" criterion Verification command Status
0 Projects and budget 3 billable projects, ≥25 APIs, 1 budget gcloud billing projects describe ⬜
0 Repository Structure + README + ≥2 commits git log --oneline ⬜
1 Remote state Bucket with versioning and no public access gcloud storage buckets describe ⬜
1 Network Own VPC, connector READY, peering active gcloud compute networks list ⬜
1 Clean plan terraform plan → No changes terraform plan ⬜
2 Identities 0 primitive roles on SAs, 0 user keys Phase 2 script ⬜
2 WIF Authenticated build with no secrets in the CI gcloud builds list ⬜
3 Database ipv4Enabled = False, backups on gcloud sql instances describe ⬜
3 Secrets Password in Secret Manager, not in the repo git log -p | grep -i password → empty ⬜
3 Fictional data ~2,000 rows, migrations recorded psql -c "SELECT count(*)..." ⬜
4 Application Service Ready, /salud/arranque = 200 curl ⬜
4 Logs Entries with jsonPayload and a trace field gcloud logging read ⬜
5 CI/CD A push deploys in <10 min gcloud builds list --limit=1 ⬜
5 Test that protects Broken test → deployment stopped Break it on purpose ⬜
6 Complete flow Booking → event → BigQuery in <2 min bq query ⬜
6 Dashboard 4 visualisations with data Report URL ⬜
6 AI Reviews with a sentiment score psql -c "SELECT ... WHERE sentimiento IS NOT NULL" ⬜
7 Domain and TLS HTTP/2 200 and a valid certificate curl -sI + openssl ⬜
8 Dashboard 4 signals visible gcloud monitoring dashboards list ⬜
8 Tested alert Email received after provoking it Screenshot of the email ⬜
8 SLO Error budget with a numeric value Monitoring console ⬜

The three rows in bold are the ones people tick without checking. Do not: they are precisely the ones that prove the system really works.

  1. The realistic warning: the first time, everything takes twice as long

It deserves its own section because it is the main cause of abandonment, and because it is not a problem with you.

Task First time Second time
Configuring WIF 2-3 h 15 min
Cloud SQL with a private IP 2 h 20 min
First Cloud Run deployment that works 3-4 h 30 min
Complete Cloud Build pipeline 4 h 45 min
Monitoring dashboard with 4 charts 2 h 30 min
Reusable Terraform module 3 h 45 min

What happens the first time and not the second: reading documentation, understanding the mental model, getting a field name wrong, waiting for propagation, discovering a permission was missing, undoing and redoing.

Three practical consequences:

  1. Plan for double. If you think phase 5 is 4 hours, set aside 8.
  2. Do not measure yourself against a tutorial. The twenty-minute video is edited and recorded by somebody who has done it fifty times.
  3. The "wasted" time is the learning. The two hours fighting with iam.serviceAccountUser are exactly why next time you take fifteen minutes, and why in an interview you will be able to answer instantly.

Common Mistakes and Tips

Starting with the application. It is mistake 2 from 08-01 and it shows up here. If your twentieth commit has nothing in infra/, you started in the wrong place.

Creating "just one little thing" by hand. It is never one. Three weeks later you have fifteen orphaned resources and a terraform plan full of noise that you no longer read. Note it, import it or revert it, always.

Applying without reading the plan. The day must be replaced appears on your database and you do not see it, you will lose the data and the weekend.

Granting roles/editor "temporarily". That "temporarily" lasts until the presentation. Start restrictive and open up with policy-troubleshoot.

Leaving the domain until the last day. DNS propagation and certificate issuance have timings you do not control. Do it as soon as phase 5 is stable.

Loading the AI model on every request. Analyse once and store the result. Calling an AI API on every page load is expensive and slow with no advantage whatsoever.

Not testing the alert. It is the quickest check to do and the most forgotten. An untested alert has a 50 % chance of not working when you need it.

Tip: write the "done" criterion first, the code afterwards. Before starting a phase, write the command you will use to verify it. It forces you to define what finishing means and avoids "I think it is done".

Tip: use terraform plan as a learning tool. Every time you write a new resource, run plan and read what it implies. It is the best documentation that exists on what each block does.

Tip: save the important outputs in the repository. A docs/evidencias/ directory with the output of the identity audit script, the screenshot of the first end-to-end flow, the alert email. It is worth points in 08-05 and it is impossible to reconstruct afterwards.

Tip: git commit at the end of every session even if it does not work. A commit WIP: VPC connector, peering still failing is useful information. The history is a deliverable.

Exercises

Exercise 1 — Build the reproducible foundation (phases 0 to 2)

Complete phases 0, 1 and 2 in your project: projects created with the names from the design, APIs enabled, budget with alerts, repository with structure; state bucket with versioning and public access prevention, remote backend with a prefix per environment, provider pinned by version, at least two of your own modules (network and identity) and per-environment variables with the same keys; service accounts per workload with roles scoped to the resource, and Workload Identity Federation with a repository condition.

Deliverable: the output of the identity verification script (0 primitive roles, 0 keys) and a terraform plan that says No changes.

Exercise 2 — Data, application and end-to-end flow (phases 3, 4 and 6)

Deploy your managed database with no public IP, with its password generated and stored in Secret Manager, its migrations applied as code and its fictional data seeded reproducibly. Containerise your application honouring the Cloud Run contract — $PORT, 0.0.0.0, non-root user, fast startup — with configuration through environment variables, secrets by reference, two different health probes and structured logs with trace_id. Connect the data flow through to the dashboard and add the AI component.

Deliverable: the complete trace of an action in your application that ends up visible in the dashboard, with the verification commands for each hop.

Exercise 3 — Automate, expose and observe (phases 5, 7 and 8)

Set up the pipeline that tests, builds, publishes, deploys to development and runs a smoke test, plus the promotion to production of the same image with an approval. Break a test on purpose and demonstrate with a screenshot that the deployment stops. Expose the system on your domain with valid HTTPS and a redirect from HTTP. Create the four signals dashboard, the alert with its first-steps documentation, the uptime check and the SLO.

Provoke the alert on purpose and save the email received. Close the exercise with the progress log table from section 13 completely ticked and verified.

Solutions

Solution 1 — RefugioReserva's foundation

After completing the first three phases, verification gives this:

$ terraform -chdir=infra/envs/dev plan
No changes. Your infrastructure matches the configuration.

$ gcloud projects get-iam-policy refugio-dev --format=json | \
  jq -r '.bindings[] | select(.role|test("roles/(owner|editor)")) |
         "\(.role): \(.members[])"'
roles/owner: user:[email protected]
# No service account. Correct.

$ for SA in $(gcloud iam service-accounts list --project=refugio-dev --format="value(email)"); do
    N=$(gcloud iam service-accounts keys list --iam-account="$SA" --managed-by=user \
        --format="value(name)" | wc -l)
    echo "$SA -> $N user keys"
  done
[email protected] -> 0
[email protected] -> 0
[email protected] -> 0
[email protected] -> 0

The three real problems that came up and what they cost:

Problem Symptom Cause Time lost
VPC connector would not be created Invalid IP CIDR range I had put a /27; it requires a /28 25 min
WIF authenticated but would not deploy PERMISSION_DENIED when deploying Missing iam.serviceAccountUser for sa-deploy on sa-web 1 h 40 min
Photos bucket rejected already exists Global name taken 10 min → random_id

The second is the module's classic. gcloud policy-troubleshoot pointed it out in two minutes; the problem was that it took me an hour and a half to remember to use it. Noted in the log so as not to repeat it.

Solution 2 — RefugioReserva's end-to-end flow

# 1. Verify the database
$ gcloud sql instances describe refugio-db --project=refugio-dev \
    --format="value(settings.ipConfiguration.ipv4Enabled, state)"
False   RUNNABLE

# 2. Verify the secret
$ gcloud secrets versions list refugio-db-password --project=refugio-dev \
    --format="value(name,state)"
1   ENABLED

# 3. Data seeded
$ psql -h 127.0.0.1 -U app -d reservas -c \
    "SELECT count(*) AS bookings, count(DISTINCT refugio_id) AS refuges FROM reserva"
 bookings | refuges
----------+---------
     2000 |      12

# 4. The application responds and reads from the DB
$ curl -s "${URL}/api/refugios" | jq '.[0]'
{"id":1,"nombre":"Refugio de Cotiella","altitud_m":2100,"capacidad":40}

# 5. Create a booking (FICTIONAL data)
$ curl -s -X POST "${URL}/api/reservas" -H 'Content-Type: application/json' \
    -d '{"refugio_id":3,"fecha":"2026-12-20","plazas":2,
         "nombre_titular":"Fictional Test","email_titular":"[email protected]"}' | jq
{"id":"7c2e...","estado":"confirmada","plazas_restantes":20}

# 6. The structured log, with its trace
$ gcloud logging read 'resource.type="cloud_run_revision" jsonPayload.ruta="/api/reservas"' \
    --limit=1 --format="value(jsonPayload.message, jsonPayload.codigo, trace)"
peticion  201  projects/refugio-dev/traces/8a1f...

# 7. The event reached BigQuery (26 seconds later)
$ bq query --use_legacy_sql=false --project_id=refugio-datos \
  'SELECT evento_id, tipo_evento, refugio_id, plazas, ocurrido_en
   FROM `refugio-datos.refugio_analitica.reservas_eventos`
   WHERE DATE(ocurrido_en) = CURRENT_DATE() ORDER BY ocurrido_en DESC LIMIT 1'
+-----------+-------------+------------+--------+---------------------+
| evento_id | tipo_evento | refugio_id | plazas |     ocurrido_en     |
+-----------+-------------+------------+--------+---------------------+
| 7c2e...   | creada      |          3 |      2 | 2026-10-14 18:42:11 |
+-----------+-------------+------------+--------+---------------------+

# 8. And the sentiment of the reviews
$ psql -c "SELECT count(*) FILTER (WHERE sentimiento IS NOT NULL) AS analysed,
           round(avg(sentimiento)::numeric,3) AS average FROM opinion"
 analysed | average
----------+---------
      400 |   0.412

Milestone M3 closed. Eight commands, eight verified hops, 26 seconds from the action to the analytical store. Screenshot saved in docs/evidencias/.

One detail that cost time and is worth pointing out: the events were appearing duplicated in BigQuery. Cause: Pub/Sub guarantees "at least once" and the function was being retried. Fix: row_ids=[evento_id] in insert_rows_json, which activates deduplication by identifier. Without that, the dashboard was showing 12 % more bookings than there really were, and the worst part is that I would not have noticed if I had not reconciled the total against the operational database. Lesson noted: always cross-check the analytical store against the operational one.

Solution 3 — RefugioReserva's delivery, exposure and observability

The proof that the pipeline protects. The capacity control test was broken on purpose:

# app/tests/test_unitarios.py — temporary change
def test_no_sobreventa():
    disponible = calcular_disponibilidad(refugio_id=1, fecha="2026-08-15")
    assert disponible == 999   # ← deliberately wrong value
$ git commit -am "TEST: breaking the capacity test on purpose" && git push
$ gcloud builds list --limit=1 --format="table(id, status, createTime)"
ID                                    STATUS   CREATE_TIME
9f2a-...                              FAILURE  2026-10-21T19:14:22

$ gcloud run revisions list --service=refugio-web --region=europe-west1 --limit=2 \
    --format="table(name, active, createTime)"
NAME                    ACTIVE  CREATE_TIME
refugio-web-a3f81c      True    2026-10-21T18:02:11   ← the previous one still serving

The deployment stopped at step 1. The previous revision carried on serving traffic. The safety net exists and is tested. Screenshot saved.

Exposure:

$ curl -sI https://dev.refugioreserva.example | head -3
HTTP/2 200
strict-transport-security: max-age=31536000; includeSubDomains
content-type: text/html; charset=utf-8

$ echo | openssl s_client -connect dev.refugioreserva.example:443 \
    -servername dev.refugioreserva.example 2>/dev/null | \
    openssl x509 -noout -dates -issuer
notBefore=Oct 20 09:14:00 2026 GMT
notAfter=Jan 18 09:13:59 2027 GMT
issuer=C = US, O = Google Trust Services, CN = WE1

$ curl -sI http://dev.refugioreserva.example | head -1
HTTP/1.1 301 Moved Permanently

After ADR-006, Cloud Run's custom domain is used: managed certificate, zero cost, RNF-5 met. The load balancer module was left written and will be deployed for a week in 08-04 for the load tests.

The alert, genuinely provoked:

$ for i in $(seq 1 300); do curl -s -o /dev/null "${URL}/api/error-de-prueba"; done
$ # 6 minutes later:
$ gcloud alpha monitoring policies list --format="value(displayName,enabled)"
5xx error rate > 5%   True

Email received at 20:41, six minutes after starting to generate errors. It contained the four first steps from the documentation block. Screenshot saved in docs/evidencias/alerta-recibida.png.

One detail that did not work first time: the first notification channel had been created but not verified, and the alerts were not arriving. There is no visible error: the policy shows as active and the incident opens, but the email never goes out. It was only discovered because the alert was provoked on purpose. That is exactly why provoking it is mandatory.

Final progress log:

Phase Criterion Status Evidence
0 3 projects, 27 APIs, €12 budget ✅ docs/evidencias/fase0.txt
1 Clean plan, own VPC, peering ✅ docs/evidencias/plan-limpio.txt
2 0 primitive roles, 0 keys, WIF ✅ docs/evidencias/identidades.txt
3 Private DB, secret, 2,000 rows ✅ docs/evidencias/datos.txt
4 Service Ready, logs with trace ✅ docs/evidencias/app.txt
5 Push deploys; a broken test stops it ✅ docs/evidencias/build-fallido.png
6 Booking → BigQuery in 26 s; 400 reviews analysed ✅ docs/evidencias/e2e.txt
7 HTTP/2 200, HSTS, valid certificate, 301 ✅ docs/evidencias/tls.txt
8 Dashboard, alert received, SLO, uptime ✅ docs/evidencias/alerta-recibida.png

Actual time invested: 63 hours against the 50 estimated. The overrun was concentrated in three places: WIF (2.5 h against the 1 estimated), the Cloud SQL peering (2 h against 0.5) and debugging the duplicated events (3 h not foreseen). It matches the warning in section 14 almost exactly.

Conclusion

Your system exists, works and rebuilds itself from scratch.

You know why the build order is what it is: from the irreversible to the reversible, from what depends on nothing to what depends on everything. The projectId in phase 0 and the dashboard in phase 8, because one can never be changed and the other is rebuilt in ten minutes. And you know each specific dependency: why identity comes before data, why the application is not automated until you have deployed it by hand once, why observability is built on the complete system.

You have phase 0 with the projects created under their definitive names, the APIs enabled in one go, the budget existing before you could spend, and the repository with its structure and its README written at the start and not at the end.

You have the foundation with Terraform: the bootstrap that solves the chicken-and-egg problem, the state bucket with versioning and no public access, the backend with a prefix per environment, the provider pinned by version so that an update on some random Tuesday does not propose destroying half your infrastructure, your own modules reused in dev and prod, variables with exactly the same keys in both, and the two private services access peering resources everybody forgets.

You have identity done properly: one service account per workload, roles scoped to the resource and not to the project — google_secret_manager_secret_iam_member, not google_project_iam_member — Workload Identity Federation with its attribute_condition, without which it accepts tokens from any repository on the planet, and iam.serviceAccountUser accounted for so that the first automated deployment does not steal an afternoon from you.

You have the data with the database with no public IP, the password generated by Terraform and stored in Secret Manager without anybody ever seeing it, buckets with enforced public access prevention and a lifecycle, idempotent migrations versioned in the repository and a reproducible fictional data generator with a fixed seed.

You have the application honouring the Cloud Run contract — $PORT, 0.0.0.0, fast startup, no state — running as non-root, with configuration through variables and secrets by reference, with two probes that deliberately do different things, with cpu_idle so as not to pay between requests, and with structured logs whose logging.googleapis.com/trace key links every line to its trace.

You have automated delivery in the right order — tests before building, smoke after deploying — and the principle that makes it reliable: the same image is promoted to production, without rebuilding, because rebuilding is deploying something you have never tested. And you have verified it in the only way that counts: by breaking a test on purpose and checking that the deployment stops.

You have the data and AI layer with events that carry no personal data, deduplication via row_ids that prevents inflated figures, partitioned tables with a mandatory filter and an AI component based on a pre-trained API that analyses once and stores the result.

You have exposure with your own domain and valid TLS, knowing how to choose between the zero-cost route and the complete one with CDN and WAF, and with rate limiting understood as what it also is: budget protection.

And you have observability with the four signals dashboard, the alert with its first steps written in the documentation block itself, the uptime check, the SLO calculating its error budget, and the alert tested by provoking it, which is the only way to discover an unverified notification channel before you need it.

On top of all that you have the cross-cutting practices: small commits, the plan always read — with must be replaced as the alarm phrase — nothing created by hand and what to do when you do it anyway, documentation on the fly, spending checked twice a week with the Friday destroy that saves money and tests your IaC at the same time, and the environments kept aligned.

And you have the five-step method for when you get stuck, the ten most frequent errors with their cause, the two-hour rule, the progress log table with verifiable criteria and the most honest warning of all: the first time, everything takes twice as long, and that time is not wasted, it is exactly the learning.

In the next lesson, 08-04, the question is no longer "does it work?" but "how do I know?". You are going to build your project's testing pyramid, test the infrastructure by recreating it from scratch — the definitive proof that your IaC is real — run through a security checklist, do an honest and cheap load test, switch off a dependency on purpose to see what happens, time a restore, rehearse a rollback, and go through your pre-launch checklist before declaring the project live.

Google Cloud Platform (GCP) Course

Module 1: Introduction to Google Cloud Platform

Module 2: Core GCP Services

Module 3: Networking and Security

Module 4: Data and Analytics

Module 5: Machine Learning and AI

Module 6: DevOps and Monitoring

Module 7: Advanced GCP Topics

Module 8: Final Project

© Copyright 2026. All rights reserved