The previous lesson ended with an uncomfortable question. You have built a continuous integration pipeline that fires on every push, that tags images with the commit SHA and that publishes to Artifact Registry everything that passes the tests. And from exactly which repository does it fire?

Because the real inventory of AlpinaShop's code at this moment is as follows. The catalogue website is in a local Git repository on Dani's laptop, with a remote pointing at a server the company stopped maintaining two years ago; it works because nobody runs git push. The startup scripts of the MIG's instance templates are pasted into the template's metadata field, and nowhere else. The Kubernetes manifests for the tienda namespace are in a shared Drive folder called "GKE definitive (v3) FINAL". Lucía's notebooks live on her Workbench instance. And the gcloud commands with which Marta created the VPC, the load balancer and the Cloud Armor rules are nowhere at all, except in her terminal history, which truncates at 10,000 commands.

Nothing built on top of that — no CI/CD, no infrastructure as code, no change review — is possible without solving this first. This lesson solves it, and it does so with a large dose of honesty about which tool you should really be using in 2026.

Contents

  1. Why version control is the foundation of everything
  2. What Cloud Source Repositories is and how it is used
  3. The honest conversation: CSR is deprecated
  4. Connecting GitHub or GitLab with GCP
  5. Workload Identity Federation: accessing GCP without JSON keys
  6. Decision table: CSR, GitHub or GitLab
  7. Branching strategy for a small team
  8. What should and should not be in the repository
  9. AlpinaShop's repository structure
  10. Code review and branch protection
  11. Commit messages and semantic versioning
  12. Quality hooks before the commit
  13. Connecting the repositories with Cloud Build triggers

  1. Why version control is the foundation of everything

It is worth stating explicitly what a version control system provides, because once you have been using one for years you forget that it solves several different problems at once:

Capability What it enables What breaks without it at AlpinaShop
History Knowing what changed, when and why Nobody knows why the catalogue's timeout is 45 s
Attribution Knowing who made each change Faced with a failure, there is nobody to ask
Rollback Going back to a known previous state The only way back is to rewrite it by hand
Branching Working in parallel without treading on each other Dani and Marta pass files to each other over Slack
Review Having someone else look before merging Nobody reviews anything
Immutable identity A SHA identifies an exact state The $COMMIT_SHA from 06-01 does not exist
Automation Reacting to repository events No triggers are possible

Look at the last two rows, because they are the ones that connect with the previous lesson. The Cloud Build pipeline tags every image with $COMMIT_SHA, and that tag is what makes it possible to know what code is in production and to roll back. With no remote repository, there is no SHA, no trigger and no pipeline. Everything built in 06-01 depends on solving this.

And there is a cultural aspect that matters as much as the technical one: version control turns code from personal property into team property. As long as the catalogue is on Dani's laptop, it is Dani's — with everything that implies when he goes on holiday, when he falls ill or when he changes jobs. As soon as it is in a shared repository with history and review, it belongs to AlpinaShop.

  1. What Cloud Source Repositories is and how it is used

Cloud Source Repositories (CSR) is the private Git repository service hosted on Google Cloud. A CSR repository is a perfectly ordinary Git repository: clone, commit, push, pull, branches and tags all work exactly the same. What it adds is its integration with the rest of GCP.

Creating and cloning a repository:

# Create the repository inside the CI/CD project
gcloud source repos create alpinashop-catalogo --project=alpinashop-cicd

# Clone it: gcloud configures authentication for you
gcloud source repos clone alpinashop-catalogo --project=alpinashop-cicd
cd alpinashop-catalogo

# From here on it is plain Git
git add .
git commit -m "Initial import of the catalogue website"
git push origin main

The gcloud source repos clone does something more than a git clone: it configures a credential helper that uses your gcloud identity to authenticate you. There are no passwords or SSH keys to manage; access is controlled with IAM, just like any other GCP resource:

# Dani can read from and write to the repository
gcloud projects add-iam-policy-binding alpinashop-cicd \
  --member='group:[email protected]' \
  --role=roles/source.writer

# The data team only reads
gcloud projects add-iam-policy-binding alpinashop-cicd \
  --member='group:[email protected]' \
  --role=roles/source.reader

If you prefer to authenticate without gcloud — for example from a machine that does not have it — there are manually generated credentials and SSH access with a key registered in your profile.

What CSR really brings, and what explains why it existed:

  • Native integration with Cloud Build: the triggers from 06-01 work without connecting anything external or installing third-party applications.
  • Permissions are IAM: the same groups, roles and conditions from 03-04 govern access to the code. There is no parallel permission system to keep in sync.
  • Code search across all the project's repositories from the console.
  • Integration with Error Reporting and Cloud Debugger: when an exception is grouped in Error Reporting, the console can show the specific line of code from the repository.
  • Automatic mirroring from GitHub or Bitbucket, to keep a copy inside GCP.
  • The data does not leave your GCP organization, which in some regulated contexts simplifies a conversation with the legal department.

All of that is real and it works. And even so, it is not what AlpinaShop is going to use.

  1. The honest conversation: CSR is deprecated

It has to be said clearly, because the name of this lesson carries the service in its title and it would be easy to give the wrong impression:

Cloud Source Repositories is deprecated. Since 2024 it has not been enabled for new customers who had not used it previously, it receives no new functionality, and Google explicitly recommends using GitHub, GitLab or another external provider connected to Cloud Build.

Existing repositories keep working and no shutdown date has been announced. But you should not start a new project with CSR in 2026, and it is worth understanding why it disappeared, because the reason is interesting and it is not technical.

CSR hosted Git perfectly. What it never had was the rest: code review with inline comments, pull request templates, discussions, issue tracking, a wiki, integration with the ecosystem of tools everybody uses. And it turns out that the repository is not the product: the product is the collaboration flow around the repository. Google competed on the easy part — storing Git objects — against platforms that had built the hard part.

That is why the dominant pattern today is: the code in GitHub or GitLab, the execution in GCP. And that is why this lesson devotes the rest of its space to getting that connection right.

So why is CSR still worth knowing about? For three practical reasons: because you will come across it in legacy infrastructures and will have to operate or migrate it; because its IAM-based permission model is still an excellent idea worth understanding; and because mirroring from GitHub is still useful when somebody demands a copy of the code inside the GCP organization.

If you have to migrate from CSR to GitHub, the procedure is short because Git is Git:

# Full clone with every branch and tag
git clone --mirror https://source.developers.google.com/p/alpinashop-cicd/r/alpinashop-catalogo
cd alpinashop-catalogo.git

# Push everything to the new destination
git remote set-url origin [email protected]:alpinashop/alpinashop-catalogo.git
git push --mirror origin

What that command does not migrate: the IAM permissions (they have to be recreated on the new provider) and the Cloud Build triggers, which have to be recreated pointing at the new connection.

  1. Connecting GitHub or GitLab with GCP

AlpinaShop chooses GitHub, and it needs Cloud Build to react to its events. The connection has evolved and in 2026 you should use the modern form.

There are two generations of repository connection, and the difference matters:

Aspect 1st generation (GitHub application) 2nd generation (Cloud Build repositories)
Mechanism GitHub application installed in the organization Managed Connection + Repository resource
Management Console only Console, gcloud and Terraform
Credentials Managed by the application Token in Secret Manager, controlled by you
Repositories Linked one by one from the console Linkable via API, in bulk
Providers GitHub, GitHub Enterprise GitHub, GitHub Enterprise, GitLab, Bitbucket
Recommendation Legacy The one to use

The second generation is the right one because linking the repository stops being an irreproducible click in a console and becomes a declarable resource — which fits with everything coming in 06-05 and 06-07.

The procedure, step by step:

# 1. Enable the required APIs
gcloud services enable cloudbuild.googleapis.com secretmanager.googleapis.com \
  --project=alpinashop-cicd

# 2. Store the GitHub personal access token in Secret Manager
#    (with permissions: repo, read:user, read:org — nothing more)
printf 'ghp_XXXXXXXXXXXXXXXXXXXX' | gcloud secrets create github-token-alpinashop \
  --data-file=- --project=alpinashop-cicd

# 3. Give the Cloud Build service account access to the secret
NUM=$(gcloud projects describe alpinashop-cicd --format='value(projectNumber)')
gcloud secrets add-iam-policy-binding github-token-alpinashop \
  --member="serviceAccount:service-${NUM}@gcp-sa-cloudbuild.iam.gserviceaccount.com" \
  --role=roles/secretmanager.secretAccessor --project=alpinashop-cicd

# 4. Create the connection (requires an interactive authorisation in the browser)
gcloud builds connections create github alpinashop-github \
  --region=europe-west1 --project=alpinashop-cicd

# 5. Link the specific repository
gcloud builds repositories create alpinashop-catalogo \
  --remote-uri=https://github.com/alpinashop/alpinashop-catalogo.git \
  --connection=alpinashop-github \
  --region=europe-west1 --project=alpinashop-cicd

Look at step 2: the GitHub token goes into Secret Manager, exactly like the database password in 03-06. A token with write permission over all your repositories is a top-tier credential, and it deserves the same treatment: encrypted storage, access granted to a specific identity and periodic rotation. Noting the token's expiry date in the team calendar avoids the classic "the pipeline stopped working on Tuesday and nobody knows why".

With the connection created, a trigger is defined against the Repository resource:

gcloud builds triggers create github \
  --name=catalogo-main \
  --repository=projects/alpinashop-cicd/locations/europe-west1/connections/alpinashop-github/repositories/alpinashop-catalogo \
  --branch-pattern='^main$' \
  --build-config=cloudbuild.yaml \
  --region=europe-west1 \
  --service-account=projects/alpinashop-cicd/serviceAccounts/[email protected] \
  --project=alpinashop-cicd

With GitLab the procedure is analogous: gcloud builds connections create gitlab, with a GitLab access token in Secret Manager. And for self-hosted GitLab behind a firewall there is the connection through a service agent in the VPC, which is exactly the case you need when the Git server is internal.

  1. Workload Identity Federation: accessing GCP without JSON keys

There is a second sense of the connection, and it is the one that causes the most security problems: letting GitHub Actions access GCP resources. AlpinaShop needs it so that a workflow can publish to Artifact Registry or query BigQuery without depending on Cloud Build.

The old way of doing it, and why it is bad:

# DO NOT DO THIS
gcloud iam service-accounts keys create clave.json \
  [email protected]
# ...and paste the contents of clave.json into a GitHub secret

That JSON key is a permanent, non-expiring, transferable credential. Whoever obtains it is that service account, forever, from anywhere in the world. There is no way of knowing whether it has been copied. It leaks into logs, into screenshots, into accidental commits, and rotating it is a manual process nobody performs. Service account JSON keys are, by a wide margin, the most common cause of serious security incidents on GCP.

Workload Identity Federation eliminates the problem at the root. The idea, which connects directly with what you saw in 03-04:

GCP trusts GitHub's token issuer. When a GitHub Action runs, GitHub hands it a short-lived OIDC token that says "I am workflow X of repository Y of organization Z". GCP verifies that token, checks that it matches a condition you have defined, and exchanges it for temporary credentials.

flowchart LR
    A[GitHub Action<br/>running] -->|1. asks for an OIDC token| B[GitHub OIDC<br/>issuer]
    B -->|2. signed token:<br/>repo, branch, workflow| A
    A -->|3. presents the token| C[Workload Identity Pool<br/>in GCP]
    C -->|4. verifies signature<br/>and attribute condition| D{Does it match?}
    D -->|Yes| E[Temporary credential<br/>for the service account]
    D -->|No| F[Rejected]
    E --> G[Access to Artifact Registry,<br/>BigQuery, etc.]

The full configuration:

PROJ=alpinashop-cicd
NUM=$(gcloud projects describe $PROJ --format='value(projectNumber)')

# 1. Create the workload identity pool
gcloud iam workload-identity-pools create github-pool \
  --location=global --display-name="GitHub Actions" --project=$PROJ

# 2. Create the OIDC provider, with the attribute condition
gcloud iam workload-identity-pools providers create-oidc github-provider \
  --location=global --workload-identity-pool=github-pool \
  --issuer-uri="https://token.actions.githubusercontent.com" \
  --attribute-mapping="google.subject=assertion.sub,attribute.repository=assertion.repository,attribute.ref=assertion.ref" \
  --attribute-condition="assertion.repository_owner == 'alpinashop'" \
  --project=$PROJ

# 3. Allow ONLY the catalogue repository to impersonate the service account
gcloud iam service-accounts add-iam-policy-binding \
  sa-github-catalogo@${PROJ}.iam.gserviceaccount.com \
  --role=roles/iam.workloadIdentityUser \
  --member="principalSet://iam.googleapis.com/projects/${NUM}/locations/global/workloadIdentityPools/github-pool/attribute.repository/alpinashop/alpinashop-catalogo" \
  --project=$PROJ

Step 3 is the one you need to understand properly. The principalSet with attribute.repository/alpinashop/alpinashop-catalogo means that only workflows from that specific repository can obtain credentials for that service account. A workflow from another repository in the same organization presents a valid token but with a different repository, it does not match the principalSet, and it is rejected.

If you also want to restrict by branch — so that only main can deploy — you use attribute.ref:

principalSet://.../workloadIdentityPools/github-pool/attribute.ref/refs/heads/main

And this is how it is consumed from the workflow:

# .github/workflows/publicar.yml
name: Publish image
on:
  push:
    branches: [main]

permissions:
  contents: read
  id-token: write        # essential: allows requesting the OIDC token

jobs:
  publicar:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4

      - id: auth
        uses: google-github-actions/auth@v2
        with:
          workload_identity_provider: 'projects/123456789/locations/global/workloadIdentityPools/github-pool/providers/github-provider'
          service_account: '[email protected]'
          # Notice: there is NO key, NO secret, NO JSON

      - uses: google-github-actions/setup-gcloud@v2
      - run: |
          gcloud auth configure-docker europe-west1-docker.pkg.dev --quiet
          docker build -t europe-west1-docker.pkg.dev/alpinashop-prod/alpinashop/catalogo:${GITHUB_SHA} .
          docker push europe-west1-docker.pkg.dev/alpinashop-prod/alpinashop/catalogo:${GITHUB_SHA}

The id-token: write line in permissions is mandatory and it gets forgotten constantly; without it, GitHub does not issue the OIDC token and authentication fails with a message that does not help much.

Aspect JSON key Workload Identity Federation
Expiry None Minutes
If it leaks Permanent access Useless out of context
Rotation Manual, nobody does it Automatic by design
Scope Anyone who holds it Only the declared repository/branch
Auditing "Account X did something" Which repository and which workflow
Maintenance work High None

The practical rule: if you are creating a service account JSON key in 2026, there is almost always a better alternative. Inside GCP, service accounts attached to the resource. Outside GCP, identity federation. JSON keys are left for residual cases, and in those cases they must have expiry, rotation and an organization policy limiting where they can be created — a subject for 07-07.

  1. Decision table: CSR, GitHub or GitLab

The comparison with no embellishment, which is what AlpinaShop needs in order to decide:

Criterion Cloud Source Repositories GitHub GitLab
Status Deprecated, no new features Active, market leader Active, strong on self-hosting
Cost Free up to a low storage and data limit Free for private repos; paid for advanced features The same; free self-hostable community edition
Code review Practically non-existent Mature pull requests, suggestions, code owners Merge requests, very complete
Its own CI/CD No, depends on Cloud Build GitHub Actions, huge ecosystem GitLab CI, very powerful and integrated
GCP integration Native, IAM permissions Excellent (2nd gen + federation) Very good (2nd gen + federation)
Ecosystem and templates None The largest by far Broad
Issues and projects No Yes Yes, DevSecOps included
Self-hosted Not applicable GitHub Enterprise Server Its strong point
Vendor lock-in High, with Google Medium (Git is portable) Low, it can be self-hosted
Hiring talent Nobody knows it Everybody knows it Very well known

AlpinaShop's decision is GitHub, for four reasons ordered by weight:

  1. Code review is the main goal. AlpinaShop's problem is not where to store bytes: it is that nobody looks at changes before they reach production. GitHub solves that; CSR does not.
  2. Anyone who joins already knows how to use it. With a technical team of three, training time counts.
  3. GCP integration is first class thanks to the 2nd generation and identity federation. CSR's supposed advantage has ceased to be one.
  4. Lock-in really is low. A Git repository is portable: git clone --mirror and you are somewhere else. What does tie you down are the issues, the pull requests and the Actions workflows, and that is why AlpinaShop's main pipeline lives in cloudbuild.yaml — a neutral file — rather than in GitHub Actions.

When you would choose GitLab instead: if you need to self-host for regulatory requirements, if you want integrated CI/CD without depending on Cloud Build, or if you value having issues, code, CI and a container registry in a single product. It is a perfectly defensible choice.

When you would stay with CSR: only if you already use it and it works. Migrating for the sake of migrating is not a priority; migrating when you need serious code review is.

  1. Branching strategy for a small team

With the repository decided, the next question is how the work is organised inside it. There are two dominant schools.

GitFlow defines long-lived branches: main (what is in production), develop (integration), feature/*, release/* and hotfix/*. A change is born in a feature branch, is merged into develop, enters a release branch, is stabilised and finally reaches main.

Trunk-based defines a single long-lived branch, main, always deployable. Changes are made in short branches — hours, or a couple of days at most — that are merged into main as soon as they pass review and CI.

Aspect GitFlow Trunk-based
Long-lived branches main and develop main only
Life of a working branch Days or weeks Hours or a couple of days
Merge pain High: they diverge a lot Low: they diverge little
Fits with CI Badly: integration happens late It is the premise of CI
Complexity High Low
Suited to Packaged releases, several maintained at once Web services with continuous deployment
Half-finished features A long branch left unmerged Feature flags

AlpinaShop chooses trunk-based, and the reasons are concrete:

  • It is a web service, not a packaged product. There are no customers on version 2.3 needing a patch while 3.0 is being developed. There is only one version: the one in alpinashop-prod.
  • The team is three people. The coordination that justifies GitFlow does not exist when you all fit on one call.
  • It is consistent with the CI from 06-01. The "I" in continuous integration means integrating continuously; a three-week branch is exactly the opposite, and the pipeline you built only delivers value if changes reach main often.
  • Merge conflicts grow non-linearly with time. Two one-day branches almost never clash; two three-week branches always do, and resolving that conflict is where the bugs slip in.

The resulting daily workflow:

git switch main && git pull                         # 1. Always start from an up-to-date main
git switch -c feat/brand-filter                     # 2. Short branch with a clear name
# ... work, small commits ...
git push -u origin feat/brand-filter                # 3. Publish and open the pull request
# 4. The CI from 06-01 runs on its own against the PR
# 5. Marta or Dani review it
# 6. Squash merge into main → it is deployed to alpinashop-dev
git switch main && git pull
git branch -d feat/brand-filter                     # 7. Delete the branch

The obvious objection: what if a feature takes three weeks? The trunk-based answer is not "keep the branch for three weeks", but merge incomplete but inactive code, protected by a switch:

# catalogo/config.py — the switch is read from the environment, not from the code
NEW_SEARCH = os.environ.get("FEATURE_NEW_SEARCH", "false").lower() == "true"

# catalogo/rutas.py
@app.get("/buscar")
def buscar():
    if NEW_SEARCH:
        return semantic_search(request.args["q"])    # 05-06, under test
    return classic_search(request.args["q"])         # what works today

The new code reaches main and production from day one, switched off. It is switched on in alpinashop-dev to test it, then for a percentage of users, and finally for everyone. And if something goes wrong, switching it off means changing an environment variable, not reverting a deployment. It is a rollback measured in seconds.

The honest trade-off: switches accumulate complexity and dead branches in the code. The discipline that goes with them is deleting them as soon as the feature is consolidated, with an expiry date noted in the very pull request that introduces them.

  1. What should and should not be in the repository

A useful rule as a starting point: if it is text describing how the system works or how it is built, it goes into the repository.

Does go into the repository Why
Application code Obvious
Dockerfile and .dockerignore The image must be reproducible from the code
cloudbuild.yaml and variants The pipeline is reviewed like the code (06-01)
Kubernetes manifests Today they are in Drive; there they are neither reviewed nor versioned
Terraform (.tf) Infrastructure as code arrives in 06-07
Tests Without them CI checks nothing
requirements.txt with pinned versions Reproducibility of dependencies
Database migrations Reviewable, ordered schema changes
Data and ML pipeline definitions The KFP components from 05-07
README and architecture decisions DA-001, DA-002 and DA-003 have to be written down
Non-sensitive per-environment configuration dev.tfvars, prod.tfvars
Does not go into the repository Where it goes Why
Passwords, tokens, API keys Secret Manager (03-06) Git never forgets
Service account JSON keys They should not exist (section 5) —
Certificates and private keys Secret Manager or KMS Same as above
Terraform state (.tfstate) Cloud Storage bucket (06-07) It contains sensitive values
Customer data, database dumps Cloud Storage with access control GDPR
Compiled artefacts, images Artifact Registry The repository is not a binary store
Large files (models, CSVs) Cloud Storage or Git LFS They bloat the repository forever
.env with real values Environment variables or Secret Manager It leaks when you run git add .

The key sentence about secrets deserves underlining: Git does not forget. If you upload a password and delete it in the next commit, it is still in the history, in every clone anybody has made and in every mirror. Rewriting history is painful and never complete. The only correct response to a leaked secret is to rotate it immediately, and only afterwards clean up the history if it is worth it.

A sensible .gitignore is the first line of defence:

# Secrets and credentials
*.env
.env.*
*-key.json
*credentials*.json
*.pem
*.p12

# Terraform state: NEVER into Git
*.tfstate
*.tfstate.*
.terraform/
*.tfvars.secret

# Local environments and artefacts
venv/
__pycache__/
*.pyc
.pytest_cache/
dist/
build/

# Data
*.csv
*.parquet
datos/

But the real defence is automatic, and it is the subject of section 12.

  1. AlpinaShop's repository structure

The monorepo versus polyrepo debate has a lot of literature and little universal answer:

Aspect Monorepo Polyrepo
A change crossing components A single atomic PR Several coordinated PRs
Per-component permissions Hard to scope Natural
CI Needs path filters Simple
Code discovery Everything in plain sight You have to know where to look
Dependencies between parts Easy to share, easy to tangle Explicit boundaries
Tooling required More Less

For a team of three, the complexity of a monorepo with advanced build tooling is not justified. But neither are forty tiny repositories. AlpinaShop settles on the middle ground: four repositories grouped by life cycle and by who touches them.

Repository Contents Who maintains it Pipeline
alpinashop-catalogo Flask, Dockerfile, tienda manifests, cloudbuild.yaml Dani catalogo-main, catalogo-pr
alpinashop-infra Terraform for the VPC, load balancer, IAM, Cloud SQL, buckets Marta plan on PR, apply with approval
alpinashop-datos Dataflow pipelines, SQL queries, DAG/Workflows from 04-06 Lucía Tests + template deployment
alpinashop-ml KFP components, training code, cleaned notebooks Lucía and Dani cloudbuild-ml.yaml from 06-01

The division criterion is not technical, it is one of pace and responsibility. The catalogue changes several times a day and Dani touches it. The infrastructure changes once a week, requires Marta's approval and its pipeline is completely different — a plan that gets reviewed, not tests that pass. Putting them together would force the infrastructure pipeline to run for every change to an HTML template.

And a warning about the infrastructure repository: it is the most dangerous of the four. Whoever can merge a PR there can modify firewall rules, IAM policies and databases. Its branch protection must be the strictest, and its mandatory reviewers must always include somebody from gcp-seguridad@.

  1. Code review and branch protection

The existence of a repository does not stop somebody pushing straight to main without anyone looking at it. What stops that is branch protection.

The rules AlpinaShop enables on main in all four repositories:

Rule Effect Why
Direct push forbidden Every change goes through a pull request Without this, the rest is decorative
At least one approval Another person has looked at the change The main goal of the lesson
CI must be green No merging with red tests It gives meaning to the pipeline from 06-01
Approvals expire when new changes arrive You do not approve one version and merge another Prevents a very common trick
Linear history Merge by squash or rebase Readable history, reverting is trivial
Nobody bypasses the rules, not even administrators No exceptions Exceptions become the norm
Secret scanning Blocks the push if it detects credentials The last net before disaster

In the alpinashop-infra repository one more rule is added: a mandatory reviewer from the security team, by means of the CODEOWNERS file:

# CODEOWNERS for the alpinashop-infra repository
*                       @alpinashop/infraestructura
/iam/                   @alpinashop/seguridad @alpinashop/infraestructura
/red/firewall.tf        @alpinashop/seguridad @alpinashop/infraestructura
/produccion/            @alpinashop/seguridad @alpinashop/infraestructura

Any PR touching IAM, firewall rules or the production directory requires approval from somebody in gcp-seguridad@. It is the code equivalent of the separation of duties from 03-04.

What to look at in a review, in order of real importance: first, is it correct? — does it do what it says, does it handle the edge cases; second, is it secure? — secrets, input validation, permissions being widened; third, is it tested? — is there a test that would fail without this change; fourth, is it understandable? — somebody is going to read it a year from now; and lastly, style, which should be automated and not argued about in the review. That last point saves an enormous amount of friction: if formatting is decided by a tool, nobody argues about commas.

And two rules of coexistence that are worth more than any list: small pull requests — a 40-line PR gets useful comments, a 2,000-line one gets an "lgtm" — and fast reviews, because a PR waiting two days blocks whoever wrote it and accumulates conflicts.

  1. Commit messages and semantic versioning

A commit history of "changes", "fix", "again" and "now it works" is useless. A good history answers why each thing was done, which is what cannot be deduced from the code.

AlpinaShop adopts Conventional Commits, a simple format with a big benefit:

<type>(<scope>): <summary in the imperative, under 72 characters>

<body: WHY it was done, not what was done — that is already in the diff>

<footer: references, breaking changes>
feat(buscador): add a brand filter to the catalogue

18 % of searches included a brand name in the free text,
which produced poor results. The filter narrows the
query to the brand before applying the ranking.

Refs: #142
fix(pedidos): retry the Cloud SQL connection after a network failure

The connection pool did not recover after a brief outage and the
application returned 500s until the pod was restarted. A retry with
exponential backoff is added, consistent with 04-04.

Refs: #158

The usual types: feat (new functionality), fix (correction), docs, refactor, test, chore (maintenance), perf and ci.

The concrete benefit: the format is machine-parseable. From it the changelog and the next version come out automatically, according to semantic versioning (MAJOR.MINOR.PATCH):

Commit type Increment Example
fix: PATCH 2.4.1 → 2.4.2
feat: MINOR 2.4.1 → 2.5.0
BREAKING CHANGE: in the footer or feat!: MAJOR 2.4.1 → 3.0.0

And here there is an honest nuance worth stating: semantic versioning is designed for artefacts that others consume — libraries, public APIs. For AlpinaShop's catalogue website, which is an internal service deployed continuously, the version matters far less than the commit SHA, which is already an exact identifier. AlpinaShop uses semantic tags to mark production candidates — that is the tag trigger from 06-01 — and to have pronounceable names in a meeting, not as a compatibility mechanism.

  1. Quality hooks before the commit

A failure caught in CI costs minutes of waiting. The same failure caught before making the commit costs seconds. Pre-commit hooks run fast checks locally before the commit is created.

# .pre-commit-config.yaml, at the root of each repository
repos:
  - repo: https://github.com/pre-commit/pre-commit-hooks
    rev: v4.6.0
    hooks:
      - id: trailing-whitespace
      - id: end-of-file-fixer
      - id: check-yaml
      - id: check-added-large-files       # stops huge models and CSVs being uploaded
        args: ['--maxkb=500']
      - id: check-merge-conflict

  - repo: https://github.com/astral-sh/ruff-pre-commit
    rev: v0.6.9
    hooks:
      - id: ruff                          # errors and style
        args: [--fix]
      - id: ruff-format                   # automatic formatting

  - repo: https://github.com/gitleaks/gitleaks
    rev: v8.21.0
    hooks:
      - id: gitleaks                      # THE MOST IMPORTANT ONE: looks for secrets
pip install pre-commit
pre-commit install          # installs the hook into .git/hooks
pre-commit run --all-files  # first pass over the whole repository

The gitleaks hook justifies the whole mechanism on its own. It looks for credential patterns — API keys, tokens, private keys, GCP JSON keys — and blocks the commit if it finds any. Remember section 8: Git does not forget. It is infinitely cheaper to block the commit than to rotate a leaked payment gateway key.

Two warnings about local hooks. The first: they are optional by nature, because anybody can bypass them with git commit --no-verify. That is why the same checks must be repeated in CI, where they cannot be dodged. Hooks are a convenience for the developer, not a security control. The second: they must be fast. A hook that takes thirty seconds makes everybody learn to use --no-verify. The slow stuff — the full test suite — goes in CI, not here.

  1. Connecting the repositories with Cloud Build triggers

The close of the lesson is joining what came in 06-01 with what is here. This is AlpinaShop's complete map:

flowchart TD
    A[Dani opens a PR<br/>feat/brand-filter] --> B[Trigger catalogo-pr]
    B --> C[cloudbuild-pr.yaml:<br/>install, pytest, ruff, bandit]
    C --> D{Green?}
    D -->|No| E[PR blocked<br/>by branch protection]
    D -->|Yes| F[Marta reviews]
    F --> G[Squash into main]
    G --> H[Trigger catalogo-main]
    H --> I[cloudbuild.yaml:<br/>build + publish :SHA]
    I --> J[Deployment to<br/>alpinashop-dev]
    J --> K{Tag v2.5.0}
    K --> L[Promotion trigger<br/>with approval]
    L --> M[alpinashop-prod]

The four triggers AlpinaShop leaves configured:

Trigger Repository Event Configuration What it does
catalogo-pr alpinashop-catalogo PR against main cloudbuild-pr.yaml Tests, does not deploy
catalogo-main alpinashop-catalogo Push to main cloudbuild.yaml Builds, publishes, deploys to dev
catalogo-prod alpinashop-catalogo Tag ^v\d+\.\d+\.\d+$ cloudbuild-prod.yaml, --require-approval Promotes to production
ml-main alpinashop-ml Push to main cloudbuild-ml.yaml Compiles and launches the Vertex AI pipeline

And a final check worth making explicit: the production trigger fires when a tag is created, so whoever can create tags can start a deployment to production. Tag protection is as necessary as branch protection:

# Tag protection rule in GitHub
Pattern: v*
Who can create: only the @alpinashop/infraestructura team

Without that rule, the manual approval from 06-01 still protects the deployment — nobody reaches production without Marta approving — but anybody could fill the queue with pending builds. With it, the path to production is fenced off from beginning to end.

Common Mistakes and Tips

Starting a new project with Cloud Source Repositories in 2026. It is deprecated, it receives no new functionality and it lacks code review, which is exactly what is needed. Know it in order to operate legacy systems; choose GitHub or GitLab for anything new.

Creating service account JSON keys for CI. Permanent, transferable credentials, the most common cause of serious incidents on GCP. Use Workload Identity Federation. And if you find old JSON keys around the infrastructure, plan their removal as if it were security debt, because it is.

Forgetting id-token: write in the GitHub Actions workflow permissions. Without that line the OIDC token is not issued and authentication fails with a confusing error. It is everybody's first stumble.

Feature branches that live for weeks. Every day a branch lives increases the probability of a painful conflict and delays the CI signal. Short branches and feature flags for what is not ready.

Uploading a secret and "fixing it" by deleting it in the next commit. It is still in the history and in every clone. Rotate the secret immediately; cleaning up the history is secondary.

Relying on pre-commit hooks alone. They are bypassed with --no-verify. Always repeat the checks in CI, where they are mandatory.

Enormous pull requests. A 2,000-line PR does not get reviewed, it gets approved. Split it into small changes with one intention each; the quality of the comments is inversely proportional to the size of the diff.

Arguing about style in reviews. Automate formatting with ruff-format or equivalent and devote the review to what a machine cannot assess: whether the approach is correct and whether it is secure.

Not protecting tags. If the production deployment is triggered by a tag, whoever creates tags starts deployments. Protect them just like branches.

A final tip: do the migration in parts. Do not try to move all four repositories to GitHub on the same day. Start with the catalogue, which is the one touched most, set up its pipeline, check that the team has adapted, and then go for the infrastructure. And start by getting the Kubernetes manifests out of Drive: it is the half hour of work with the best benefit/cost ratio in the whole module.

Exercises

Exercise 1: rescue the scattered code

Draw up a concrete three-phase plan to bring all of AlpinaShop's scattered assets under version control: Dani's local repository, the startup scripts in the instance template metadata, the Kubernetes manifests in Drive, Lucía's notebooks and Marta's gcloud commands. For each asset, state which repository it goes to, what needs to be checked before uploading it and what specific risk that upload carries.

Exercise 2: configure GitHub Actions access with least privilege

AlpinaShop wants a GitHub Actions workflow in alpinashop-datos that, on push to main, validates the SQL queries by running a dry run against BigQuery in alpinashop-datos and publishes a report to the alpinashop-datalake bucket. Write the Workload Identity Federation configuration and the exact IAM roles, and explain what stops a workflow from the alpinashop-catalogo repository using that same identity.

Exercise 3: choose a branching strategy with an awkward case

AlpinaShop signs a contract with a chain of shops that needs a customised version of the catalogue, with its own design and some different pricing rules, maintained for at least two years alongside the public version. Dani proposes creating a long-lived client-x branch. Analyse the proposal, say whether trunk-based is still valid and propose the solution you would recommend, with its trade-offs.

Solutions

Solution 1

Phase 1 — Rescue what risks being lost (this week).

Asset Destination Check first Risk
Dani's local repository alpinashop-catalogo The full history, looking for secrets High: if there is a password in some old commit, uploading it publishes it to the team
Manifests from Drive alpinashop-catalogo/k8s/ Which version is the one actually applied Medium: there are three versions and nobody knows which one is running
Startup scripts alpinashop-infra/scripts/ Credentials embedded in the script High: startup scripts are a classic place to paste passwords

For Dani's repository, the correct procedure is to scan the full history, not just the current state:

git clone --mirror /path/local/catalogo catalogo-auditoria.git
gitleaks detect --source=catalogo-auditoria.git --report-path=hallazgos.json

If secrets show up, there are two paths. The pragmatic one: rotate every credential found and upload the history as it is, accepting that it contains values that are already invalid. The purist one: rewrite the history with git filter-repo. For AlpinaShop I recommend the pragmatic one — rotating is mandatory either way, and rewriting adds little if the values no longer work — unless there is personal data, where rewriting is indeed necessary.

For the manifests, the essential check is what is really applied, not which file is called FINAL:

kubectl get deployment catalogo-web -n tienda -o yaml > k8s/catalogo-web.actual.yaml

You strip out the fields generated by the cluster (status, metadata.uid, resourceVersion, last-applied-configuration annotations) and that is the version that goes into the repository. The truth is in the cluster, not in Drive.

Phase 2 — Structure and protect (the following week). Create the four repositories, enable branch protection and CODEOWNERS, install the pre-commit hooks with gitleaks and connect the Cloud Build triggers.

Phase 3 — The hard part (the weeks after).

Asset Destination How Risk
Lucía's notebooks alpinashop-ml/notebooks/ Strip outputs with nbstripout; extract the reusable parts into modules High: notebooks often contain customer data in their outputs
Marta's gcloud commands alpinashop-infra/ as Terraform Do not transcribe commands: export the real state, 06-05 and 06-07 Medium: laborious, but nothing is lost

The detail about the notebooks deserves emphasis: a cell run with df.head() over alpinashop_analitica leaves names, email addresses and postal addresses saved inside the .ipynb. Uploading a notebook without cleaning the outputs is a leak of personal data into the repository, with Git forgetting nothing. nbstripout as a pre-commit hook solves it automatically and is non-negotiable in alpinashop-ml.

And the observation about Marta's commands: transcribing them into a script would be the mistake. That history is incomplete and contains commands that were run and later undone. The right thing to do is to export the real configuration of the live infrastructure, which is exactly the procedure in 06-05.

Solution 2

Federation configuration, reusing the pool already created:

PROJ=alpinashop-datos
NUM=$(gcloud projects describe $PROJ --format='value(projectNumber)')

gcloud iam service-accounts create sa-gha-datos \
  --display-name="GitHub Actions - SQL validation" --project=$PROJ

# Only the alpinashop-datos repository, and only the main branch
gcloud iam service-accounts add-iam-policy-binding \
  sa-gha-datos@${PROJ}.iam.gserviceaccount.com \
  --role=roles/iam.workloadIdentityUser \
  --member="principalSet://iam.googleapis.com/projects/${NUM}/locations/global/workloadIdentityPools/github-pool/attribute.repository/alpinashop/alpinashop-datos" \
  --project=$PROJ

Exact roles, applying least privilege for real:

Role Scope Why that one and not another
roles/bigquery.jobUser alpinashop-datos project Allows launching jobs; a dry run does not read data
roles/bigquery.metadataViewer alpinashop_analitica dataset The dry run needs the schemas in order to validate
roles/storage.objectCreator alpinashop-datalake bucket, reports prefix Create, not read or delete

Notice what it does not carry: no bigquery.dataViewer — a dry run validates syntax and schema, it neither runs the query nor reads a single row — and no storage.objectAdmin — creating reports does not require being able to delete existing ones. And the objectCreator scoped to the prefix by means of an IAM condition (03-04) stops a compromised workflow overwriting data lake data.

What stops a workflow from alpinashop-catalogo using this identity: the principalSet is tied to attribute.repository/alpinashop/alpinashop-datos. The OIDC token GitHub issues carries the name of the originating repository, signed, and GitHub does not allow it to be forged — it is issued by GitHub's own service, not by the workflow. A workflow from alpinashop-catalogo would present a token with repository: alpinashop/alpinashop-catalogo, which does not match the principalSet, and GCP would reject the exchange. That is the property that makes federation substantially more secure than a JSON key: the identity is certified by a trusted third party and tied to the execution context, not to possession of a file.

And one further improvement: the provider's attribute-condition (assertion.repository_owner == 'alpinashop') acts as a second barrier at provider level. Without it, any GitHub repository in the world could at least attempt the exchange, and although the principalSet would reject it anyway, it is better to filter it earlier. Defence in depth.

Solution 3

The long-lived branch proposal is a bad one, and it is worth explaining why before proposing an alternative. A client-x branch maintained for two years produces, with complete certainty:

  • Growing divergence: every security fix has to be applied twice, and at some point one will be forgotten. That forgotten one will be exactly the important one.
  • Increasingly expensive merges: six months in, bringing changes from main into client-x is a project in itself.
  • Duplication of everything else: two pipelines, two sets of images, two deployments, two sets of metrics.
  • A system nobody tests as a whole: CI would test main, and client-x would be left with old tests.

But trunk-based is still valid, because the real problem is not one of branching strategy: it is one of architecture. The right question is not "how do I maintain two branches?", but "how do I make a single code base serve two customers?".

The recommended solution: multi-tenancy by configuration. A single main branch, a single image, and the different behaviour resolved at runtime:

# catalogo/inquilino.py
@dataclass(frozen=True)
class Tenant:
    id: str
    theme: str                 # templates and stylesheets
    pricing_rules: str         # pricing strategy to apply
    domain: str

TENANTS = {
    "public":   Tenant("public",   "alpina",  "standard",  "www.alpinashop.example"),
    "client-x": Tenant("client-x", "clientx", "wholesale", "tienda.client-x.example"),
}

def current_tenant() -> Tenant:
    # Resolution by domain: the same deployment serves both
    return TENANTS.get(DOMAINS.get(request.host), TENANTS["public"])

What changes per tenant is data and configuration: the visual theme, the pricing rules, the domain, perhaps the subset of the catalogue. None of that justifies a branch.

Approach Maintenance cost Risk of divergence Code complexity Recommended?
Long-lived branch Very high and growing Certain Low No
Multi-tenancy by configuration Low None Medium Yes
Separate repository and product Very high N/A: they are different products Low Only if they genuinely diverge
Shared core as a library Medium Medium High If there are many customers

The trade-offs, stated honestly. Multi-tenancy is not free: it introduces conditionals into the code, both paths have to be tested, and there is a new and serious risk — a leak between tenants, client X seeing the public data or the other way round. That demands that the tenant identifier be applied in the data access layer and tested explicitly, rather than trusting that every query remembers to filter.

And there is a breaking point that has to be recognised: if in a year's time client X asks for a completely different purchase flow, a catalogue with another structure and a billing process of its own, they are no longer two configurations of the same product but two products. At that moment separation is correct, and it is done as separate products with a common library, not as a branch.

The rule to take away from this: branches serve to separate work in time, not to separate variants in space. When somebody proposes a long-lived branch for a product variant, the signal is that there is an architecture decision waiting to be taken. And the conversation with the sales side matters too: if the contract with client X allows some flexibility in the design and in the pricing rules, multi-tenancy comes almost free; if it demands a completely different product, the real cost of that contract is much higher than it looked when it was signed, and that is worth knowing in advance.

Conclusion

AlpinaShop finally has a foundation to build on.

You know what version control provides — history, attribution, rollback, branching, review, immutable identity and automation — and that the last two are what make everything in 06-01 possible: with no remote repository there is no $COMMIT_SHA, no trigger and no pipeline.

You know Cloud Source Repositories: private Git repositories on GCP, gcloud source repos clone with authentication by gcloud identity, permissions governed by IAM with the same groups from 03-04, code search and native integration with Cloud Build and Error Reporting. And you know the uncomfortable truth: it is deprecated, it takes no new customers, it receives no functionality and it never had code review, which is what is really needed. You know it in order to operate legacy infrastructures and to migrate them with git clone --mirror, not to start anything new.

You know how to connect GitHub or GitLab with GCP using the 2nd generation of Cloud Build repositories — connection and repository as declarable resources, token in Secret Manager — and why it is preferable to the first generation's GitHub application. And above all you know how to use Workload Identity Federation so that GitHub Actions access GCP with no JSON key at all: the pool, the OIDC provider, the attribute-condition, the principalSet tied to a repository and a branch, and the id-token: write everybody forgets. With the rule that sums the section up: if in 2026 you are creating a service account JSON key, there is almost always a better alternative.

You have the honest decision table between CSR, GitHub and GitLab, and AlpinaShop's decision — GitHub — with its four reasons, alongside the cases in which GitLab would be the better choice. You have the branching strategy: trunk-based with short branches, because it is a web service with a single live version, because the team is three people, because it is the very premise of continuous integration and because conflicts grow non-linearly with time. With feature flags as the answer to "this will take three weeks", and with the discipline of deleting them afterwards.

You know what goes into the repository — code, Dockerfile, pipelines, manifests, Terraform, tests, migrations, architecture decisions — and what does not — secrets, Terraform state, data, artefacts, large files —, with the warning that is worth the whole lesson: Git does not forget, and faced with a leaked secret the only correct response is to rotate it. You have AlpinaShop's four-repository structure divided by pace and responsibility, with alpinashop-infra singled out as the most dangerous. You have branch protection with its seven rules, CODEOWNERS demanding a security review for IAM and firewall, conventional commit messages, semantic versioning with its honest nuance, and the pre-commit hooks with gitleaks as a safety net that does not replace CI.

And you have the four triggers connected from beginning to end: a pull request that tests, main that builds and deploys to development, a protected tag that promotes to production with approval, and the ML pipeline compiling itself.

With this, the catalogue builds, tests and deploys itself. But there are still pieces of the system that are neither a web application nor a data pipeline: they are reactions to events. When a message arrives on the imagenes-subidas topic, something has to wake up, call the Vision API, write to BigQuery and go back to sleep. Standing up a virtual machine or a pod for that would be absurd.

In 06-03 come Cloud Functions, and with them the loose end left by module 5 is finally tied off.

Google Cloud Platform (GCP) Course

Module 1: Introduction to Google Cloud Platform

Module 2: Core GCP Services

Module 3: Networking and Security

Module 4: Data and Analytics

Module 5: Machine Learning and AI

Module 6: DevOps and Monitoring

Module 7: Advanced GCP Topics

Module 8: Final Project

© Copyright 2026. All rights reserved