The previous lesson ended with an uncomfortable question. You have built a continuous integration pipeline that fires on every push, that tags images with the commit SHA and that publishes to Artifact Registry everything that passes the tests. And from exactly which repository does it fire?
Because the real inventory of AlpinaShop's code at this moment is as follows. The catalogue website is in a local Git repository on Dani's laptop, with a remote pointing at a server the company stopped maintaining two years ago; it works because nobody runs git push. The startup scripts of the MIG's instance templates are pasted into the template's metadata field, and nowhere else. The Kubernetes manifests for the tienda namespace are in a shared Drive folder called "GKE definitive (v3) FINAL". Lucía's notebooks live on her Workbench instance. And the gcloud commands with which Marta created the VPC, the load balancer and the Cloud Armor rules are nowhere at all, except in her terminal history, which truncates at 10,000 commands.
Nothing built on top of that — no CI/CD, no infrastructure as code, no change review — is possible without solving this first. This lesson solves it, and it does so with a large dose of honesty about which tool you should really be using in 2026.
Contents
- Why version control is the foundation of everything
- What Cloud Source Repositories is and how it is used
- The honest conversation: CSR is deprecated
- Connecting GitHub or GitLab with GCP
- Workload Identity Federation: accessing GCP without JSON keys
- Decision table: CSR, GitHub or GitLab
- Branching strategy for a small team
- What should and should not be in the repository
- AlpinaShop's repository structure
- Code review and branch protection
- Commit messages and semantic versioning
- Quality hooks before the commit
- Connecting the repositories with Cloud Build triggers
- Why version control is the foundation of everything
It is worth stating explicitly what a version control system provides, because once you have been using one for years you forget that it solves several different problems at once:
| Capability | What it enables | What breaks without it at AlpinaShop |
|---|---|---|
| History | Knowing what changed, when and why | Nobody knows why the catalogue's timeout is 45 s |
| Attribution | Knowing who made each change | Faced with a failure, there is nobody to ask |
| Rollback | Going back to a known previous state | The only way back is to rewrite it by hand |
| Branching | Working in parallel without treading on each other | Dani and Marta pass files to each other over Slack |
| Review | Having someone else look before merging | Nobody reviews anything |
| Immutable identity | A SHA identifies an exact state | The $COMMIT_SHA from 06-01 does not exist |
| Automation | Reacting to repository events | No triggers are possible |
Look at the last two rows, because they are the ones that connect with the previous lesson. The Cloud Build pipeline tags every image with $COMMIT_SHA, and that tag is what makes it possible to know what code is in production and to roll back. With no remote repository, there is no SHA, no trigger and no pipeline. Everything built in 06-01 depends on solving this.
And there is a cultural aspect that matters as much as the technical one: version control turns code from personal property into team property. As long as the catalogue is on Dani's laptop, it is Dani's — with everything that implies when he goes on holiday, when he falls ill or when he changes jobs. As soon as it is in a shared repository with history and review, it belongs to AlpinaShop.
- What Cloud Source Repositories is and how it is used
Cloud Source Repositories (CSR) is the private Git repository service hosted on Google Cloud. A CSR repository is a perfectly ordinary Git repository: clone, commit, push, pull, branches and tags all work exactly the same. What it adds is its integration with the rest of GCP.
Creating and cloning a repository:
# Create the repository inside the CI/CD project
gcloud source repos create alpinashop-catalogo --project=alpinashop-cicd
# Clone it: gcloud configures authentication for you
gcloud source repos clone alpinashop-catalogo --project=alpinashop-cicd
cd alpinashop-catalogo
# From here on it is plain Git
git add .
git commit -m "Initial import of the catalogue website"
git push origin mainThe gcloud source repos clone does something more than a git clone: it configures a credential helper that uses your gcloud identity to authenticate you. There are no passwords or SSH keys to manage; access is controlled with IAM, just like any other GCP resource:
# Dani can read from and write to the repository
gcloud projects add-iam-policy-binding alpinashop-cicd \
--member='group:[email protected]' \
--role=roles/source.writer
# The data team only reads
gcloud projects add-iam-policy-binding alpinashop-cicd \
--member='group:[email protected]' \
--role=roles/source.readerIf you prefer to authenticate without gcloud — for example from a machine that does not have it — there are manually generated credentials and SSH access with a key registered in your profile.
What CSR really brings, and what explains why it existed:
- Native integration with Cloud Build: the triggers from 06-01 work without connecting anything external or installing third-party applications.
- Permissions are IAM: the same groups, roles and conditions from 03-04 govern access to the code. There is no parallel permission system to keep in sync.
- Code search across all the project's repositories from the console.
- Integration with Error Reporting and Cloud Debugger: when an exception is grouped in Error Reporting, the console can show the specific line of code from the repository.
- Automatic mirroring from GitHub or Bitbucket, to keep a copy inside GCP.
- The data does not leave your GCP organization, which in some regulated contexts simplifies a conversation with the legal department.
All of that is real and it works. And even so, it is not what AlpinaShop is going to use.
- The honest conversation: CSR is deprecated
It has to be said clearly, because the name of this lesson carries the service in its title and it would be easy to give the wrong impression:
Cloud Source Repositories is deprecated. Since 2024 it has not been enabled for new customers who had not used it previously, it receives no new functionality, and Google explicitly recommends using GitHub, GitLab or another external provider connected to Cloud Build.
Existing repositories keep working and no shutdown date has been announced. But you should not start a new project with CSR in 2026, and it is worth understanding why it disappeared, because the reason is interesting and it is not technical.
CSR hosted Git perfectly. What it never had was the rest: code review with inline comments, pull request templates, discussions, issue tracking, a wiki, integration with the ecosystem of tools everybody uses. And it turns out that the repository is not the product: the product is the collaboration flow around the repository. Google competed on the easy part — storing Git objects — against platforms that had built the hard part.
That is why the dominant pattern today is: the code in GitHub or GitLab, the execution in GCP. And that is why this lesson devotes the rest of its space to getting that connection right.
So why is CSR still worth knowing about? For three practical reasons: because you will come across it in legacy infrastructures and will have to operate or migrate it; because its IAM-based permission model is still an excellent idea worth understanding; and because mirroring from GitHub is still useful when somebody demands a copy of the code inside the GCP organization.
If you have to migrate from CSR to GitHub, the procedure is short because Git is Git:
# Full clone with every branch and tag
git clone --mirror https://source.developers.google.com/p/alpinashop-cicd/r/alpinashop-catalogo
cd alpinashop-catalogo.git
# Push everything to the new destination
git remote set-url origin [email protected]:alpinashop/alpinashop-catalogo.git
git push --mirror originWhat that command does not migrate: the IAM permissions (they have to be recreated on the new provider) and the Cloud Build triggers, which have to be recreated pointing at the new connection.
- Connecting GitHub or GitLab with GCP
AlpinaShop chooses GitHub, and it needs Cloud Build to react to its events. The connection has evolved and in 2026 you should use the modern form.
There are two generations of repository connection, and the difference matters:
| Aspect | 1st generation (GitHub application) | 2nd generation (Cloud Build repositories) |
|---|---|---|
| Mechanism | GitHub application installed in the organization | Managed Connection + Repository resource |
| Management | Console only | Console, gcloud and Terraform |
| Credentials | Managed by the application | Token in Secret Manager, controlled by you |
| Repositories | Linked one by one from the console | Linkable via API, in bulk |
| Providers | GitHub, GitHub Enterprise | GitHub, GitHub Enterprise, GitLab, Bitbucket |
| Recommendation | Legacy | The one to use |
The second generation is the right one because linking the repository stops being an irreproducible click in a console and becomes a declarable resource — which fits with everything coming in 06-05 and 06-07.
The procedure, step by step:
# 1. Enable the required APIs
gcloud services enable cloudbuild.googleapis.com secretmanager.googleapis.com \
--project=alpinashop-cicd
# 2. Store the GitHub personal access token in Secret Manager
# (with permissions: repo, read:user, read:org — nothing more)
printf 'ghp_XXXXXXXXXXXXXXXXXXXX' | gcloud secrets create github-token-alpinashop \
--data-file=- --project=alpinashop-cicd
# 3. Give the Cloud Build service account access to the secret
NUM=$(gcloud projects describe alpinashop-cicd --format='value(projectNumber)')
gcloud secrets add-iam-policy-binding github-token-alpinashop \
--member="serviceAccount:service-${NUM}@gcp-sa-cloudbuild.iam.gserviceaccount.com" \
--role=roles/secretmanager.secretAccessor --project=alpinashop-cicd
# 4. Create the connection (requires an interactive authorisation in the browser)
gcloud builds connections create github alpinashop-github \
--region=europe-west1 --project=alpinashop-cicd
# 5. Link the specific repository
gcloud builds repositories create alpinashop-catalogo \
--remote-uri=https://github.com/alpinashop/alpinashop-catalogo.git \
--connection=alpinashop-github \
--region=europe-west1 --project=alpinashop-cicdLook at step 2: the GitHub token goes into Secret Manager, exactly like the database password in 03-06. A token with write permission over all your repositories is a top-tier credential, and it deserves the same treatment: encrypted storage, access granted to a specific identity and periodic rotation. Noting the token's expiry date in the team calendar avoids the classic "the pipeline stopped working on Tuesday and nobody knows why".
With the connection created, a trigger is defined against the Repository resource:
gcloud builds triggers create github \
--name=catalogo-main \
--repository=projects/alpinashop-cicd/locations/europe-west1/connections/alpinashop-github/repositories/alpinashop-catalogo \
--branch-pattern='^main$' \
--build-config=cloudbuild.yaml \
--region=europe-west1 \
--service-account=projects/alpinashop-cicd/serviceAccounts/[email protected] \
--project=alpinashop-cicdWith GitLab the procedure is analogous: gcloud builds connections create gitlab, with a GitLab access token in Secret Manager. And for self-hosted GitLab behind a firewall there is the connection through a service agent in the VPC, which is exactly the case you need when the Git server is internal.
- Workload Identity Federation: accessing GCP without JSON keys
There is a second sense of the connection, and it is the one that causes the most security problems: letting GitHub Actions access GCP resources. AlpinaShop needs it so that a workflow can publish to Artifact Registry or query BigQuery without depending on Cloud Build.
The old way of doing it, and why it is bad:
# DO NOT DO THIS
gcloud iam service-accounts keys create clave.json \
[email protected]
# ...and paste the contents of clave.json into a GitHub secretThat JSON key is a permanent, non-expiring, transferable credential. Whoever obtains it is that service account, forever, from anywhere in the world. There is no way of knowing whether it has been copied. It leaks into logs, into screenshots, into accidental commits, and rotating it is a manual process nobody performs. Service account JSON keys are, by a wide margin, the most common cause of serious security incidents on GCP.
Workload Identity Federation eliminates the problem at the root. The idea, which connects directly with what you saw in 03-04:
GCP trusts GitHub's token issuer. When a GitHub Action runs, GitHub hands it a short-lived OIDC token that says "I am workflow X of repository Y of organization Z". GCP verifies that token, checks that it matches a condition you have defined, and exchanges it for temporary credentials.
flowchart LR
A[GitHub Action<br/>running] -->|1. asks for an OIDC token| B[GitHub OIDC<br/>issuer]
B -->|2. signed token:<br/>repo, branch, workflow| A
A -->|3. presents the token| C[Workload Identity Pool<br/>in GCP]
C -->|4. verifies signature<br/>and attribute condition| D{Does it match?}
D -->|Yes| E[Temporary credential<br/>for the service account]
D -->|No| F[Rejected]
E --> G[Access to Artifact Registry,<br/>BigQuery, etc.]
The full configuration:
PROJ=alpinashop-cicd
NUM=$(gcloud projects describe $PROJ --format='value(projectNumber)')
# 1. Create the workload identity pool
gcloud iam workload-identity-pools create github-pool \
--location=global --display-name="GitHub Actions" --project=$PROJ
# 2. Create the OIDC provider, with the attribute condition
gcloud iam workload-identity-pools providers create-oidc github-provider \
--location=global --workload-identity-pool=github-pool \
--issuer-uri="https://token.actions.githubusercontent.com" \
--attribute-mapping="google.subject=assertion.sub,attribute.repository=assertion.repository,attribute.ref=assertion.ref" \
--attribute-condition="assertion.repository_owner == 'alpinashop'" \
--project=$PROJ
# 3. Allow ONLY the catalogue repository to impersonate the service account
gcloud iam service-accounts add-iam-policy-binding \
sa-github-catalogo@${PROJ}.iam.gserviceaccount.com \
--role=roles/iam.workloadIdentityUser \
--member="principalSet://iam.googleapis.com/projects/${NUM}/locations/global/workloadIdentityPools/github-pool/attribute.repository/alpinashop/alpinashop-catalogo" \
--project=$PROJStep 3 is the one you need to understand properly. The principalSet with attribute.repository/alpinashop/alpinashop-catalogo means that only workflows from that specific repository can obtain credentials for that service account. A workflow from another repository in the same organization presents a valid token but with a different repository, it does not match the principalSet, and it is rejected.
If you also want to restrict by branch — so that only main can deploy — you use attribute.ref:
And this is how it is consumed from the workflow:
# .github/workflows/publicar.yml
name: Publish image
on:
push:
branches: [main]
permissions:
contents: read
id-token: write # essential: allows requesting the OIDC token
jobs:
publicar:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- id: auth
uses: google-github-actions/auth@v2
with:
workload_identity_provider: 'projects/123456789/locations/global/workloadIdentityPools/github-pool/providers/github-provider'
service_account: '[email protected]'
# Notice: there is NO key, NO secret, NO JSON
- uses: google-github-actions/setup-gcloud@v2
- run: |
gcloud auth configure-docker europe-west1-docker.pkg.dev --quiet
docker build -t europe-west1-docker.pkg.dev/alpinashop-prod/alpinashop/catalogo:${GITHUB_SHA} .
docker push europe-west1-docker.pkg.dev/alpinashop-prod/alpinashop/catalogo:${GITHUB_SHA}The id-token: write line in permissions is mandatory and it gets forgotten constantly; without it, GitHub does not issue the OIDC token and authentication fails with a message that does not help much.
| Aspect | JSON key | Workload Identity Federation |
|---|---|---|
| Expiry | None | Minutes |
| If it leaks | Permanent access | Useless out of context |
| Rotation | Manual, nobody does it | Automatic by design |
| Scope | Anyone who holds it | Only the declared repository/branch |
| Auditing | "Account X did something" | Which repository and which workflow |
| Maintenance work | High | None |
The practical rule: if you are creating a service account JSON key in 2026, there is almost always a better alternative. Inside GCP, service accounts attached to the resource. Outside GCP, identity federation. JSON keys are left for residual cases, and in those cases they must have expiry, rotation and an organization policy limiting where they can be created — a subject for 07-07.
- Decision table: CSR, GitHub or GitLab
The comparison with no embellishment, which is what AlpinaShop needs in order to decide:
| Criterion | Cloud Source Repositories | GitHub | GitLab |
|---|---|---|---|
| Status | Deprecated, no new features | Active, market leader | Active, strong on self-hosting |
| Cost | Free up to a low storage and data limit | Free for private repos; paid for advanced features | The same; free self-hostable community edition |
| Code review | Practically non-existent | Mature pull requests, suggestions, code owners | Merge requests, very complete |
| Its own CI/CD | No, depends on Cloud Build | GitHub Actions, huge ecosystem | GitLab CI, very powerful and integrated |
| GCP integration | Native, IAM permissions | Excellent (2nd gen + federation) | Very good (2nd gen + federation) |
| Ecosystem and templates | None | The largest by far | Broad |
| Issues and projects | No | Yes | Yes, DevSecOps included |
| Self-hosted | Not applicable | GitHub Enterprise Server | Its strong point |
| Vendor lock-in | High, with Google | Medium (Git is portable) | Low, it can be self-hosted |
| Hiring talent | Nobody knows it | Everybody knows it | Very well known |
AlpinaShop's decision is GitHub, for four reasons ordered by weight:
- Code review is the main goal. AlpinaShop's problem is not where to store bytes: it is that nobody looks at changes before they reach production. GitHub solves that; CSR does not.
- Anyone who joins already knows how to use it. With a technical team of three, training time counts.
- GCP integration is first class thanks to the 2nd generation and identity federation. CSR's supposed advantage has ceased to be one.
- Lock-in really is low. A Git repository is portable:
git clone --mirrorand you are somewhere else. What does tie you down are the issues, the pull requests and the Actions workflows, and that is why AlpinaShop's main pipeline lives incloudbuild.yaml— a neutral file — rather than in GitHub Actions.
When you would choose GitLab instead: if you need to self-host for regulatory requirements, if you want integrated CI/CD without depending on Cloud Build, or if you value having issues, code, CI and a container registry in a single product. It is a perfectly defensible choice.
When you would stay with CSR: only if you already use it and it works. Migrating for the sake of migrating is not a priority; migrating when you need serious code review is.
- Branching strategy for a small team
With the repository decided, the next question is how the work is organised inside it. There are two dominant schools.
GitFlow defines long-lived branches: main (what is in production), develop (integration), feature/*, release/* and hotfix/*. A change is born in a feature branch, is merged into develop, enters a release branch, is stabilised and finally reaches main.
Trunk-based defines a single long-lived branch, main, always deployable. Changes are made in short branches — hours, or a couple of days at most — that are merged into main as soon as they pass review and CI.
| Aspect | GitFlow | Trunk-based |
|---|---|---|
| Long-lived branches | main and develop |
main only |
| Life of a working branch | Days or weeks | Hours or a couple of days |
| Merge pain | High: they diverge a lot | Low: they diverge little |
| Fits with CI | Badly: integration happens late | It is the premise of CI |
| Complexity | High | Low |
| Suited to | Packaged releases, several maintained at once | Web services with continuous deployment |
| Half-finished features | A long branch left unmerged | Feature flags |
AlpinaShop chooses trunk-based, and the reasons are concrete:
- It is a web service, not a packaged product. There are no customers on version 2.3 needing a patch while 3.0 is being developed. There is only one version: the one in
alpinashop-prod. - The team is three people. The coordination that justifies GitFlow does not exist when you all fit on one call.
- It is consistent with the CI from 06-01. The "I" in continuous integration means integrating continuously; a three-week branch is exactly the opposite, and the pipeline you built only delivers value if changes reach
mainoften. - Merge conflicts grow non-linearly with time. Two one-day branches almost never clash; two three-week branches always do, and resolving that conflict is where the bugs slip in.
The resulting daily workflow:
git switch main && git pull # 1. Always start from an up-to-date main
git switch -c feat/brand-filter # 2. Short branch with a clear name
# ... work, small commits ...
git push -u origin feat/brand-filter # 3. Publish and open the pull request
# 4. The CI from 06-01 runs on its own against the PR
# 5. Marta or Dani review it
# 6. Squash merge into main → it is deployed to alpinashop-dev
git switch main && git pull
git branch -d feat/brand-filter # 7. Delete the branchThe obvious objection: what if a feature takes three weeks? The trunk-based answer is not "keep the branch for three weeks", but merge incomplete but inactive code, protected by a switch:
# catalogo/config.py — the switch is read from the environment, not from the code
NEW_SEARCH = os.environ.get("FEATURE_NEW_SEARCH", "false").lower() == "true"
# catalogo/rutas.py
@app.get("/buscar")
def buscar():
if NEW_SEARCH:
return semantic_search(request.args["q"]) # 05-06, under test
return classic_search(request.args["q"]) # what works todayThe new code reaches main and production from day one, switched off. It is switched on in alpinashop-dev to test it, then for a percentage of users, and finally for everyone. And if something goes wrong, switching it off means changing an environment variable, not reverting a deployment. It is a rollback measured in seconds.
The honest trade-off: switches accumulate complexity and dead branches in the code. The discipline that goes with them is deleting them as soon as the feature is consolidated, with an expiry date noted in the very pull request that introduces them.
- What should and should not be in the repository
A useful rule as a starting point: if it is text describing how the system works or how it is built, it goes into the repository.
| Does go into the repository | Why |
|---|---|
| Application code | Obvious |
Dockerfile and .dockerignore |
The image must be reproducible from the code |
cloudbuild.yaml and variants |
The pipeline is reviewed like the code (06-01) |
| Kubernetes manifests | Today they are in Drive; there they are neither reviewed nor versioned |
Terraform (.tf) |
Infrastructure as code arrives in 06-07 |
| Tests | Without them CI checks nothing |
requirements.txt with pinned versions |
Reproducibility of dependencies |
| Database migrations | Reviewable, ordered schema changes |
| Data and ML pipeline definitions | The KFP components from 05-07 |
README and architecture decisions |
DA-001, DA-002 and DA-003 have to be written down |
| Non-sensitive per-environment configuration | dev.tfvars, prod.tfvars |
| Does not go into the repository | Where it goes | Why |
|---|---|---|
| Passwords, tokens, API keys | Secret Manager (03-06) | Git never forgets |
| Service account JSON keys | They should not exist (section 5) | — |
| Certificates and private keys | Secret Manager or KMS | Same as above |
Terraform state (.tfstate) |
Cloud Storage bucket (06-07) | It contains sensitive values |
| Customer data, database dumps | Cloud Storage with access control | GDPR |
| Compiled artefacts, images | Artifact Registry | The repository is not a binary store |
| Large files (models, CSVs) | Cloud Storage or Git LFS | They bloat the repository forever |
.env with real values |
Environment variables or Secret Manager | It leaks when you run git add . |
The key sentence about secrets deserves underlining: Git does not forget. If you upload a password and delete it in the next commit, it is still in the history, in every clone anybody has made and in every mirror. Rewriting history is painful and never complete. The only correct response to a leaked secret is to rotate it immediately, and only afterwards clean up the history if it is worth it.
A sensible .gitignore is the first line of defence:
# Secrets and credentials
*.env
.env.*
*-key.json
*credentials*.json
*.pem
*.p12
# Terraform state: NEVER into Git
*.tfstate
*.tfstate.*
.terraform/
*.tfvars.secret
# Local environments and artefacts
venv/
__pycache__/
*.pyc
.pytest_cache/
dist/
build/
# Data
*.csv
*.parquet
datos/But the real defence is automatic, and it is the subject of section 12.
- AlpinaShop's repository structure
The monorepo versus polyrepo debate has a lot of literature and little universal answer:
| Aspect | Monorepo | Polyrepo |
|---|---|---|
| A change crossing components | A single atomic PR | Several coordinated PRs |
| Per-component permissions | Hard to scope | Natural |
| CI | Needs path filters | Simple |
| Code discovery | Everything in plain sight | You have to know where to look |
| Dependencies between parts | Easy to share, easy to tangle | Explicit boundaries |
| Tooling required | More | Less |
For a team of three, the complexity of a monorepo with advanced build tooling is not justified. But neither are forty tiny repositories. AlpinaShop settles on the middle ground: four repositories grouped by life cycle and by who touches them.
| Repository | Contents | Who maintains it | Pipeline |
|---|---|---|---|
alpinashop-catalogo |
Flask, Dockerfile, tienda manifests, cloudbuild.yaml |
Dani | catalogo-main, catalogo-pr |
alpinashop-infra |
Terraform for the VPC, load balancer, IAM, Cloud SQL, buckets | Marta | plan on PR, apply with approval |
alpinashop-datos |
Dataflow pipelines, SQL queries, DAG/Workflows from 04-06 | Lucía | Tests + template deployment |
alpinashop-ml |
KFP components, training code, cleaned notebooks | Lucía and Dani | cloudbuild-ml.yaml from 06-01 |
The division criterion is not technical, it is one of pace and responsibility. The catalogue changes several times a day and Dani touches it. The infrastructure changes once a week, requires Marta's approval and its pipeline is completely different — a plan that gets reviewed, not tests that pass. Putting them together would force the infrastructure pipeline to run for every change to an HTML template.
And a warning about the infrastructure repository: it is the most dangerous of the four. Whoever can merge a PR there can modify firewall rules, IAM policies and databases. Its branch protection must be the strictest, and its mandatory reviewers must always include somebody from gcp-seguridad@.
- Code review and branch protection
The existence of a repository does not stop somebody pushing straight to main without anyone looking at it. What stops that is branch protection.
The rules AlpinaShop enables on main in all four repositories:
| Rule | Effect | Why |
|---|---|---|
| Direct push forbidden | Every change goes through a pull request | Without this, the rest is decorative |
| At least one approval | Another person has looked at the change | The main goal of the lesson |
| CI must be green | No merging with red tests | It gives meaning to the pipeline from 06-01 |
| Approvals expire when new changes arrive | You do not approve one version and merge another | Prevents a very common trick |
| Linear history | Merge by squash or rebase | Readable history, reverting is trivial |
| Nobody bypasses the rules, not even administrators | No exceptions | Exceptions become the norm |
| Secret scanning | Blocks the push if it detects credentials | The last net before disaster |
In the alpinashop-infra repository one more rule is added: a mandatory reviewer from the security team, by means of the CODEOWNERS file:
# CODEOWNERS for the alpinashop-infra repository * @alpinashop/infraestructura /iam/ @alpinashop/seguridad @alpinashop/infraestructura /red/firewall.tf @alpinashop/seguridad @alpinashop/infraestructura /produccion/ @alpinashop/seguridad @alpinashop/infraestructura
Any PR touching IAM, firewall rules or the production directory requires approval from somebody in gcp-seguridad@. It is the code equivalent of the separation of duties from 03-04.
What to look at in a review, in order of real importance: first, is it correct? — does it do what it says, does it handle the edge cases; second, is it secure? — secrets, input validation, permissions being widened; third, is it tested? — is there a test that would fail without this change; fourth, is it understandable? — somebody is going to read it a year from now; and lastly, style, which should be automated and not argued about in the review. That last point saves an enormous amount of friction: if formatting is decided by a tool, nobody argues about commas.
And two rules of coexistence that are worth more than any list: small pull requests — a 40-line PR gets useful comments, a 2,000-line one gets an "lgtm" — and fast reviews, because a PR waiting two days blocks whoever wrote it and accumulates conflicts.
- Commit messages and semantic versioning
A commit history of "changes", "fix", "again" and "now it works" is useless. A good history answers why each thing was done, which is what cannot be deduced from the code.
AlpinaShop adopts Conventional Commits, a simple format with a big benefit:
<type>(<scope>): <summary in the imperative, under 72 characters> <body: WHY it was done, not what was done — that is already in the diff> <footer: references, breaking changes>
feat(buscador): add a brand filter to the catalogue 18 % of searches included a brand name in the free text, which produced poor results. The filter narrows the query to the brand before applying the ranking. Refs: #142
fix(pedidos): retry the Cloud SQL connection after a network failure The connection pool did not recover after a brief outage and the application returned 500s until the pod was restarted. A retry with exponential backoff is added, consistent with 04-04. Refs: #158
The usual types: feat (new functionality), fix (correction), docs, refactor, test, chore (maintenance), perf and ci.
The concrete benefit: the format is machine-parseable. From it the changelog and the next version come out automatically, according to semantic versioning (MAJOR.MINOR.PATCH):
| Commit type | Increment | Example |
|---|---|---|
fix: |
PATCH | 2.4.1 → 2.4.2 |
feat: |
MINOR | 2.4.1 → 2.5.0 |
BREAKING CHANGE: in the footer or feat!: |
MAJOR | 2.4.1 → 3.0.0 |
And here there is an honest nuance worth stating: semantic versioning is designed for artefacts that others consume — libraries, public APIs. For AlpinaShop's catalogue website, which is an internal service deployed continuously, the version matters far less than the commit SHA, which is already an exact identifier. AlpinaShop uses semantic tags to mark production candidates — that is the tag trigger from 06-01 — and to have pronounceable names in a meeting, not as a compatibility mechanism.
- Quality hooks before the commit
A failure caught in CI costs minutes of waiting. The same failure caught before making the commit costs seconds. Pre-commit hooks run fast checks locally before the commit is created.
# .pre-commit-config.yaml, at the root of each repository
repos:
- repo: https://github.com/pre-commit/pre-commit-hooks
rev: v4.6.0
hooks:
- id: trailing-whitespace
- id: end-of-file-fixer
- id: check-yaml
- id: check-added-large-files # stops huge models and CSVs being uploaded
args: ['--maxkb=500']
- id: check-merge-conflict
- repo: https://github.com/astral-sh/ruff-pre-commit
rev: v0.6.9
hooks:
- id: ruff # errors and style
args: [--fix]
- id: ruff-format # automatic formatting
- repo: https://github.com/gitleaks/gitleaks
rev: v8.21.0
hooks:
- id: gitleaks # THE MOST IMPORTANT ONE: looks for secretspip install pre-commit
pre-commit install # installs the hook into .git/hooks
pre-commit run --all-files # first pass over the whole repositoryThe gitleaks hook justifies the whole mechanism on its own. It looks for credential patterns — API keys, tokens, private keys, GCP JSON keys — and blocks the commit if it finds any. Remember section 8: Git does not forget. It is infinitely cheaper to block the commit than to rotate a leaked payment gateway key.
Two warnings about local hooks. The first: they are optional by nature, because anybody can bypass them with git commit --no-verify. That is why the same checks must be repeated in CI, where they cannot be dodged. Hooks are a convenience for the developer, not a security control. The second: they must be fast. A hook that takes thirty seconds makes everybody learn to use --no-verify. The slow stuff — the full test suite — goes in CI, not here.
- Connecting the repositories with Cloud Build triggers
The close of the lesson is joining what came in 06-01 with what is here. This is AlpinaShop's complete map:
flowchart TD
A[Dani opens a PR<br/>feat/brand-filter] --> B[Trigger catalogo-pr]
B --> C[cloudbuild-pr.yaml:<br/>install, pytest, ruff, bandit]
C --> D{Green?}
D -->|No| E[PR blocked<br/>by branch protection]
D -->|Yes| F[Marta reviews]
F --> G[Squash into main]
G --> H[Trigger catalogo-main]
H --> I[cloudbuild.yaml:<br/>build + publish :SHA]
I --> J[Deployment to<br/>alpinashop-dev]
J --> K{Tag v2.5.0}
K --> L[Promotion trigger<br/>with approval]
L --> M[alpinashop-prod]
The four triggers AlpinaShop leaves configured:
| Trigger | Repository | Event | Configuration | What it does |
|---|---|---|---|---|
catalogo-pr |
alpinashop-catalogo |
PR against main |
cloudbuild-pr.yaml |
Tests, does not deploy |
catalogo-main |
alpinashop-catalogo |
Push to main |
cloudbuild.yaml |
Builds, publishes, deploys to dev |
catalogo-prod |
alpinashop-catalogo |
Tag ^v\d+\.\d+\.\d+$ |
cloudbuild-prod.yaml, --require-approval |
Promotes to production |
ml-main |
alpinashop-ml |
Push to main |
cloudbuild-ml.yaml |
Compiles and launches the Vertex AI pipeline |
And a final check worth making explicit: the production trigger fires when a tag is created, so whoever can create tags can start a deployment to production. Tag protection is as necessary as branch protection:
# Tag protection rule in GitHub Pattern: v* Who can create: only the @alpinashop/infraestructura team
Without that rule, the manual approval from 06-01 still protects the deployment — nobody reaches production without Marta approving — but anybody could fill the queue with pending builds. With it, the path to production is fenced off from beginning to end.
Common Mistakes and Tips
Starting a new project with Cloud Source Repositories in 2026. It is deprecated, it receives no new functionality and it lacks code review, which is exactly what is needed. Know it in order to operate legacy systems; choose GitHub or GitLab for anything new.
Creating service account JSON keys for CI. Permanent, transferable credentials, the most common cause of serious incidents on GCP. Use Workload Identity Federation. And if you find old JSON keys around the infrastructure, plan their removal as if it were security debt, because it is.
Forgetting id-token: write in the GitHub Actions workflow permissions. Without that line the OIDC token is not issued and authentication fails with a confusing error. It is everybody's first stumble.
Feature branches that live for weeks. Every day a branch lives increases the probability of a painful conflict and delays the CI signal. Short branches and feature flags for what is not ready.
Uploading a secret and "fixing it" by deleting it in the next commit. It is still in the history and in every clone. Rotate the secret immediately; cleaning up the history is secondary.
Relying on pre-commit hooks alone. They are bypassed with --no-verify. Always repeat the checks in CI, where they are mandatory.
Enormous pull requests. A 2,000-line PR does not get reviewed, it gets approved. Split it into small changes with one intention each; the quality of the comments is inversely proportional to the size of the diff.
Arguing about style in reviews. Automate formatting with ruff-format or equivalent and devote the review to what a machine cannot assess: whether the approach is correct and whether it is secure.
Not protecting tags. If the production deployment is triggered by a tag, whoever creates tags starts deployments. Protect them just like branches.
A final tip: do the migration in parts. Do not try to move all four repositories to GitHub on the same day. Start with the catalogue, which is the one touched most, set up its pipeline, check that the team has adapted, and then go for the infrastructure. And start by getting the Kubernetes manifests out of Drive: it is the half hour of work with the best benefit/cost ratio in the whole module.
Exercises
Exercise 1: rescue the scattered code
Draw up a concrete three-phase plan to bring all of AlpinaShop's scattered assets under version control: Dani's local repository, the startup scripts in the instance template metadata, the Kubernetes manifests in Drive, Lucía's notebooks and Marta's gcloud commands. For each asset, state which repository it goes to, what needs to be checked before uploading it and what specific risk that upload carries.
Exercise 2: configure GitHub Actions access with least privilege
AlpinaShop wants a GitHub Actions workflow in alpinashop-datos that, on push to main, validates the SQL queries by running a dry run against BigQuery in alpinashop-datos and publishes a report to the alpinashop-datalake bucket. Write the Workload Identity Federation configuration and the exact IAM roles, and explain what stops a workflow from the alpinashop-catalogo repository using that same identity.
Exercise 3: choose a branching strategy with an awkward case
AlpinaShop signs a contract with a chain of shops that needs a customised version of the catalogue, with its own design and some different pricing rules, maintained for at least two years alongside the public version. Dani proposes creating a long-lived client-x branch. Analyse the proposal, say whether trunk-based is still valid and propose the solution you would recommend, with its trade-offs.
Solutions
Solution 1
Phase 1 — Rescue what risks being lost (this week).
| Asset | Destination | Check first | Risk |
|---|---|---|---|
| Dani's local repository | alpinashop-catalogo |
The full history, looking for secrets | High: if there is a password in some old commit, uploading it publishes it to the team |
| Manifests from Drive | alpinashop-catalogo/k8s/ |
Which version is the one actually applied | Medium: there are three versions and nobody knows which one is running |
| Startup scripts | alpinashop-infra/scripts/ |
Credentials embedded in the script | High: startup scripts are a classic place to paste passwords |
For Dani's repository, the correct procedure is to scan the full history, not just the current state:
git clone --mirror /path/local/catalogo catalogo-auditoria.git
gitleaks detect --source=catalogo-auditoria.git --report-path=hallazgos.jsonIf secrets show up, there are two paths. The pragmatic one: rotate every credential found and upload the history as it is, accepting that it contains values that are already invalid. The purist one: rewrite the history with git filter-repo. For AlpinaShop I recommend the pragmatic one — rotating is mandatory either way, and rewriting adds little if the values no longer work — unless there is personal data, where rewriting is indeed necessary.
For the manifests, the essential check is what is really applied, not which file is called FINAL:
You strip out the fields generated by the cluster (status, metadata.uid, resourceVersion, last-applied-configuration annotations) and that is the version that goes into the repository. The truth is in the cluster, not in Drive.
Phase 2 — Structure and protect (the following week). Create the four repositories, enable branch protection and CODEOWNERS, install the pre-commit hooks with gitleaks and connect the Cloud Build triggers.
Phase 3 — The hard part (the weeks after).
| Asset | Destination | How | Risk |
|---|---|---|---|
| Lucía's notebooks | alpinashop-ml/notebooks/ |
Strip outputs with nbstripout; extract the reusable parts into modules |
High: notebooks often contain customer data in their outputs |
Marta's gcloud commands |
alpinashop-infra/ as Terraform |
Do not transcribe commands: export the real state, 06-05 and 06-07 | Medium: laborious, but nothing is lost |
The detail about the notebooks deserves emphasis: a cell run with df.head() over alpinashop_analitica leaves names, email addresses and postal addresses saved inside the .ipynb. Uploading a notebook without cleaning the outputs is a leak of personal data into the repository, with Git forgetting nothing. nbstripout as a pre-commit hook solves it automatically and is non-negotiable in alpinashop-ml.
And the observation about Marta's commands: transcribing them into a script would be the mistake. That history is incomplete and contains commands that were run and later undone. The right thing to do is to export the real configuration of the live infrastructure, which is exactly the procedure in 06-05.
Solution 2
Federation configuration, reusing the pool already created:
PROJ=alpinashop-datos
NUM=$(gcloud projects describe $PROJ --format='value(projectNumber)')
gcloud iam service-accounts create sa-gha-datos \
--display-name="GitHub Actions - SQL validation" --project=$PROJ
# Only the alpinashop-datos repository, and only the main branch
gcloud iam service-accounts add-iam-policy-binding \
sa-gha-datos@${PROJ}.iam.gserviceaccount.com \
--role=roles/iam.workloadIdentityUser \
--member="principalSet://iam.googleapis.com/projects/${NUM}/locations/global/workloadIdentityPools/github-pool/attribute.repository/alpinashop/alpinashop-datos" \
--project=$PROJExact roles, applying least privilege for real:
| Role | Scope | Why that one and not another |
|---|---|---|
roles/bigquery.jobUser |
alpinashop-datos project |
Allows launching jobs; a dry run does not read data |
roles/bigquery.metadataViewer |
alpinashop_analitica dataset |
The dry run needs the schemas in order to validate |
roles/storage.objectCreator |
alpinashop-datalake bucket, reports prefix |
Create, not read or delete |
Notice what it does not carry: no bigquery.dataViewer — a dry run validates syntax and schema, it neither runs the query nor reads a single row — and no storage.objectAdmin — creating reports does not require being able to delete existing ones. And the objectCreator scoped to the prefix by means of an IAM condition (03-04) stops a compromised workflow overwriting data lake data.
What stops a workflow from alpinashop-catalogo using this identity: the principalSet is tied to attribute.repository/alpinashop/alpinashop-datos. The OIDC token GitHub issues carries the name of the originating repository, signed, and GitHub does not allow it to be forged — it is issued by GitHub's own service, not by the workflow. A workflow from alpinashop-catalogo would present a token with repository: alpinashop/alpinashop-catalogo, which does not match the principalSet, and GCP would reject the exchange. That is the property that makes federation substantially more secure than a JSON key: the identity is certified by a trusted third party and tied to the execution context, not to possession of a file.
And one further improvement: the provider's attribute-condition (assertion.repository_owner == 'alpinashop') acts as a second barrier at provider level. Without it, any GitHub repository in the world could at least attempt the exchange, and although the principalSet would reject it anyway, it is better to filter it earlier. Defence in depth.
Solution 3
The long-lived branch proposal is a bad one, and it is worth explaining why before proposing an alternative. A client-x branch maintained for two years produces, with complete certainty:
- Growing divergence: every security fix has to be applied twice, and at some point one will be forgotten. That forgotten one will be exactly the important one.
- Increasingly expensive merges: six months in, bringing changes from
mainintoclient-xis a project in itself. - Duplication of everything else: two pipelines, two sets of images, two deployments, two sets of metrics.
- A system nobody tests as a whole: CI would test
main, andclient-xwould be left with old tests.
But trunk-based is still valid, because the real problem is not one of branching strategy: it is one of architecture. The right question is not "how do I maintain two branches?", but "how do I make a single code base serve two customers?".
The recommended solution: multi-tenancy by configuration. A single main branch, a single image, and the different behaviour resolved at runtime:
# catalogo/inquilino.py
@dataclass(frozen=True)
class Tenant:
id: str
theme: str # templates and stylesheets
pricing_rules: str # pricing strategy to apply
domain: str
TENANTS = {
"public": Tenant("public", "alpina", "standard", "www.alpinashop.example"),
"client-x": Tenant("client-x", "clientx", "wholesale", "tienda.client-x.example"),
}
def current_tenant() -> Tenant:
# Resolution by domain: the same deployment serves both
return TENANTS.get(DOMAINS.get(request.host), TENANTS["public"])What changes per tenant is data and configuration: the visual theme, the pricing rules, the domain, perhaps the subset of the catalogue. None of that justifies a branch.
| Approach | Maintenance cost | Risk of divergence | Code complexity | Recommended? |
|---|---|---|---|---|
| Long-lived branch | Very high and growing | Certain | Low | No |
| Multi-tenancy by configuration | Low | None | Medium | Yes |
| Separate repository and product | Very high | N/A: they are different products | Low | Only if they genuinely diverge |
| Shared core as a library | Medium | Medium | High | If there are many customers |
The trade-offs, stated honestly. Multi-tenancy is not free: it introduces conditionals into the code, both paths have to be tested, and there is a new and serious risk — a leak between tenants, client X seeing the public data or the other way round. That demands that the tenant identifier be applied in the data access layer and tested explicitly, rather than trusting that every query remembers to filter.
And there is a breaking point that has to be recognised: if in a year's time client X asks for a completely different purchase flow, a catalogue with another structure and a billing process of its own, they are no longer two configurations of the same product but two products. At that moment separation is correct, and it is done as separate products with a common library, not as a branch.
The rule to take away from this: branches serve to separate work in time, not to separate variants in space. When somebody proposes a long-lived branch for a product variant, the signal is that there is an architecture decision waiting to be taken. And the conversation with the sales side matters too: if the contract with client X allows some flexibility in the design and in the pricing rules, multi-tenancy comes almost free; if it demands a completely different product, the real cost of that contract is much higher than it looked when it was signed, and that is worth knowing in advance.
Conclusion
AlpinaShop finally has a foundation to build on.
You know what version control provides — history, attribution, rollback, branching, review, immutable identity and automation — and that the last two are what make everything in 06-01 possible: with no remote repository there is no $COMMIT_SHA, no trigger and no pipeline.
You know Cloud Source Repositories: private Git repositories on GCP, gcloud source repos clone with authentication by gcloud identity, permissions governed by IAM with the same groups from 03-04, code search and native integration with Cloud Build and Error Reporting. And you know the uncomfortable truth: it is deprecated, it takes no new customers, it receives no functionality and it never had code review, which is what is really needed. You know it in order to operate legacy infrastructures and to migrate them with git clone --mirror, not to start anything new.
You know how to connect GitHub or GitLab with GCP using the 2nd generation of Cloud Build repositories — connection and repository as declarable resources, token in Secret Manager — and why it is preferable to the first generation's GitHub application. And above all you know how to use Workload Identity Federation so that GitHub Actions access GCP with no JSON key at all: the pool, the OIDC provider, the attribute-condition, the principalSet tied to a repository and a branch, and the id-token: write everybody forgets. With the rule that sums the section up: if in 2026 you are creating a service account JSON key, there is almost always a better alternative.
You have the honest decision table between CSR, GitHub and GitLab, and AlpinaShop's decision — GitHub — with its four reasons, alongside the cases in which GitLab would be the better choice. You have the branching strategy: trunk-based with short branches, because it is a web service with a single live version, because the team is three people, because it is the very premise of continuous integration and because conflicts grow non-linearly with time. With feature flags as the answer to "this will take three weeks", and with the discipline of deleting them afterwards.
You know what goes into the repository — code, Dockerfile, pipelines, manifests, Terraform, tests, migrations, architecture decisions — and what does not — secrets, Terraform state, data, artefacts, large files —, with the warning that is worth the whole lesson: Git does not forget, and faced with a leaked secret the only correct response is to rotate it. You have AlpinaShop's four-repository structure divided by pace and responsibility, with alpinashop-infra singled out as the most dangerous. You have branch protection with its seven rules, CODEOWNERS demanding a security review for IAM and firewall, conventional commit messages, semantic versioning with its honest nuance, and the pre-commit hooks with gitleaks as a safety net that does not replace CI.
And you have the four triggers connected from beginning to end: a pull request that tests, main that builds and deploys to development, a protected tag that promotes to production with approval, and the ML pipeline compiling itself.
With this, the catalogue builds, tests and deploys itself. But there are still pieces of the system that are neither a web application nor a data pipeline: they are reactions to events. When a message arrives on the imagenes-subidas topic, something has to wake up, call the Vision API, write to BigQuery and go back to sleep. Standing up a virtual machine or a pod for that would be absurd.
In 06-03 come Cloud Functions, and with them the loose end left by module 5 is finally tied off.
Google Cloud Platform (GCP) Course
Module 1: Introduction to Google Cloud Platform
- What is Google Cloud Platform?
- Setting Up Your GCP Account
- A Tour of the GCP Console
- Projects, Resource Hierarchy and Billing
- Regions, Zones and the Shared Responsibility Model
- Cloud Shell and the gcloud CLI
Module 2: Core GCP Services
- Compute Engine: Virtual Machines on Google Cloud
- Cloud Storage: Object Storage
- Cloud SQL: Managed Relational Databases
- App Engine: Platform as a Service
- Google Kubernetes Engine (GKE)
- NoSQL Databases: Firestore, Bigtable and Spanner
- How to Choose the Right Compute Service
Module 3: Networking and Security
- VPC Networks
- Cloud Load Balancing
- Cloud CDN
- Identity and Access Management (IAM)
- Cloud Armor
- Secrets and Encryption: Secret Manager and Cloud KMS
- Cloud DNS, TLS Certificates and Publishing Services Securely
Module 4: Data and Analytics
- BigQuery: The Analytical Data Warehouse
- Cloud Dataflow: Batch and Streaming Data Processing
- Cloud Dataproc: Managed Spark and Hadoop
- Cloud Pub/Sub: Asynchronous Messaging
- Cloud Data Fusion: Code-Free Data Integration
- Orchestrating Pipelines with Cloud Composer and Workflows
- Data Governance and Dashboards with Dataplex and Looker Studio
Module 5: Machine Learning and AI
- Vertex AI: The Machine Learning Platform on GCP
- AutoML: Custom Models Without Writing Code
- TensorFlow on GCP: Training and Serving Models
- Natural Language API
- Vision API
- Generative AI on Vertex AI: Gemini Models and Embeddings
- MLOps: From Model to Product with Vertex AI Pipelines
Module 6: DevOps and Monitoring
- Cloud Build: Continuous Integration on GCP
- Cloud Source Repositories and Source Code Management
- Cloud Functions: Serverless Functions
- Cloud Monitoring (formerly Stackdriver): Metrics, Dashboards and Alerts
- Cloud Deployment Manager and Native Infrastructure as Code
- Cloud Logging and Cloud Trace: Logs, Traces and Diagnostics
- Terraform on GCP: Infrastructure as Code in Practice
Module 7: Advanced GCP Topics
- Hybrid and Multicloud with Anthos
- Serverless Computing with Cloud Run
- Advanced Networking: Shared VPC, Peering and Hybrid Connectivity
- Security Best Practices
- Cost Management and Optimization
- Reliability: SLOs, High Availability and Disaster Recovery
- Governance at Scale: Organization, Policies and Auditing
