Throughout this course we have been leaving loose ends, and they are all the same loose end. In 03-01 we wrote a firewall rule allowing SSH from the IAP range, and we said the firewall does not distinguish between people: making sure "only Marta and Lucía can get in" is solved somewhere else. In 02-02 we limited Lucía to the exportaciones/ prefix of the bucket with a condition we never explained. In 02-05 we bound the GKE pods to a service account with no key files. In 03-03 we said that invalidating the cache requires a role that is best not handed around. That somewhere else, that condition and that role are IAM.
IAM is the system that answers a single question, asked millions of times a day in every project: can this identity perform this action on this resource? It is, without exaggeration, the most important service in Google Cloud. A badly sized VM costs money; a badly granted permission costs the company. And it is also the service most people configure wrongly, almost always in the same way: giving Editor to everybody so that things work and promising themselves they will tidy it up later.
In this lesson you are going to build AlpinaShop's real permission model: you will understand the identity + role + resource equation, you will know why permissions are granted to groups and never to people, you will create a custom role for Lucía, you will design the complete access matrix for Marta, Dani and Lucía across the four projects, you will understand service accounts and why a downloaded JSON key is a time bomb, you will learn to restrict access with conditions, to diagnose why somebody cannot do something, and to publish internal applications and open SSH sessions with no VPN thanks to IAP.
Warning. Permission design has security and, depending on the sector, regulatory implications. What follows is a solid, applicable teaching model, but before taking into production an access scheme governing personal or financial data, it must be reviewed by a security and compliance professional.
Contents
- The IAM equation: identity + role + resource
- Permissions, roles and policies: the three objects to tell apart
- Types of identity
- Google groups: the decision that simplifies everything most
- Types of role: basic, predefined and custom
- A real custom role: "catalogue analyst" for Lucía
- Inheritance down the hierarchy and the effective policy
- AlpinaShop's access design
- Service accounts: identities for software
- Impersonation and short-lived credentials
- JSON keys: why they are the problem and what to use instead
- IAM conditions: access limited in time and by resource
- Diagnosis:
policy-troubleshooter, effective policy and Recommender - Least privilege and separation of duties, step by step
- IAP: SSH with no public IP and internal applications with no VPN
- The IAM equation: identity + role + resource
The whole of IAM boils down to one sentence:
An allow policy binds who (identity) with what they can do (role) on where (resource).
flowchart LR
subgraph QUIEN["WHO - identity"]
U["User<br/>[email protected]"]
G["Group<br/>[email protected]"]
SA["Service account<br/>sa-catalogo-web@..."]
F["Federated identity<br/>GitHub Actions, Entra ID"]
end
subgraph QUE["WHAT - role"]
R["roles/storage.objectViewer<br/>= set of permissions<br/>storage.objects.get<br/>storage.objects.list"]
end
subgraph DONDE["WHERE - resource"]
ORG["Organization alpinashop.example"]
CAR["Folder produccion"]
PRO["Project alpinashop-prod"]
BUC["Bucket alpinashop-catalogo"]
end
QUIEN -->|binding| QUE
QUE -->|applied on| DONDE
ORG --> CAR --> PRO --> BUC
Three immediate consequences of this model:
- Everything is denied by default. An identity with no bindings can do absolutely nothing. You do not have to "remove" permissions: you have to grant them.
- The policy is attached to the resource, not to the person. There is no such thing as "Lucía's permissions"; there is "the policy of the
alpinashop-datosproject", in which Lucía appears. That is why you have to know where to look. - Permissions are additive and are inherited downwards. A role granted on the
produccionfolder applies to every project inside it. And unless you use deny policies, nothing takes away what has been granted higher up.
- Permissions, roles and policies: the three objects to tell apart
| Object | What it is | Example | Is it assigned? |
|---|---|---|---|
| Permission | The atomic unit, in the form service.resource.verb |
storage.objects.get |
No. Individual permissions are never assigned |
| Role | A named set of permissions | roles/storage.objectViewer |
Yes. It is the only thing that is assigned |
| Binding | The pair (identity, role) | group:gcp-datos@… → roles/bigquery.dataViewer |
It is what you create with add-iam-policy-binding |
| Policy | The set of bindings of a resource | The policy of alpinashop-prod |
It is read in full with get-iam-policy |
Permissions correspond almost one to one with API calls. When the console tells you "you do not have permission to perform this action", what is missing is a specific permission; the work consists of finding which role contains it.
# Which permissions a role contains
gcloud iam roles describe roles/storage.objectViewer \
--format="value(includedPermissions)"
# Which predefined roles contain a specific permission
gcloud iam roles list --filter="includedPermissions:compute.urlMaps.invalidateCache" \
--format="table(name, title)"That second command is pure gold: it turns "I need to be able to invalidate the CDN cache" into "I need roles/compute.loadBalancerAdmin" without searching the documentation.
- Types of identity
In IAM jargon, an identity is called a principal and is written with a prefix indicating its type:
| Prefix | Type | Example | Use |
|---|---|---|---|
user: |
A person with a Google account | user:[email protected] |
Only for justified exceptions |
group: |
A Google group | group:[email protected] |
The normal way of giving permissions to people |
serviceAccount: |
A workload's identity | serviceAccount:[email protected] |
Applications, VMs, pods, pipelines |
domain: |
The whole domain | domain:alpinashop.example |
Very broad; use with great care |
principalSet:// |
A federated set | A GitHub repository, a group from an external IdP | Identity federation |
allUsers / allAuthenticatedUsers |
Public | — | Practically only for explicitly public content |
The two forms of federation deserve to be told apart properly, because they get confused:
- Workforce Identity Federation — for people who already have an identity with an external provider (Microsoft Entra ID, Okta, any OIDC/SAML). It lets employees sign in to Google Cloud with their corporate credentials without creating Google accounts. It is configured at organization level.
- Workload Identity Federation — for external machines and processes: a GitHub Actions pipeline, a Kubernetes cluster outside Google, a workload in another cloud. The external process presents its own token (for example, the OIDC token GitHub issues for each run) and Google exchanges it for temporary credentials of a service account.
This second one is today the correct answer to "how do I deploy from GitHub without uploading a JSON key to the repository?". We will see it applied in 06-01, but it is worth knowing it exists from now on, because it eliminates the worst risk in this lesson at the root.
- Google groups: the decision that simplifies everything most
It is tempting to grant permissions directly to people. It is also the decision that makes access ungovernable within a year. Compare:
| Permissions to people | Permissions to groups | |
|---|---|---|
| Onboarding an employee | Repeat N bindings across M projects | Add them to 1-2 groups |
| Offboarding an employee | Hunt down their bindings across the whole hierarchy and hope you miss none | Take them out of the group. Done |
| Change of role | A full manual audit | Change group |
| Auditing | "Who can read production?" is a search across the whole hierarchy | Look at who is in the group |
| Reproducibility in Terraform (06-07) | The code changes with every person | The code names groups, stable for years |
The rule, with no practical exceptions:
Roles are granted to groups. People are added to groups. A
user:binding in a project's policy is a design smell: either it is temporary and carries an expiry condition, or it is wrong.
AlpinaShop's groups are created in Google Workspace (or Cloud Identity, which is free for this use) and must have names that say what they are, not who is in them:
[email protected] → infrastructure and networking [email protected] → application development [email protected] → analytics and data [email protected] → cost visibility [email protected] → review and auditing
One important nuance: IAM does not create or manage groups. Groups live in Cloud Identity / Workspace and IAM merely references them. Membership management is done by the identity administrator, who may be a different person from the cloud administrator. Far from being a drawback, that is a healthy separation of duties.
- Types of role: basic, predefined and custom
| Type | Examples | Granularity | Verdict |
|---|---|---|---|
| Basic (formerly "primitive") | roles/owner, roles/editor, roles/viewer |
Thousands of permissions, the whole project | Anti-pattern. Only viewer in toy environments |
| Predefined | roles/storage.objectViewer, roles/cloudsql.client, roles/compute.networkAdmin |
Per service and per task | The default choice. Google maintains them |
| Custom | alpinashop.analistaCatalogo |
Whatever you decide | When no predefined role fits. You maintain them |
Why the basic ones are an anti-pattern, with specifics rather than dogma:
roles/editorincludes the ability to create service accounts and grant yourself permissions through them in many scenarios, as well as to modify or delete practically any resource in the project. It is a privilege escalation waiting to happen.roles/owneradds the ability to modify the IAM policy: whoever has it can give themselves and anyone else everything else, and take it away from you.- All three predate the existence of most services. When Google launches a new product, its permissions are automatically folded into
editor. In other words, your users' permissions grow without anyone deciding anything. - They break any audit: "who can delete the production database?" is answered with a list of twenty people.
A reasonable exception: in a personal lab project, roles/owner for yourself. In alpinashop-prod, never.
The predefined ones are the answer 90 % of the time. They are grouped by service and usually come in three flavours: viewer (read), user/developer (use), admin (manage). Find them like this:
# Cloud SQL predefined roles
gcloud iam roles list --filter="name:roles/cloudsql" \
--format="table(name, title)"
- A real custom role: "catalogue analyst" for Lucía
In alpinashop-datos, Lucía needs to be able to:
- Query the catalogue tables and run queries.
- Read the exports from the bucket, but not write or delete.
- See the state of the instances so she knows whether a report failed because the VM was stopped, but not stop or start them.
No predefined role says exactly that. roles/viewer would give her read access to the whole project, including the IAM policies and the network configurations. This is the textbook case for a custom role.
First, discover which permissions exist and which can be used in a custom role:
# Permissions applicable in the project (useful for exploring)
gcloud iam list-testable-permissions \
//cloudresourcemanager.googleapis.com/projects/alpinashop-datos \
--filter="customRolesSupportLevel!=NOT_SUPPORTED" \
--format="value(name)" | grep -E "^(bigquery|storage|compute\.instances)"Now the role, in a YAML file that must live in the infrastructure repository, because a role is code:
# roles/analista-catalogo.yaml
title: "Catalogue analyst"
description: "Querying catalogue data and reading exports. No writing."
stage: "GA"
includedPermissions:
# --- BigQuery: read data and run queries ---
- bigquery.datasets.get
- bigquery.tables.get
- bigquery.tables.list
- bigquery.tables.getData
- bigquery.jobs.create # needed to RUN queries
- bigquery.jobs.list
# --- Cloud Storage: read the exports ---
- storage.buckets.get
- storage.objects.get
- storage.objects.list
# --- Compute: see the state, nothing else ---
- compute.instances.get
- compute.instances.list
# --- Basic observability ---
- monitoring.timeSeries.list# Create the role IN THE PROJECT (it can also be created at organization level)
gcloud iam roles create analistaCatalogo \
--project=alpinashop-datos \
--file=roles/analista-catalogo.yaml
# Grant it to the group, not to Lucia
gcloud projects add-iam-policy-binding alpinashop-datos \
--member="group:[email protected]" \
--role="projects/alpinashop-datos/roles/analistaCatalogo"
# Update the role later on (same syntax, update verb)
gcloud iam roles update analistaCatalogo \
--project=alpinashop-datos \
--file=roles/analista-catalogo.yamlFive things to know about custom roles that are only learned by suffering them:
bigquery.jobs.createis the forgotten permission. Without it, Lucía sees the tables but any query fails. Reading data and running jobs are different permissions.stagecan beALPHA,BETA,GAorDISABLED. SettingDISABLEDis the way to switch a role off without deleting it, useful for checking whether anybody depended on it.- They are created in a project or in an organization, not in a folder. If several projects are going to use the role, create it at organization level (
--organization=ORG_ID); otherwise it gets duplicated. - You maintain them. When Google adds a new permission to a service, the predefined roles pick it up by themselves; your custom role does not. Review them periodically.
- Start by copying a predefined role and removing, rather than starting from scratch:
gcloud iam roles copy --source=roles/bigquery.dataViewer --destination=... --dest-project=....
- Inheritance down the hierarchy and the effective policy
AlpinaShop's hierarchy, which you defined in 01-04:
flowchart TD
O["Organization<br/>alpinashop.example"]
F1["Folder produccion"]
F2["Folder desarrollo"]
F3["Folder compartido"]
P1["alpinashop-prod"]
P2["alpinashop-dev"]
P3["alpinashop-datos"]
P4["alpinashop-cicd"]
R["Bucket alpinashop-catalogo<br/>Instance alpinashop-pedidos"]
O --> F1 --> P1 --> R
O --> F2 --> P2
O --> F3 --> P3
F3 --> P4
A resource's effective policy is the union of the policies of all its ancestors plus its own. Consequences:
- A role granted at the organization applies to all four projects and to all their resources. That is why there are very few things that should be granted there.
- You cannot subtract by inheriting. If somebody has
roles/editoron theproduccionfolder, there is no way of "taking it away" inalpinashop-prodwith an allow policy. The only tool that subtracts is deny policies (gcloud iam policies), which are evaluated before the allow ones and always win; they are the exception, not the everyday mechanism, and they are covered alongside governance at scale in 07-07. - The practical rule: grant at the lowest level that solves the problem. If the permission is only needed on a bucket, grant it on the bucket, not on the project.
Many resources accept a policy of their own. Compare the granularity:
# Project level
gcloud projects add-iam-policy-binding alpinashop-datos \
--member="group:[email protected]" \
--role="roles/bigquery.jobUser"
# Bucket level: much finer
gcloud storage buckets add-iam-policy-binding gs://alpinashop-catalogo \
--member="serviceAccount:[email protected]" \
--role="roles/storage.objectViewer"
# Secret level (03-06): the finest possible
gcloud secrets add-iam-policy-binding db-password-catalogo \
--member="serviceAccount:[email protected]" \
--role="roles/secretmanager.secretAccessor"
- AlpinaShop's access design
This is the outcome of the design exercise, and the one you should be able to justify line by line:
| Group | Members | alpinashop-prod |
alpinashop-dev |
alpinashop-datos |
alpinashop-cicd |
|---|---|---|---|---|---|
gcp-infra@ |
Marta | compute.admin, compute.networkAdmin, compute.loadBalancerAdmin, cloudsql.admin, iap.tunnelResourceAccessor |
owner |
viewer |
editor |
gcp-desarrollo@ |
Dani | compute.viewer, logging.viewer, errorreporting.viewer, iap.tunnelResourceAccessor |
editor |
bigquery.dataViewer |
cloudbuild.builds.viewer, artifactregistry.writer |
gcp-datos@ |
Lucía | (none directly) | — | analistaCatalogo (custom), bigquery.jobUser |
— |
gcp-facturacion@ |
Marta, management | billing.viewer at organization level |
|||
gcp-seguridad@ |
Marta (and external audit) | iam.securityReviewer at organization level |
Read it carefully, because the interesting decisions are the ones that do not appear:
- Nobody has
ownerinalpinashop-prod. Not even Marta. Administering production's IAM policy is done through an explicit temporary elevation (section 12), not with a permanent role. - Dani cannot modify production. He can look — metrics, logs, instance state — because he needs to diagnose incidents, and he can get in over SSH via IAP to debug. But he cannot deploy by hand: changes go in through the
alpinashop-cicdpipeline (06-01). That is not distrust, it is traceability: every change in production has a commit behind it. - Marta is
ownerinalpinashop-dev. The development environment gets broken and rebuilt; friction adds nothing there. - Lucía has absolutely nothing in production. Her data reaches
alpinashop-datosthrough replication and export. If she needed to read directly from thealpinashop-pedidos-replica-informesreplica, the permission would beroles/cloudsql.clienton that specific instance, with a condition, and not a project-level role. - Billing is viewed at organization level, not project level, because cost questions are cross-cutting (07-05).
roles/iam.securityReviewerallows reading every IAM policy without being able to modify any. It is the right role for auditing and for answering "who has access to what?".
Applied with gcloud, the design is written like this:
ORG_ID=$(gcloud organizations list --format="value(name)" | head -1)
# --- Infrastructure in production: administration, not ownership ---
for ROLE in roles/compute.admin roles/compute.networkAdmin \
roles/compute.loadBalancerAdmin roles/cloudsql.admin \
roles/iap.tunnelResourceAccessor; do
gcloud projects add-iam-policy-binding alpinashop-prod \
--member="group:[email protected]" \
--role="$ROLE" --condition=None
done
# --- Development: read in production, control in dev ---
for ROLE in roles/compute.viewer roles/logging.viewer \
roles/iap.tunnelResourceAccessor; do
gcloud projects add-iam-policy-binding alpinashop-prod \
--member="group:[email protected]" \
--role="$ROLE" --condition=None
done
gcloud projects add-iam-policy-binding alpinashop-dev \
--member="group:[email protected]" \
--role="roles/editor" --condition=None
# --- Security and billing at organization level ---
gcloud organizations add-iam-policy-binding "$ORG_ID" \
--member="group:[email protected]" \
--role="roles/iam.securityReviewer"
gcloud organizations add-iam-policy-binding "$ORG_ID" \
--member="group:[email protected]" \
--role="roles/billing.viewer"
--condition=Noneis mandatory when the policy already contains conditional bindings; without it,gcloudasks interactively and the loop stops.
- Service accounts: identities for software
A service account is an identity that belongs to the application, not to a person. It has an email of the form [email protected] and it is two things at once, which is disconcerting at first:
- An identity: roles are granted to it, as to a user.
- A resource: it has its own IAM policy, which says who can use it.
That second facet is the one almost nobody sees at first and the one that explains half of all permission errors.
# Create one service account per workload
gcloud iam service-accounts create sa-informes-nocturnos \
--display-name="Nightly reporting process" \
--description="Reads the orders replica and writes to BigQuery"
SA="[email protected]"
# 1) As an IDENTITY: what it can do
gcloud projects add-iam-policy-binding alpinashop-datos \
--member="serviceAccount:$SA" --role="roles/bigquery.dataEditor"
gcloud projects add-iam-policy-binding alpinashop-prod \
--member="serviceAccount:$SA" --role="roles/cloudsql.client"
# 2) As a RESOURCE: who can act as it
gcloud iam service-accounts add-iam-policy-binding "$SA" \
--member="group:[email protected]" \
--role="roles/iam.serviceAccountTokenCreator"AlpinaShop's service accounts so far, and why there are several and not one:
| Account | Workload | Roles |
|---|---|---|
sa-catalogo-web |
The Flask application on the MIG | storage.objectViewer on the bucket, cloudsql.client, secretmanager.secretAccessor on two secrets |
sa-migracion-catalogo |
The one-off initial upload process | storage.objectCreator on the bucket |
sa-catalogo-gke |
The pods in the tienda namespace |
The same as sa-catalogo-web, via Workload Identity |
sa-informes-nocturnos |
Lucía's batch process | Read from the replica, write to BigQuery |
One service account per workload. Not one per project, not one per team. When a credential is compromised, the blast radius is exactly that of that workload. And when you review the logs, you know who did what with no ambiguity.
Two specific warnings:
- Never use the default Compute Engine service account. It is created automatically, it comes with
roles/editoron the project and every VM uses it unless you say otherwise. In other words: any code on any VM can do almost anything in the project. Create your own and assign it explicitly in the instance template. roles/iam.serviceAccountUseron a service account is more powerful than it looks. It allows launching resources as that identity, and therefore inheriting its permissions. Granting it on a powerful account amounts to granting those permissions.
- Impersonation and short-lived credentials
The correct way for a person to run something with a service account's permissions is not to download its key: it is to impersonate it. Google issues a short-lived token — typically one hour — and there is no credentials file to steal.
# Requirement: having roles/iam.serviceAccountTokenCreator on that account
gcloud storage ls gs://alpinashop-catalogo/exportaciones/ \
--impersonate-service-account=sa-informes-nocturnos@alpinashop-datos.iam.gserviceaccount.com
# Set it for the whole session
gcloud config set auth/impersonate_service_account \
[email protected]
# Get a one-off token (for curl against an API)
gcloud auth print-access-token \
--impersonate-service-account=sa-informes-nocturnos@alpinashop-datos.iam.gserviceaccount.comAnd for the Python client libraries, impersonation also works with no keys:
from google.auth import default, impersonated_credentials
from google.cloud import storage
base_credentials, _ = default()
credentials = impersonated_credentials.Credentials(
source_credentials=base_credentials,
target_principal="[email protected]",
target_scopes=["https://www.googleapis.com/auth/cloud-platform"],
lifetime=3600, # seconds; the usual maximum is one hour
)
client = storage.Client(credentials=credentials, project="alpinashop-datos")The advantages over a downloaded key are worth listing, because this is the argument you will have to make to somebody:
- It expires by itself. An hour later it is worthless.
- It leaves a trail. The audit logs record that
lucia@impersonatedsa-informes-nocturnos, so the action has a human owner. - It is revoked by removing a role, not by chasing files across the team's laptops.
- It chains well: one process can impersonate another account if the policy allows it, forming auditable chains.
Related, and very useful from an organizational point of view: Privileged Access Manager allows granting temporary elevations with approval and automatic expiry — "Marta needs to be an IAM administrator in production for two hours to fix this" — without anybody holding that role permanently. It is the modern answer to "nobody has owner in production, but somebody has to be able to fix it at 3 in the morning".
- JSON keys: why they are the problem and what to use instead
This command exists, it works, and you should treat it as the last resort:
# What NOT to do unless there is no alternative
gcloud iam service-accounts keys create clave.json \
--iam-account=sa-informes-nocturnos@alpinashop-datos.iam.gserviceaccount.comThat file contains an RSA private key that never expires and that authenticates the service account from anywhere on the internet. No second factor, no network restriction, no expiry. What happens next is always the same story: it ends up in a Git repository, in a Slack channel, in a .env copied to three laptops, or in a container image published to a registry.
Alternatives, in order of preference:
| Scenario | Correct solution | Keys |
|---|---|---|
| Code on a Compute Engine VM | Assign the service account to the VM; the libraries detect it on their own | None |
| Pods on GKE | Workload Identity (02-05) | None |
| Cloud Run, Cloud Functions, App Engine | Assign the service account to the service | None |
| CI/CD on GitHub Actions, GitLab, Azure DevOps | Workload Identity Federation | None |
| A person running something one-off | Impersonation | None |
| A legacy system outside Google that does not support OIDC | A JSON key, with forced rotation and minimal scope | One, watched |
If you end up in the last row, at least:
# Inventory: which user keys exist and since when?
for SA in $(gcloud iam service-accounts list --format="value(email)"); do
gcloud iam service-accounts keys list --iam-account="$SA" \
--managed-by=user --format="table[no-heading](name.basename(), validAfterTime)" \
| sed "s|^|$SA |"
doneAnd block key creation preventively with the organization policy constraint constraints/iam.disableServiceAccountKeyCreation, which is explained in 07-07. It is one of the three or four measures with the best effort/risk ratio on the whole platform.
- IAM conditions: access limited in time and by resource
A conditional binding only grants the role if a CEL expression is satisfied. It serves two very practical purposes.
Case 1: temporary access. An external consultant needs to see production logs for a week:
gcloud projects add-iam-policy-binding alpinashop-prod \
--member="user:[email protected]" \
--role="roles/logging.viewer" \
--condition='expression=request.time < timestamp("2026-08-20T00:00:00Z"),
title=acceso-temporal-auditoria-agosto,
description=Expires by itself on 20 August 2026'This eliminates access debt by construction: the permission withdraws itself, even if nobody remembers. It is the correct way of granting any exceptional access.
Case 2: access limited by resource. This is the condition we left pending in 02-02, the one restricting Lucía to the exportaciones/ prefix:
gcloud storage buckets add-iam-policy-binding gs://alpinashop-catalogo \
--member="group:[email protected]" \
--role="roles/storage.objectViewer" \
--condition='expression=resource.name.startsWith("projects/_/buckets/alpinashop-catalogo/objects/exportaciones/"),
title=solo-prefijo-exportaciones,
description=Read access limited to exportaciones/'Other attributes available in the expressions:
| Attribute | Example | Use |
|---|---|---|
request.time |
request.time < timestamp("...") |
Expiry |
request.time.getHours("Europe/Madrid") |
>= 8 && <= 20 |
Time windows |
resource.name |
.startsWith(...) / .endsWith(...) |
A prefix, a specific resource |
resource.type |
== "compute.googleapis.com/Instance" |
A resource type |
request.auth.claims |
The token's claims | Federation |
Limitations to know about before designing on top of this:
- Not every service supports conditions on
resource.name. Cloud Storage, Compute Engine, Secret Manager and BigQuery do, to varying degrees; others ignore the attribute and the condition is never satisfied, leaving the person with no access and you confused. Always check it in the specific service's documentation. - Conditions cannot be used with basic roles.
- They complicate diagnosis. A person with the right role and a condition that is not satisfied sees exactly the same error as a person without the role. That is why the next section exists.
- Diagnosis:
policy-troubleshooter, effective policy and Recommender
policy-troubleshooter, effective policy and RecommenderAny cloud administrator's most frequent question is "why can't this user do this?". There is a tool that answers it directly:
gcloud policy-troubleshoot iam \
//cloudresourcemanager.googleapis.com/projects/alpinashop-prod \
[email protected] \
--permission=compute.instances.setMetadataThe answer is not a yes or a no: it is the list of every binding examined across the whole hierarchy, with the verdict on each one and the reason. There you can see whether the role is missing, whether it is there but in the wrong project, or whether it is there with a condition that is not satisfied. It is the difference between diagnosing and guessing.
The other three tools in the kit:
# 1) A resource's complete policy
gcloud projects get-iam-policy alpinashop-prod --format=yaml
# 2) Which roles ONE identity has in a project (the inverse query)
gcloud projects get-iam-policy alpinashop-prod \
--flatten="bindings[].members" \
--filter="bindings.members:[email protected]" \
--format="table(bindings.role, bindings.condition.title)"
# 3) Search for an identity across the WHOLE organization (Cloud Asset Inventory)
gcloud asset search-all-iam-policies \
--scope="organizations/$ORG_ID" \
--query="policy:[email protected]" \
--format="table(resource, policy.bindings.role)"Command 3 is the one that really answers "what access does this person have?" before an offboarding or an audit. Number 2 only looks at one project; number 3 sweeps the organization, folders, projects and resources.
IAM Recommender closes the loop: it analyses ninety days of real usage and proposes removing permissions nobody has exercised.
gcloud recommender recommendations list \
--project=alpinashop-prod \
--location=global \
--recommender=google.iam.policy.Recommender \
--format="table(content.overview.member, content.overview.removedRole, priority)"Two common-sense caveats about these recommendations: there are permissions used once a year — the accounting close process, restoring a backup — and ninety days of analysis do not see them. Review before applying, and apply first to service accounts, where the usage pattern is far more stable than a person's.
- Least privilege and separation of duties, step by step
Least privilege: each identity has exactly the permissions it needs, not one more. In practice it is a procedure, not a virtue:
- Start from zero. Never from
Editorwith the intention of trimming; that trimming never arrives. - Let it fail. Run the real job and note which permission each error demands.
- Translate the permission into a role with
gcloud iam roles list --filter="includedPermissions:...". - Grant the smallest predefined role that contains it; if it is still too big, a custom role.
- Grant it at the lowest possible level: resource before project, project before folder.
- Review after three months with Recommender and with
search-all-iam-policies.
Separation of duties: no single identity can complete a critical chain on its own. Applied to AlpinaShop:
| Critical chain | How it is separated |
|---|---|
| Writing code → deploying to production | Dani writes and approves PRs; the pipeline deploys with sa-despliegue, not Dani |
| Creating a permission → using it | Whoever administers IAM (gcp-infra@) is not whoever operates the data (gcp-datos@) |
| Generating spend → approving it | gcp-facturacion@ sees the cost; it cannot create resources |
| Acting → auditing | gcp-seguridad@ has securityReviewer: it reads policies, it does not modify them |
And the corollary that is already in section 8's table: nobody is a permanent owner of production. When it is needed, they are elevated temporarily, with a reason and with an expiry.
- IAP: SSH with no public IP and internal applications with no VPN
Identity-Aware Proxy applies IAM to application access, not just to the API. It is the bridge between what you learned about networking in 03-01 and what you have just learned about identity. Its idea is zero trust: authorisation depends on who you are, not where you are.
It has two uses, and at AlpinaShop we use both.
Use 1: TCP forwarding for SSH. You already wrote the fw-allow-ssh-iap rule allowing port 22 only from 35.235.240.0/20. What was missing was the permission:
# Who can open tunnels towards the project's VMs
gcloud projects add-iam-policy-binding alpinashop-prod \
--member="group:[email protected]" \
--role="roles/iap.tunnelResourceAccessor"
# And also, being able to log in to the OS (OS Login, 02-01)
gcloud projects add-iam-policy-binding alpinashop-prod \
--member="group:[email protected]" \
--role="roles/compute.osLogin"
# Connection, with no public IP on the VM
gcloud compute ssh alpinashop-informes-1 --zone=europe-west1-b --tunnel-through-iapHere at last is the answer to the exercise left pending in 03-01: "only Marta and Lucía can get in over SSH" is implemented with roles/iap.tunnelResourceAccessor and roles/compute.osLogin, not with a firewall rule. The firewall opens the door to the IAP range; IAM decides who walks through it. And if we wanted Lucía to reach only the reporting VM and not the catalogue ones, it would be a conditional binding on resource.name.
Use 2: publishing internal applications with no VPN. Marta wants the internal reporting dashboard to be reachable from home, with no VPN, only for authorised staff. The traditional solution would be a VPN; IAP's consists of putting the dashboard behind an external load balancer and enabling IAP on its backend service:
gcloud compute backend-services update bs-informes-internos --global \
--iap=enabled
# Who can see the application
gcloud iap web add-iam-policy-binding \
--resource-type=backend-services \
--service=bs-informes-internos \
--member="group:[email protected]" \
--role="roles/iap.httpsResourceAccessor"The flow of a request:
sequenceDiagram
participant L as Lucía (from home)
participant LB as Global load balancer
participant IAP as Identity-Aware Proxy
participant G as Google (login + 2FA)
participant B as bs-informes-internos
L->>LB: GET https://informes.alpinashop.example/
LB->>IAP: check authorisation
IAP-->>L: redirect to sign-in
L->>G: authentication + second factor
G-->>IAP: identity verified
IAP->>IAP: does she have iap.httpsResourceAccessor?
IAP->>B: request + header with the signed identity
B-->>L: reporting dashboard
Three details that make this genuinely secure:
- The application never sees unauthenticated traffic. IAP filters first.
- IAP injects the identity into signed headers (
X-Goog-IAP-JWT-Assertion), which the application must validate cryptographically. If you just read the email from a header without verifying the signature, anyone who can reach the backend directly can impersonate whoever they like. - For that very reason, the firewall must still stop anyone reaching the backend around the load balancer: IAP does not replace the network rules from 03-01, it complements them.
And it combines naturally with what comes next: Cloud Armor (03-05) filters by reputation and patterns before IAP asks for credentials, so that automated traffic does not even reach the sign-in screen.
Common Mistakes and Tips
- Granting
roles/editor"to make it work". It is the most expensive mistake in the ecosystem. Spend fifteen minutes finding the right predefined role. - Giving permissions to people instead of to groups. It works for the first month and becomes impossible to audit within the first year.
- Using the default Compute Engine service account. It ships with
roles/editorand is inherited by every VM that does not say otherwise. - Downloading JSON keys. They do not expire, they have no second factor and they end up in Git. Impersonation or Workload Identity Federation, almost always.
- Forgetting that a service account is also a resource. "It has the right roles but cannot use it" almost always means
serviceAccountUserorserviceAccountTokenCreatoris missing from the account's policy. - Granting at the organization what is only needed in one project. It is inherited downwards and cannot be subtracted.
- Believing that an inherited permission can be "removed" with another binding. Permissions are additive; only deny policies subtract (07-07).
- Conditions on services that do not support them. The binding is created without error and never works. Verify the service's support.
- Confusing authentication with authorisation.
gcloud auth loginsays who you are; IAM says what you can do. A permission error is not fixed by authenticating again. - Applying the Recommender's recommendations blindly. Ninety days do not see an annual process.
- Forgetting
--condition=Nonein scripts. The command sits waiting for interactive input and the loop hangs. - Tip: manage IAM as code. Terraform (06-07) for the bindings, versioned YAML for the custom roles. A permission granted by hand in the console is a permission nobody will remember the reason for.
- Tip: document every exception in the condition's own
description. Your six-months-from-now self will thank you.
Exercises
Exercise 1 — "Dani cannot deploy"
Dani tries to update the MIG's template in alpinashop-prod from his laptop and gets PERMISSION_DENIED: compute.instanceGroupManagers.update. Write the diagnostic sequence and decide what to do. Bear in mind the access matrix from section 8 before answering "grant him the role".
Exercise 2 — Custom role for the nightly process
Lucía's batch process runs every night on alpinashop-informes-1 and must:
- Read the
alpinashop-pedidos-replica-informesreplica through the Auth Proxy. - Write a CSV file to
gs://alpinashop-catalogo/exportaciones/<date>/. - Load that file into a BigQuery table in the
alpinashop-datosproject. - Read the replica's password from Secret Manager.
- It must not be able to delete anything, nor read the catalogue images, nor see other secrets.
Design the complete solution: service account, roles (predefined or custom), conditions and at which level of the hierarchy each thing is granted. Write the commands.
Exercise 3 — Temporary access for an external audit
An external consultancy will carry out a security audit for two weeks, from 1 to 15 September 2026. They need to read the network configurations, the IAM policies and the load balancer logs of alpinashop-prod, but they must not see customer data or modify anything. Design the access and justify each decision.
Solutions
Solution 1
Diagnosis:
# 1) What does the troubleshooter say?
gcloud policy-troubleshoot iam \
//cloudresourcemanager.googleapis.com/projects/alpinashop-prod \
[email protected] \
--permission=compute.instanceGroupManagers.update
# 2) Which roles does Dani have there?
gcloud projects get-iam-policy alpinashop-prod \
--flatten="bindings[].members" \
--filter="bindings.members:[email protected] OR bindings.members:[email protected]" \
--format="table(bindings.role, bindings.condition.title)"
# 3) Which role would contain that permission?
gcloud iam roles list \
--filter="includedPermissions:compute.instanceGroupManagers.update" \
--format="value(name)"Expected result: Dani belongs to gcp-desarrollo@, which in alpinashop-prod only has read roles. The permission is in roles/compute.instanceGroupManagerAdmin and in roles/compute.admin.
What to do: do not grant it to him. The matrix in section 8 is a deliberate separation-of-duties decision: changes in production go in through the pipeline, with a commit, a review and traceability. Granting him the role would fix the symptom and destroy the property that makes the system trustworthy.
The correct answers, in order:
- Deploy through the pipeline (06-01): the change is made in the repository, reviewed, and applied by
sa-despliegue, which does have the role inalpinashop-prod. - If it is a genuine emergency, a temporary elevation with Privileged Access Manager, or failing that a conditional binding with an expiry of hours and a description explaining the incident:
gcloud projects add-iam-policy-binding alpinashop-prod \
--member="user:[email protected]" \
--role="roles/compute.instanceGroupManagerAdmin" \
--condition='expression=request.time < timestamp("2026-08-06T06:00:00Z"),
title=incidente-INC-2026-08-05,
description=Emergency access; expires at 06:00 UTC'- If Dani needs this every week, what is failing is not the permissions but the pipeline. Fix the pipeline.
Solution 2
SA="[email protected]"
gcloud iam service-accounts create sa-informes-nocturnos \
--project=alpinashop-datos \
--display-name="Nightly reporting process"
# 1) Read the replica: client role IN THE PROJECT where Cloud SQL lives
gcloud projects add-iam-policy-binding alpinashop-prod \
--member="serviceAccount:$SA" \
--role="roles/cloudsql.client" --condition=None
# 2) Write ONLY to exportaciones/: creator, not admin, and with a condition
gcloud storage buckets add-iam-policy-binding gs://alpinashop-catalogo \
--member="serviceAccount:$SA" \
--role="roles/storage.objectCreator" \
--condition='expression=resource.name.startsWith("projects/_/buckets/alpinashop-catalogo/objects/exportaciones/"),
title=solo-escribe-exportaciones'
# 3) Load into BigQuery: write data + be able to run jobs
gcloud projects add-iam-policy-binding alpinashop-datos \
--member="serviceAccount:$SA" \
--role="roles/bigquery.dataEditor" --condition=None
gcloud projects add-iam-policy-binding alpinashop-datos \
--member="serviceAccount:$SA" \
--role="roles/bigquery.jobUser" --condition=None
# 4) A single secret, at SECRET level and not at project level
gcloud secrets add-iam-policy-binding db-password-informes \
--project=alpinashop-prod \
--member="serviceAccount:$SA" \
--role="roles/secretmanager.secretAccessor"
# 5) Assign the account to the VM (no keys)
gcloud compute instances set-service-account alpinashop-informes-1 \
--zone=europe-west1-b \
--service-account="$SA" \
--scopes=https://www.googleapis.com/auth/cloud-platformThe key decisions and why:
objectCreatorinstead ofobjectAdmin. It allows creating objects but not deleting or overwriting them. It meets the "delete nothing" requirement by construction, not by trust.- The prefix condition stops the process writing into
productos/even if it had a programming bug. And since it has noobjectViewer, it cannot read the images either. - The secret is granted at secret level.
roles/secretmanager.secretAccessoron the project would give access to every secret, including the payment gateway's key. cloudsql.clientgoes inalpinashop-prod, which is where the instance lives, even though the service account belongs toalpinashop-datos. An identity from one project can hold roles in another; it is completely normal.- No JSON key. The VM carries the account assigned to it and the Python libraries detect it on their own.
- A custom role adds nothing here: the predefined ones, narrowed with conditions and applied to the right resource, are already tight enough. Do not create custom roles if a well-placed predefined one solves the case.
Solution 3
# A dedicated group, not permissions to individual email addresses
GROUP="group:[email protected]"
EXPIRES='request.time < timestamp("2026-09-16T00:00:00Z")'
for ROLE in roles/compute.networkViewer \
roles/iam.securityReviewer \
roles/logging.viewer \
roles/monitoring.viewer; do
gcloud projects add-iam-policy-binding alpinashop-prod \
--member="$GROUP" --role="$ROLE" \
--condition="expression=$EXPIRES,
title=auditoria-externa-sept-2026,
description=Security audit; expires on 16-09-2026"
doneThe justification:
- A group of their own, even for three outside people. They are added to and removed from the group, and the policy is not touched.
- An expiry on every binding. Access disappears by itself on the 16th. This is the canonical use of conditions and it avoids the classic "the consultant still has access two years later".
roles/iam.securityReviewerallows reading the IAM policies without being able to modify them: exactly what an audit needs.roles/compute.networkViewerinstead ofroles/compute.viewer: they see the network, the firewall rules and the load balancer, but not the instance metadata, which can contain sensitive information.- No role over data. No
cloudsql.viewer, nostorage.objectViewer, nobigquery.dataViewer: they can audit the database's configuration without reading a single customer row. If they needed to see data, the data protection officer would have to be involved and anonymised data would probably be enough. - No access to
alpinashop-devoralpinashop-datos, unless the audit's scope requires it in writing. - In addition, check in 07-07 that data access audit logs are enabled during the period, so that there is a record of what the consultancy queried.
Conclusion
IAM has stopped being the place you go to "give permissions" and has become the design that holds up everything else. You have internalised the equation identity + role + resource = allow policy, with its three properties: everything is denied by default, the policy is attached to the resource, and permissions are inherited downwards and cannot be subtracted. You tell permission, role, binding and policy apart, and you know how to translate a PERMISSION_DENIED into the role that resolves it with gcloud iam roles list --filter="includedPermissions:...".
You know the types of identity — users, groups, service accounts and the two federations, Workforce for people and Workload for machines — and you have adopted the rule that most simplifies long-term operation: roles are granted to groups and people move between groups. You know why Owner, Editor and Viewer are an anti-pattern (they grow by themselves, they allow escalation, they ruin auditing) and you have built AlpinaShop's real design: gcp-infra@, gcp-desarrollo@, gcp-datos@, gcp-facturacion@ and gcp-seguridad@ with different roles in each of the four projects, with nobody being owner of production. You have created the analistaCatalogo custom role for Lucía in versioned YAML, with the bigquery.jobs.create permission that almost everybody forgets.
You have understood the dual nature of service accounts — identity and resource at the same time — the rule of one per workload, and why the default Compute Engine account must never be used. You know how to impersonate accounts with --impersonate-service-account to obtain credentials that expire in an hour and leave a trail, and you have the ordered list of alternatives to the downloaded JSON key, which is the worst credential in existence: it does not expire, it has no second factor and it ends up in Git. You have limited access in time and by resource prefix with CEL conditions, finally resolving Lucía's restriction to exportaciones/ that was left pending in 02-02. And you have the four diagnostic tools: policy-troubleshoot for "why can't they", get-iam-policy for the snapshot, asset search-all-iam-policies for sweeping the whole organization before an offboarding, and the Recommender for trimming what nobody uses.
Finally, IAP has closed the circle between network and identity: the fw-allow-ssh-iap rule from 03-01 opens the door to the 35.235.240.0/20 range, and roles/iap.tunnelResourceAccessor decides who walks through it; and Lucía's reporting dashboard can be published on the internet with no VPN because IAP requires a corporate identity and a second factor before the request reaches the backend.
But IAM protects against whoever has an identity. It says nothing about the anonymous visitor crawling the catalogue at a thousand requests per second to copy the prices, nor about the one trying ten thousand passwords against the sign-in form, nor about the one injecting ' OR 1=1-- into the search box. That traffic reaches the load balancer with no identity at all and has to be filtered before, at the edge. In the next lesson, 03-05, Cloud Armor, we put an application firewall in front of the global load balancer: rules by IP and geography, protection against the OWASP Top 10, rate limiting against brute force, the indispensable preview mode so you do not block your own customers, and the pol-catalogo-web policy applied to bs-catalogo-web just in time for the autumn campaign.
Google Cloud Platform (GCP) Course
Module 1: Introduction to Google Cloud Platform
- What is Google Cloud Platform?
- Setting Up Your GCP Account
- A Tour of the GCP Console
- Projects, Resource Hierarchy and Billing
- Regions, Zones and the Shared Responsibility Model
- Cloud Shell and the gcloud CLI
Module 2: Core GCP Services
- Compute Engine: Virtual Machines on Google Cloud
- Cloud Storage: Object Storage
- Cloud SQL: Managed Relational Databases
- App Engine: Platform as a Service
- Google Kubernetes Engine (GKE)
- NoSQL Databases: Firestore, Bigtable and Spanner
- How to Choose the Right Compute Service
Module 3: Networking and Security
- VPC Networks
- Cloud Load Balancing
- Cloud CDN
- Identity and Access Management (IAM)
- Cloud Armor
- Secrets and Encryption: Secret Manager and Cloud KMS
- Cloud DNS, TLS Certificates and Publishing Services Securely
Module 4: Data and Analytics
- BigQuery: The Analytical Data Warehouse
- Cloud Dataflow: Batch and Streaming Data Processing
- Cloud Dataproc: Managed Spark and Hadoop
- Cloud Pub/Sub: Asynchronous Messaging
- Cloud Data Fusion: Code-Free Data Integration
- Orchestrating Pipelines with Cloud Composer and Workflows
- Data Governance and Dashboards with Dataplex and Looker Studio
Module 5: Machine Learning and AI
- Vertex AI: The Machine Learning Platform on GCP
- AutoML: Custom Models Without Writing Code
- TensorFlow on GCP: Training and Serving Models
- Natural Language API
- Vision API
- Generative AI on Vertex AI: Gemini Models and Embeddings
- MLOps: From Model to Product with Vertex AI Pipelines
Module 6: DevOps and Monitoring
- Cloud Build: Continuous Integration on GCP
- Cloud Source Repositories and Source Code Management
- Cloud Functions: Serverless Functions
- Cloud Monitoring (formerly Stackdriver): Metrics, Dashboards and Alerts
- Cloud Deployment Manager and Native Infrastructure as Code
- Cloud Logging and Cloud Trace: Logs, Traces and Diagnostics
- Terraform on GCP: Infrastructure as Code in Practice
Module 7: Advanced GCP Topics
- Hybrid and Multicloud with Anthos
- Serverless Computing with Cloud Run
- Advanced Networking: Shared VPC, Peering and Hybrid Connectivity
- Security Best Practices
- Cost Management and Optimization
- Reliability: SLOs, High Availability and Disaster Recovery
- Governance at Scale: Organization, Policies and Auditing
