AlpinaShop already builds, deploys and observes itself. Module 7 begins with the part the course has been dodging: there are AlpinaShop systems that are not on Google Cloud and are not going to be. In the physical shop in Sabadell there is a server running an ERP on MySQL that manages inventory, the till and the invoices; and in the warehouse there is a radio-frequency terminal that talks to that ERP over the local network. Nobody has migrated them, nobody plans to migrate them this year, and yet the online shop needs to know the real stock.
That is a hybrid architecture, even if nobody has ever called it that. And it is not a rare case: it is the normal situation of almost any company with more than five years of history. This lesson is about how Google Cloud tackles that scenario, and about the product that represented it for years: Anthos.
With one honest warning up front, because it sets the tone for the whole module: by the end of the lesson, AlpinaShop is not going to adopt Anthos. And understanding why is as valuable as knowing how to configure it, because the skill that separates a good cloud architect from a good reciter of services is knowing when the big tool is not needed. You will learn what it is, how it works, what problems it genuinely solves and how much it costs — and with those three facts you will be able to make the decision yourself, at your own company, on your own judgement.
A note on names first, because this is the Google Cloud product that has changed name the most times and the documentation you will find is spread across all of them. It is set out in section 5.
Contents
- What "hybrid" and "multicloud" mean exactly
- The real reasons for not being on one cloud alone
- The hidden costs the brochure does not mention
- The AlpinaShop case: Sabadell, the ERP and the warehouse terminal
- What Anthos is and what it has become
- The fleet: the unit of management
- The components, one by one
- Config Sync: configuration as code with GitOps
- Policy Controller: policy as code
- Practical example: registering
alpinashop-clusterand applying a policy - The service mesh explained without the hype
- Binary Authorization and the chain of trust
- Unified fleet observability
- Licensing and cost model: what decides it in practice
- Lighter alternatives to full hybrid
- When yes and when no: the decision table
- DA-004: AlpinaShop's honest decision
- What "hybrid" and "multicloud" mean exactly
The two terms are used as synonyms and they are not. The difference matters because the problems they pose are different.
| Term | Definition | Typical example |
|---|---|---|
| Single public cloud | The whole workload on one provider | AlpinaShop until today: everything on GCP |
| Hybrid | Public cloud + your own infrastructure (data centre, office, shop, factory) | GCP + the Sabadell ERP |
| Multicloud | Two or more public clouds | GCP for the catalogue, AWS for the SaaS ERP |
| Edge | Compute close to the point of use: shop, plant, vehicle | The warehouse terminal running local logic |
And a fourth category almost nobody names but which is the most frequent of all:
- Accidental multicloud: the company is on GCP, but it also has a CRM on Salesforce, email on Microsoft 365, invoicing on a SaaS product and a couple of machines at a cheap provider somebody signed up for in 2019. Nobody decided it; it happened. And yet it has to be integrated, protected and paid for.
flowchart TB
subgraph GCP["Google Cloud - europe-west1"]
A["alpinashop-prod<br/>catalogue, GKE, Cloud SQL"]
B["alpinashop-datos<br/>BigQuery"]
end
subgraph SAB["Sabadell shop - on site"]
C["ERP server<br/>MySQL 5.7"]
D["Warehouse terminal<br/>radio frequency"]
end
subgraph SAAS["Third party SaaS"]
E["Payment gateway"]
F["Email and office"]
end
A <-->|"stock and orders"| C
C --- D
A --> E
B <-.->|"reports"| F
style GCP fill:#e8f0fe,stroke:#4285f4
style SAB fill:#fce8e6,stroke:#ea4335
style SAAS fill:#fef7e0,stroke:#fbbc04
AlpinaShop, without ever having decided it, is already hybrid and is already accidentally multicloud. The architectural work is not about avoiding that, but about managing it deliberately.
- The real reasons for not being on one cloud alone
Sales presentations always give the same five reasons. They are true, but each one has an honest version worth knowing.
Prior investment in the data centre
The brochure reason: "make the most of your existing investment".
The real version: if eighteen months ago the company bought servers with a five-year depreciation schedule, the finance director is not going to sign off on throwing them in the skip, and quite rightly. The cost of those servers is already paid: switching them off does not give the money back, it only saves the electricity and the maintenance. Migrating before they are written off means paying twice for the same capacity.
It is the most common reason and the most legitimate, and also the one that expires by itself: in three years' time that investment no longer exists and the conversation is a different one. That is why many hybrid architectures are in reality slow migrations dressed up as a permanent strategy.
Latency with an on-site system
The brochure reason: "process the data where it is generated".
The real version: there are processes where 30 milliseconds matter and others where they do not. A warehouse terminal scanning barcodes against a local server answers in 2 ms; against europe-west1 it will answer in 15-25 ms from Sabadell. For scanning a parcel, neither of the two is annoying. For the robotic arm on a production line adjusting its position a thousand times per second, the cloud is not an option at any latency.
The right question is not "is there latency?" — there always is — but "what happens if the link goes down for two hours?". If the answer is "the shop cannot take payments", the system has to be able to work offline, and that forces you to have local compute. That, and not the milliseconds, is the strong argument.
Data residency and regulatory requirements
The brochure reason: "meet your regulatory obligations".
The real version: there are three very different levels that are constantly confused:
| Level | What it demands | Does it force hybrid? |
|---|---|---|
| Residency | The data is stored in a specific region | No: europe-west1 or europe-southwest1 satisfy it |
| Sovereignty | The data is beyond the legal reach of third countries | Sometimes: sovereign cloud offerings exist |
| Physical control | Nobody but you touches the hardware | Yes, by definition |
The vast majority of Spanish companies that say "we cannot go to the cloud because of the law" are at level 1, which is solved by choosing a region. The GDPR does not prohibit the public cloud: it requires appropriate processing, a data processor under contract, technical measures and control of international transfers. Only specific sectors — defence, certain public administrations, healthcare in some circumstances — reach level 3.
Warning. Interpreting the regulatory requirements that apply to your data is not a technical decision. Any design handling personal, health or financial data must be reviewed by a security and compliance professional before going to production.
Avoiding vendor lock-in
The brochure reason: "do not tie yourself to one vendor".
The real version, and the most misinterpreted of the five: lock-in is not binary, it has degrees, and avoiding it entirely costs more than it saves.
| Level of coupling | Example | Cost of leaving |
|---|---|---|
| Very low | Standard containers on Kubernetes | Days |
| Low | Cloud Run, Cloud Storage (S3-compatible API elsewhere) | Weeks |
| Medium | Cloud SQL PostgreSQL, Pub/Sub | Months |
| High | BigQuery, Spanner, Vertex AI | Rewrite |
| Total | Classic App Engine standard | Rewrite |
The sensible strategy is not "use only what is portable", because that means giving up BigQuery, which is probably the best reason to be on Google Cloud. The sensible strategy is to know which level you are at for each piece and for that to be a conscious decision. AlpinaShop is heavily coupled in analytics (BigQuery) and barely coupled in the application (containers). That is fine: the analytics would be rebuilt in six months if it came to it, the shop would move in two weeks.
And an uncomfortable fact: most companies that build a multicloud architecture "so as not to depend on anybody" end up depending on the layer that unifies both clouds — which is usually another vendor — and adding a dependency rather than removing one.
Business continuity
The brochure reason: "if one cloud goes down, you keep running".
The real version: total outages of a provider are extraordinarily rare; regional outages are far more frequent, and you do not need another cloud for that: another region is enough. Building active-active multicloud to withstand a global Google outage means paying double the complexity all year round to cover an event that may never happen. It is covered in detail in 07-06, with RTO and RPO.
- The hidden costs the brochure does not mention
Every model has a cost that does not show up in the price comparison.
| Model | Obvious cost | Hidden costs |
|---|---|---|
| Single cloud | Monthly bills | Coupling; weak negotiating position |
| Hybrid | Cloud + hardware + links | Two operational disciplines; staff with two profiles; the link is a single point of failure; hardware patching; your own 24×7 on-call |
| Multicloud | Two bills | Two different IAM systems; two ways of doing networking; two billing models; egress between clouds (expensive); the team half-masters two ecosystems instead of one properly |
| Edge | Devices | Physical deployment; remote updates; physical security; inventory |
The most expensive hidden cost of all is human. AlpinaShop has one infrastructure person. A genuinely managed hybrid model — with its service mesh, its replicated policies and its redundant link — requires that person to be an expert in Kubernetes, in enterprise networking, in the on-site hardware and in Google Cloud, and to be available when it fails at three in the morning. That is why most small-business hybrid architectures fail, and it has nothing to do with the technology.
- The AlpinaShop case: Sabadell, the ERP and the warehouse terminal
Let us put the specific case on the table, with the facts Marta gathered.
| System | What it is | Why it is not migrating now |
|---|---|---|
| ERP | Desktop application with MySQL 5.7 on a server in the back room | Vendor licence tied to the physical server; the vendor charges for the cloud version and there is no budget this year |
| Warehouse terminal | Radio-frequency reader that talks over the LAN to the ERP | Closed firmware, it only knows how to talk to a local IP |
| Shop till system | Point of sale integrated with the ERP | Must be able to take payments even if the internet goes down |
What the online shop needs out of all that is far less than it seems:
- The real stock every few minutes, so as not to sell what is no longer there.
- The web orders pushed to the ERP, so the accounts add up.
- Nothing else. Not the supplier catalogue, not the payroll, not the history back to 2012.
And that is the observation that decides the whole lesson: you do not need to unify two platforms in order to move two data flows. If the requirement were "run the same containerised applications in the shop and in the cloud, with the same policies and the same mesh", Anthos would be the answer. The real requirement is "synchronise two tables and send some orders", and that is solved with a tunnel and a scheduled process.
Even so, we are going to study Anthos thoroughly. Because the decision not to use it is only worth anything if it is made knowing what is being turned down.
- What Anthos is and what it has become
Anthos was launched in 2019 as Google's application platform for hybrid and multicloud environments. Its central idea was powerful and still is:
If your applications run on Kubernetes, the physical place where that Kubernetes runs stops mattering. Google gives you a single control plane to manage them all — on GCP, in your data centre, on AWS or on Azure — with the same configuration, the same policies and the same observability.
Kubernetes as the common denominator. That is the entire thesis.
The evolution of the names
Here is the part that confuses everybody. The name "Anthos" has been retired progressively and the offering has been reorganised into two families:
| Name | Era | What it is today |
|---|---|---|
| Anthos | 2019-2023 | The original brand. It still appears in documentation, certifications, blogs and in this course's syllabus |
| GKE Enterprise | Since 2023 | The enterprise tier of GKE: fleets, Config Sync, Policy Controller, Cloud Service Mesh, multi-cluster dashboards. It is "Anthos" turned into an edition of GKE |
| Google Distributed Cloud (GDC) | Since 2023 | Google's software and hardware outside its own data centres: in your data centre, at the edge, or in air-gapped versions for sovereign environments |
| Cloud Service Mesh | Since 2024 | The managed service mesh, previously "Anthos Service Mesh" and before that "Istio on GKE" |
| Anthos clusters on AWS/Azure | — | Today GKE Multi-Cloud: Google-managed GKE clusters running on another cloud's machines |
The practical way to remember it:
- If you are talking about managing clusters in a unified way, the current name is GKE Enterprise.
- If you are talking about Google hardware or software running outside Google, the current name is Google Distributed Cloud.
- If somebody says "Anthos", they mean one of the two and probably learned the subject before 2023.
In this lesson we will use the current naming, flagging the old name where it helps you find documentation.
flowchart TB
A["Anthos<br/>2019-2023"] --> B["GKE Enterprise<br/>fleets, policies, mesh"]
A --> C["Google Distributed Cloud<br/>connected, air gapped, edge"]
A --> D["GKE Multi-Cloud<br/>GKE on AWS and Azure"]
B --> E["Cloud Service Mesh<br/>previously Anthos Service Mesh"]
- The fleet: the unit of management
The concept you have to understand before any other is the fleet.
A fleet is a logical group of Kubernetes clusters that are managed together and that trust one another.
A fleet:
- Belongs to one project, called the fleet host project.
- Can contain GKE clusters, GKE on-prem clusters, GKE Multi-Cloud clusters, EKS, AKS, or any conformant Kubernetes cluster registered with the
gke-connectagent. - Is the boundary for enforcing policy, identity and configuration.
The most important and least obvious consequence is the one about fleet namespaces:
In a fleet, the namespace
tiendameans the same thing in every cluster. It is "the shop team", not "a folder inside one particular cluster".
This is called namespace sameness and it has enormous practical effects:
- A permission granted on the
tiendanamespace at fleet level applies across all ten clusters. - A
catalogoservice in thetiendanamespace can be discovered from any cluster in the fleet. - Nobody can create a
tiendanamespace on another cluster for "something else": the name is reserved for that team across the whole fleet.
flowchart TB
subgraph FLOTA["Fleet - project alpinashop-prod"]
direction LR
subgraph C1["alpinashop-cluster (GKE, europe-west1)"]
N1["ns tienda"]
N2["ns pagos"]
end
subgraph C2["cluster-sabadell (GDC, physical shop)"]
N3["ns tienda"]
N4["ns pagos"]
end
end
P["Fleet policies<br/>and configuration"] --> N1
P --> N2
P --> N3
P --> N4
Without a fleet, each cluster is an island with its own RBAC, its own policies and its own criteria. With a fleet, there is one shared mental model. That is the product's real value, more than any specific feature.
- The components, one by one
GKE Enterprise is not a service: it is a set of pieces that are enabled separately. These are they, with their exact function.
| Component | Current name | What it does | Without it, what happens? |
|---|---|---|---|
| Fleet registration | Fleet / Connect | Registers the cluster and opens a secure outbound channel to Google | Each cluster is managed by hand |
| Config Sync | Config Sync | Applies to every cluster the configuration declared in a Git repository | Configuration is applied with kubectl and drifts |
| Policy Controller | Policy Controller | Rejects resources that breach policy (based on Gatekeeper/OPA) | Somebody deploys a privileged pod and nobody notices |
| Service mesh | Cloud Service Mesh | mTLS, routing, retries, telemetry between services | Every application implements that on its own |
| Observability | Cloud Logging / Monitoring | Logs and metrics from every cluster in one place | One dashboard per cluster |
| Supply chain | Binary Authorization | Only signed and approved images run | Any image can reach production |
| Identity management | Fleet Workload Identity | Pods on any cluster obtain GCP credentials with no keys | JSON keys handed around |
| Multi-cluster dashboard | GKE Enterprise console | Status, compliance and cost of the whole fleet | Ten open tabs |
Of all of them, the two people adopt first — and rightly so — are Config Sync and Policy Controller. And they are also the two you can take advantage of on a single cluster, with no hybrid involved.
- Config Sync: configuration as code with GitOps
Config Sync is an agent that runs inside the cluster and does one thing, very well:
It watches a Git repository and guarantees that the cluster's state matches what is in that repository. If somebody changes something by hand, it reverts it.
That is GitOps, and it is worth understanding why it is different from a deployment pipeline like the one AlpinaShop built in 06-01.
| Deployment pipeline (Cloud Build) | GitOps (Config Sync) | |
|---|---|---|
| Direction | Push: the pipeline runs kubectl apply from outside |
Pull: the agent reads Git from inside |
| Credentials | The pipeline needs permissions on the cluster | The cluster exposes credentials to nobody |
| Drift | If somebody changes something by hand, it stays changed | It is reverted within minutes |
| New clusters | They have to be added to the pipeline | They register and synchronise themselves |
| Auditing | The history lives in the pipeline | The history is Git's |
The decisive advantage in a hybrid scenario is the third column of the first row: the cluster in the Sabadell shop does not need to be reachable from the internet. It reaches out to Google, reads its configuration and applies it. There is no need to open inbound ports towards the shop, which is exactly what you do not want to do.
The typical repository structure, which at AlpinaShop would live in alpinashop-infra:
alpinashop-infra/
└── config-flota/
├── cluster/ # applies to ALL clusters
│ ├── politicas/
│ │ ├── no-privilegiados.yaml
│ │ ├── solo-artifact-registry.yaml
│ │ └── etiquetas-obligatorias.yaml
│ └── rbac/
│ └── grupo-desarrollo.yaml
├── namespaces/
│ ├── tienda/
│ │ ├── namespace.yaml
│ │ ├── cuota-recursos.yaml
│ │ └── politica-red.yaml
│ └── pagos/
│ ├── namespace.yaml
│ └── politica-red.yaml
└── system/
└── reposync.yamlTwo key folders: cluster/ for what applies to all of them, and namespaces/<name>/ for what applies to a specific namespace across every cluster in the fleet. The namespace sameness of section 6 becomes concrete here.
- Policy Controller: policy as code
Policy Controller is the managed implementation of Gatekeeper, which in turn builds on OPA (Open Policy Agent). It works as an admission webhook: every time somebody tries to create or modify a resource in the cluster, the request goes through it and can be rejected.
It has two objects:
ConstraintTemplate: the template of the rule, written in a language called Rego. It defines the kind of check.Constraint: the instance of that template with specific parameters and its scope.
Google publishes a template library with more than a hundred ready-made rules, so in practice you almost never have to write Rego. Examples from the library:
| Template | What it prevents |
|---|---|
K8sPSPPrivilegedContainer |
Privileged containers |
K8sAllowedRepos |
Images from unauthorised registries |
K8sRequiredLabels |
Resources without the mandatory labels |
K8sContainerLimits |
Containers without CPU or memory limits |
K8sBlockNodePort |
Services of type NodePort |
K8sRequiredProbes |
Containers without liveness and readiness probes |
And it has a mechanism you must always use at the beginning:
enforcementAction |
Effect |
|---|---|
warn |
Lets it through and warns in the response |
dryrun |
Lets it through and records the violation so you can measure the impact |
deny |
Rejects creation of the resource |
The right pattern is the same one already used with Cloud Armor in 03-05: dryrun first, measure for a week, and only then deny. Enabling a policy in deny without measuring is the fastest way to stop the development team deploying on a Friday afternoon and to get Policy Controller uninstalled on the Monday.
- Practical example: registering
alpinashop-cluster and applying a policy
alpinashop-cluster and applying a policyWe are going to do it for real on the cluster AlpinaShop has had since 02-05. It is a useful exercise even if you never adopt GKE Enterprise, because Config Sync and Policy Controller work on a single GKE Autopilot cluster and add value on their own.
Step 1: enable the required APIs
gcloud config set project alpinashop-prod
gcloud services enable \
gkehub.googleapis.com \
anthosconfigmanagement.googleapis.com \
gkeconnect.googleapis.comgkehub.googleapis.comis the fleets API. "Hub" was the original internal name and it has stuck in the API.anthosconfigmanagement.googleapis.commanages Config Sync and Policy Controller (note the old name, which survives in the API even though the product is called something else).gkeconnect.googleapis.comis the outbound connection channel for registered clusters.
Step 2: register the cluster in the fleet
A GKE cluster in the same project as the fleet is registered with a single command:
gcloud container fleet memberships register alpinashop-cluster \
--gke-cluster=europe-west1/alpinashop-cluster \
--enable-workload-identityExplained piece by piece:
memberships registercreates a membership: the object that represents that cluster within the fleet.alpinashop-clusteris the name of the membership. It is a good idea for it to match the cluster's name so you do not lose your mind.--gke-cluster=REGION/NAMEidentifies a GKE cluster. For an external cluster you would use--kubeconfigand--context.--enable-workload-identityis what lets pods obtain Google credentials with no key files, exactly as in 03-04. On an Autopilot cluster it is already active; registration integrates it with the fleet.
Check:
Step 3: enable Config Sync pointing at alpinashop-infra
The feature's configuration is declared in a file and applied to the fleet:
# config-management.yaml
applySpecVersion: 1
spec:
configSync:
enabled: true
sourceFormat: unstructured
git:
syncRepo: https://github.com/alpinashop/alpinashop-infra
syncBranch: main
policyDir: config-flota
secretType: token
syncWait: 30
policyController:
enabled: true
templateLibraryInstalled: true
referentialRulesEnabled: true
auditIntervalSeconds: 60Field by field:
sourceFormat: unstructuredlets you organise the repository however you like. The alternative,hierarchy, imposes the strict structure seen in section 8. To begin with,unstructuredcauses fewer headaches.syncRepo/syncBranch/policyDir: repository, branch and subdirectory from which it synchronises. Anything outsideconfig-flotais ignored.secretType: tokenstates how it authenticates against GitHub. In production you would use a GitHub app or, better,gcpserviceaccountagainst a repository hosted on GCP.syncWait: 30is the number of seconds between checks. Thirty seconds is a good balance; lowering it a lot only generates API calls.templateLibraryInstalled: trueinstalls Google's library of more than a hundred templates. Without this, you have to write the Rego by hand.auditIntervalSeconds: 60is how often Policy Controller reviews the resources that already existed before the policy. This matters: the webhook only sees new things; the periodic audit is what uncovers the old things that breach it.
It is applied like this:
gcloud beta container fleet config-management apply \
--membership=alpinashop-cluster \
--config=config-management.yaml \
--project=alpinashop-prodAnd the status is queried with:
Name Status Last_Synced_Token Sync_Branch Last_Synced_Time alpinashop-cluster SYNCED a3f91c2 main 2026-08-05T09:14:22Z
Last_Synced_Token is the hash of the applied commit. That single fact answers in one second the question "which version of the configuration is running on this cluster?", which in a push model usually requires archaeology in the pipeline logs.
Step 4: the conformance policy
AlpinaShop wants to guarantee something that until now depended on goodwill: only images from its own Artifact Registry run on the cluster. No docker.io/whatever:latest.
The file is added to the alpinashop-infra repository, under config-flota/cluster/politicas/:
# config-flota/cluster/politicas/solo-artifact-registry.yaml
apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sAllowedRepos
metadata:
name: solo-artifact-registry-alpinashop
spec:
enforcementAction: dryrun # FIRST measure, then deny
match:
kinds:
- apiGroups: [""]
kinds: ["Pod"]
excludedNamespaces:
- kube-system
- gke-system
- config-management-system
- gatekeeper-system
parameters:
repos:
- "europe-west1-docker.pkg.dev/alpinashop-prod/alpinashop/"The details that matter:
enforcementAction: dryrun: for a week it records violations without blocking anything.excludedNamespaces: the system namespaces use Google images. If you do not exclude them, you break the cluster. This is the number one mistake with Policy Controller.parameters.reposaccepts prefixes. The trailing slash matters: without it,.../alpinashop-maliciosowould get through too.
And a second policy, this one about the labels AlpinaShop has been using since 01-04:
# config-flota/cluster/politicas/etiquetas-obligatorias.yaml
apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sRequiredLabels
metadata:
name: etiquetas-obligatorias-alpinashop
spec:
enforcementAction: dryrun
match:
kinds:
- apiGroups: ["apps"]
kinds: ["Deployment"]
namespaces: ["tienda", "pagos"]
parameters:
labels:
- key: "aplicacion"
- key: "equipo"
allowedRegex: "^(infra|desarrollo|datos)$"
- key: "entorno"
allowedRegex: "^(prod|dev)$"Note that allowedRegex does not just require the label: it requires its value to be one of the valid ones. Without that, somebody puts equipo: varios and the cost allocation of 07-05 becomes useless again.
Step 5: measure before denying
After a week, the recorded violations are queried:
kubectl get constraint solo-artifact-registry-alpinashop -o json \
| jq '.status.violations[] | {ns: .namespace, name: .name, msg: .message}'{"ns":"tienda","name":"redis-cache-7c9f","msg":"container <redis> has an invalid image repo <docker.io/redis:7>"}
{"ns":"tienda","name":"exportador-metricas-2d1a","msg":"container <exp> has an invalid image repo <quay.io/prometheus/node-exporter>"}Two real findings, and neither of them is an attack: they are two legitimate public images that somebody deployed without thinking about it. The right response is not to give up and remove the policy, but to:
- Copy those two images into your own Artifact Registry (which is also the correct thing to do for availability and vulnerability scanning, as you will see in 07-04).
- Update the deployments.
- Now, change
dryruntodeny, commit and let Config Sync propagate it.
sed -i 's/enforcementAction: dryrun/enforcementAction: deny/' \
config-flota/cluster/politicas/solo-artifact-registry.yaml
git commit -am "Policy: require images from Artifact Registry (deny)"
git pushIn under a minute, the cluster rejects any pod with a foreign image. And the important part: the change is in a reviewed pull request, not in somebody's memory.
- The service mesh explained without the hype
The service mesh is the part of GKE Enterprise that is sold the most and understood the worst. Here is the honest explanation.
What it is technically
A proxy (Envoy) is injected alongside each application container — the sidecar pattern — or on each node (ambient, the more recent mode that avoids the sidecar). All traffic between services passes through those proxies. Because the proxy sees all the traffic and a central control plane configures it, you can do things the application does not know how to do.
flowchart LR
subgraph P1["Pod catalogo"]
A["App catalogo"] <--> B["Envoy proxy"]
end
subgraph P2["Pod pedidos"]
C["Envoy proxy"] <--> D["App pedidos"]
end
B <-->|"automatic mTLS"| C
CP["Control plane<br/>Cloud Service Mesh"] -.->|"configuration"| B
CP -.->|"configuration"| C
B -.->|"telemetry"| CP
C -.->|"telemetry"| CP
What it genuinely solves
| Capability | The problem it solves | Can you do it without a mesh? |
|---|---|---|
| Automatic mTLS | Encryption and authentication between services without touching the code | Yes, but by implementing it in every application and rotating certificates by hand |
| Service-to-service authorisation | "Only catalogo may call pedidos" |
With NetworkPolicy, but at IP level, not identity level |
| Retries and timeouts | Retrying the transient failure without code | Yes, with a library in each language |
| Percentage-based canary deployment | Sending 5 % of traffic to the new version | Yes, and in Cloud Run it is one flag (07-02) |
| Uniform telemetry | Latency and error rate of every call, with no instrumentation | Partially, with OpenTelemetry (06-06) |
| Circuit breaker | Stop calling the service that is down | Yes, with a library |
The two values that are genuinely hard to replicate are automatic mTLS with per-service identities and uniform telemetry without touching the code. The rest is achievable in other ways, often simpler ones.
What complexity it adds
And here is the part the presentations leave out:
- You double the containers. Every pod ends up with two processes. Additional CPU and memory consumption of 10-30 %, and on Autopilot that is straight onto the bill.
- You add a network hop to every call. Additional latency of between 0.5 and 3 ms per hop, multiplied by the depth of the call chain.
- You debug twice as much. When something fails, the question becomes "is it the application or is it the proxy?". Envoy's error codes (
UF,UO,NR,URX) have to be learned. - You introduce new concepts:
VirtualService,DestinationRule,PeerAuthentication,AuthorizationPolicy,Gateway,Sidecar. That is six more objects to master, and how they interact is not obvious. - Start-up order matters. A container that calls the network before its proxy is ready fails in baffling ways.
- Upgrading it is a delicate operation, because it touches the traffic path of everything running on the cluster.
The honest rule: a service mesh starts to pay off from roughly twenty services calling one another, maintained by more than three teams, in more than one language. Below that, the cost of operating it exceeds what it contributes.
AlpinaShop has one catalogue service, an image function and a handful of batch processes. Putting a mesh there is like building an airport for a cycle lane. It is not that it is bad technology: it is that there is no problem to solve.
- Binary Authorization and the chain of trust
Binary Authorization answers a different question: does this specific image have permission to run in this environment? It also works as admission control, but instead of looking at the shape of the resource, it checks cryptographic signatures.
The flow is:
- Cloud Build builds the image and publishes it to Artifact Registry (06-01).
- A verification process — the vulnerability scan, or Marta's manual approval step — creates an attestation: a signature saying "this image, identified by its digest, has passed this check".
- On deployment, Binary Authorization checks that the attestations required by the policy exist. If not, it rejects the pod.
# Minimal policy: require an attestation from the attestor "aprobacion-marta" in production
gcloud container binauthz policy import - <<'EOF'
defaultAdmissionRule:
evaluationMode: REQUIRE_ATTESTATION
enforcementMode: ENFORCED_BLOCK_AND_AUDIT_LOG
requireAttestationsBy:
- projects/alpinashop-cicd/attestors/aprobacion-marta
clusterAdmissionRules:
europe-west1.alpinashop-cluster:
evaluationMode: REQUIRE_ATTESTATION
enforcementMode: ENFORCED_BLOCK_AND_AUDIT_LOG
requireAttestationsBy:
- projects/alpinashop-cicd/attestors/aprobacion-marta
admissionWhitelistPatterns:
- namePattern: europe-west1-docker.pkg.dev/alpinashop-prod/alpinashop/*
EOFImportant details:
ENFORCED_BLOCK_AND_AUDIT_LOGblocks and logs. There isDRYRUN_AUDIT_LOG_ONLY, which is where you should start, just as with Policy Controller.admissionWhitelistPatternsis a list of exceptions by name pattern. Use it carefully: every exception is a permanent hole.- Binary Authorization is not limited to GKE Enterprise: it works on standard GKE and also on Cloud Run, which makes it relevant to AlpinaShop even if it rules out the fleet. It comes back in 07-04 as a piece of the supply chain.
- Unified fleet observability
Registered clusters send their metrics and logs to Cloud Monitoring and Cloud Logging in the fleet project, including those outside Google Cloud. This means the campaign dashboard AlpinaShop built in 06-04 could show, on the same screen, the catalogue latency in europe-west1 and that of the process running in the back room in Sabadell.
It is a real and often underestimated advantage: in a badly built hybrid, 80 % of the time spent on an incident goes on opening three different tools and comparing clocks that are not in sync.
The GKE Enterprise console also adds a fleet view with:
- Synchronisation status of each cluster (the
Last_Synced_Tokenfrom section 10). - Policy compliance percentage and list of violations.
- Kubernetes versions and end-of-support warnings per cluster.
- Estimated cost grouped by team, if the labels are set correctly.
- Licensing and cost model: what decides it in practice
This is where the real decisions get made, and where it pays to be very clear.
GKE Enterprise is not billed by spot consumption like a VM: it is billed as a GKE edition, with a price per node vCPU of the clusters included in the fleet, monthly or with an annual commitment. Google publishes the prices and they change; the order of magnitude to keep in your head is:
| Item | Order of magnitude |
|---|---|
| GKE Standard: control plane | A few tens of euros per month per cluster (with one free zonal cluster per billing account on the free tier) |
| GKE Enterprise: surcharge | Of the order of several euros per vCPU per month, across all the fleet's vCPUs |
| Managed Cloud Service Mesh | Included in the Enterprise edition |
| Google Distributed Cloud | Subscription per node, plus the hardware |
Always verify prices in the official documentation and with the pricing calculator. The figures above are orders of magnitude for reasoning, not quotations.
Let us do the sums for AlpinaShop, which is what decides it:
- The
alpinashop-clustercluster is Autopilot and carries the catalogue with the equivalent of about 8 vCPUs during normal hours. - A surcharge of the order of €5 per vCPU per month comes to about €40 a month in licensing alone, not counting the compute.
- On top of that you would have to add the extra consumption of the mesh sidecars (10-30 % more CPU and memory) and, if you really wanted hybrid, a cluster in Sabadell with its hardware, its subscription and its maintenance.
Forty euros a month does not ruin anybody. The problem is not the money: it is that you are paying for capabilities you are not going to use. And there is a much bigger cost hidden behind it: Marta's hours learning, operating and debugging six new Istio objects for a single service.
- Lighter alternatives to full hybrid
Between "do nothing" and "adopt GKE Enterprise" there is a spectrum. These are the intermediate options, ordered from least to most commitment.
| Option | What it solves | Cost and complexity |
|---|---|---|
| HA Cloud VPN + scheduled process | Private connectivity with the on-site system and data synchronisation | Very low. It is what AlpinaShop needs (07-03) |
| Pub/Sub as a bridge | The on-site system publishes events and the cloud consumes them, with no inbound connectivity | Low. Excellent when the on-site side can reach the internet |
| Config Sync + Policy Controller on a single cluster | Configuration and policy as code, with no hybrid fleet | Low. Adds value with just one cluster |
| Config Connector | Managing GCP resources (buckets, Cloud SQL) from Kubernetes manifests | Medium. Interesting if your team lives in kubectl; if it lives in Terraform (06-07), no |
| Multi-provider Terraform | Managing GCP + another cloud + GitHub + DNS in one language | Low. AlpinaShop already has it |
| GKE Multi-Cloud | Google-managed GKE on AWS or Azure machines | High |
| Google Distributed Cloud | Google software on your hardware or at the edge | Very high |
Two notes:
- "Cloud Run for Anthos", which appears in older documentation and exams, allowed you to run the Cloud Run API on your own cluster via Knative. It has been deprecated in favour of Cloud Run plain and simple (07-02) and the managed mesh. If you see it in a certification syllabus, now you know what was being talked about.
- Config Connector deserves honest consideration: it is a real alternative to Terraform if — and only if — your team already works entirely in Kubernetes. For AlpinaShop, having just spent module 6 investing in Terraform, adding a second system for the same thing would be a step backwards.
- When yes and when no: the decision table
This is the table you take away from the lesson.
| Situation | Managed hybrid / GKE Enterprise? | Alternative |
|---|---|---|
| Data centre with unamortised investment and dozens of applications | Yes | — |
| Legal requirement for physical control of the hardware | Yes, with Google Distributed Cloud | — |
| Processes that must keep working offline (factory, shop, ship) | Yes, at the edge | Bespoke local compute |
| More than 20 microservices, several teams, several languages | Yes for the mesh | — |
| Company merger: one on GCP and the other on AWS, with both teams staying | Probably yes | — |
| I need uniform security policies across 10+ clusters | Yes | — |
| A single cluster and I want configuration as code | No | Config Sync on its own |
| I need to read an on-site database from the cloud | No | HA Cloud VPN (07-03) |
| I want to "avoid vendor lock-in" in the abstract | No | Standard containers and exportable data |
| An infrastructure team of one to three people | No | The complexity eats the team |
| A single web service with irregular traffic | No | Cloud Run (07-02) |
And the single question that resolves most cases:
Do I have workloads running outside Google Cloud that need the same management as the ones inside?
If the answer is "no, I only need to talk to an outside system", you do not need Anthos: you need networking.
- DA-004: AlpinaShop's honest decision
Marta, Dani and Lucía sat down with this information and wrote the decision, following the same format as DA-001, DA-002 and DA-003. It is in the README of alpinashop-infra, as the discipline of 06-02 requires.
DA-004 — Hybrid strategy and integration with the Sabadell shop Date: 2026-08-05 · Authors: Marta (infrastructure), Dani (backend) · Status: accepted
1. Context. The physical shop in Sabadell runs an ERP with MySQL 5.7 on a local server, with a radio-frequency terminal and a till system tied to it over the local network. The ERP cannot migrate this year because of vendor licensing and budget. The online shop needs two things from the ERP: reading the stock every few minutes and pushing web orders across. The infrastructure team is one person.
2. Options considered.
| Option | Advantage | Drawback |
|---|---|---|
| GKE Enterprise with a cluster in Sabadell | Unified management, policies and mesh on both sides | New hardware in the shop; per-vCPU licence; mesh complexity for one service; one person cannot operate it |
| Google Distributed Cloud in the shop | Google's platform on site, Google support | Cost far greater than the problem; oversized for two data flows |
| Migrate the ERP to Cloud SQL now | Removes the hybrid | Blocked by vendor licensing and budget; the till must take payments offline |
| HA Cloud VPN + scheduled synchronisation | Low cost, low complexity, solves the two real flows | Does not unify management; the ERP remains an on-site responsibility |
3. Decision. AlpinaShop does not adopt GKE Enterprise or Google Distributed Cloud. Private connectivity with the shop is established via HA Cloud VPN (two tunnels with BGP over Cloud Router, detailed in 07-03) and the integration is resolved as follows:
- Stock: a scheduled job with Cloud Scheduler + Workflows (04-06) reads the stock tables from the Sabadell MySQL through the tunnel every 10 minutes and updates the catalogue.
- Orders: each web order publishes to the
pedidos-nuevostopic (04-04); a consumer inserts them into the ERP. If the link is down, the messages wait in the subscription and are processed when it comes back — the online shop does not stop because Sabadell is unavailable. - Offline: the till and the terminal keep working against the local ERP exactly as they do today. No new cloud dependency is introduced into the operation of the physical shop.
4. Consequences.
- Positive: marginal cost (that of two VPN tunnels); no new technology to learn; the decoupling via Pub/Sub means a link outage is a delay, not an outage.
- Negative: management remains twofold — the ERP is patched by hand in Sabadell — and there are no uniform policies across the two sides. This is accepted: it is one server, not a fleet.
- One piece is partially adopted: Policy Controller in
dryrunmode onalpinashop-clusterfor as long as the catalogue stays there, because it adds value with a single cluster and without an Enterprise licence at its basic tier. It will be reviewed when the catalogue moves to Cloud Run in 07-02.
5. Review. This decision is reviewed if any of these three things happens: two more physical shops open, the ERP migrates to the cloud, or the number of services calling one another exceeds ten.
Common Mistakes and Tips
- Confusing "hybrid" with "having a VPN". Having private connectivity to an on-site system does not make you a managed hybrid architecture. That is good to know: you probably do not need the big platform.
- Adopting the service mesh "because it is best practice". It is best practice beyond a certain scale. Below that, it is pure operational overhead.
- Enabling Policy Controller straight into
deny. You will block legitimate deployments on day one.dryrunfor a week, measure, fix, and thendeny. - Forgetting
excludedNamespacesin the policies. If the policy applies tokube-system, the cluster stops working. Always excludekube-system,gke-system,gatekeeper-systemandconfig-management-system. - Searching for documentation under the old name and ending up with obsolete instructions. "Anthos Service Mesh" and "Cloud Service Mesh" are the same product in two eras, and the commands have changed. Always check the date on the page.
- Believing that multicloud reduces vendor lock-in. It often increases it: it adds an abstraction layer that also belongs to somebody, and it forces you to use the lowest common denominator of both clouds.
- Ignoring egress between clouds. Moving data from GCP to AWS and back is charged per gigabyte in both directions. An "elegant" multicloud architecture that crosses data constantly can double the bill. It is covered in 07-05.
- Registering clusters in the fleet without
--enable-workload-identity. You will end up handing out JSON keys, exactly what 03-04 taught you to avoid. - Tip: use Config Sync even if you are not going to be hybrid. GitOps on a single cluster already eliminates configuration drift, and it is one of the best value-for-effort improvements in the catalogue.
- Tip: if you are torn between hybrid and migrating, work out the hardware's write-off date. Very often the right answer is "hybrid for 18 months and then cloud", and that completely changes how much platform is worth building.
- Tip: write the decision down. DA-004 exists so that in two years' time nobody has to reconstruct why it was not adopted. And so that, when the review conditions are met, it actually gets reviewed.
Exercises
Exercise 1 — Deciding with judgement across four scenarios
For each of these cases, decide whether GKE Enterprise / Google Distributed Cloud is appropriate, or a lighter alternative, or nothing at all, and justify it in three or four lines.
(a) A chain of 60 supermarkets. Each shop has tills that must be able to take payments even if the internet goes down, and an electronic shelf-labelling system that updates every hour. Head office with a systems team of 8 people.
(b) A 12-person startup with a backend of 4 microservices on GKE Autopilot, all in europe-west1. The CTO wants to "be ready for multicloud".
(c) An insurer that has just bought a smaller company. The buyer is on Google Cloud; the acquired company has 40 applications in its own data centre with a hosting contract running for another 3 years. Both teams are staying.
(d) A manufacturer with a plant in Zaragoza. A computer vision system inspects parts at 30 images per second and must decide in under 50 ms whether to reject the part. They also want to train models on the historical images.
Exercise 2 — Writing and deploying a conformance policy
AlpinaShop wants to guarantee two more things on alpinashop-cluster:
- No container may run as
root. - Every container must declare CPU and memory limits (so Autopilot does not bill surprises and so a runaway pod does not starve the others).
Write the two Constraint objects using the template library (K8sPSPAllowPrivilegeEscalationContainer and K8sContainerLimits), state where to place them in the alpinashop-infra repository, and describe the full safe deployment procedure. Also explain what you would check before moving to deny.
Exercise 3 — Rebutting a vendor's proposal
An integrator presents this proposal to AlpinaShop's management:
"We recommend deploying Google Distributed Cloud in the Sabadell shop and adopting GKE Enterprise with Cloud Service Mesh. That way you will unify management, have end-to-end mTLS between the shop and the cloud, homogeneous policies and centralised observability, leaving you ready for multicloud growth. Investment: €18,000 the first year plus €900/month."
Write Marta's technical response to management: which parts of the proposal are true, which are irrelevant to AlpinaShop, what questions should be put to the integrator, and what counter-proposal is put forward with its approximate cost.
Solutions
Solution 1 — Deciding with judgement across four scenarios
(a) Chain of 60 supermarkets: YES, and it is the textbook case.
All three decisive factors coincide at once. Offline operation: the tills must take payments with the link down, so there has to be local compute out of obligation, not preference. Scale: 60 sites means 60 clusters, and managing them one by one is unworkable — here the fleet stops being a luxury and becomes the only way to stay sane. Team: 8 people can sustain the platform.
The solution is Google Distributed Cloud in each shop, all registered in one fleet, with Config Sync distributing the configuration from Git. Note the advantage of the pull model: the 60 shops do not need to be reachable from outside; they reach out for their own configuration. Updating the shelf-labelling software across 60 shops becomes one commit.
The service mesh, by contrast, would be evaluated separately and probably later: the value here is in fleet management, not in mTLS between services.
(b) 12-person startup: NO, clearly.
"Being ready for multicloud" with 4 microservices on a single cluster is not an architecture: it is a worry with no use case. The cost of the preparation is paid today, every month, and the benefit may never arrive.
What is worth doing, and is free:
- Keep the applications containerised and free of proprietary dependencies on the critical path, which they already are.
- Manage the infrastructure with Terraform, which is multi-provider by design.
- Make the data exportable (standard formats,
pg_dump, Parquet). - Write in the
READMEwhich services are deliberately coupled and why.
With that, the day there is a real reason to move, the exit takes weeks. And in the meantime you make the most of what is good about GCP instead of giving it up. If they also want configuration as code, Config Sync on the cluster they already have is the improvement with the best return.
(c) Insurer after a merger: PROBABLY YES, and it is the most nuanced case.
Several factors coincide: unamortised investment (a 3-year hosting contract, which is committed money), 40 applications (a volume that justifies a platform), and two teams that are staying (there are hands available). Moreover, in a regulated sector, having demonstrably uniform security policies on both sides has audit value, not just technical value.
The sensible strategy is hybrid with an expiry date: register the clusters on both sides in one fleet, unify policy with Policy Controller and observability with Cloud Logging/Monitoring, and use those 3 years to migrate in waves the applications that make sense. The mistake would be to adopt it as a permanent state with no convergence plan; the other mistake would be to try to migrate 40 applications in one go.
The questions that would settle the detail: how many of those 40 applications are containerised? If it is 3, the fleet does not help the other 37 and the real problem is modernisation, not platform.
(d) Manufacturer with computer vision: YES, but at the edge and for a different reason.
A total budget of 50 ms to capture, infer and decide makes the cloud impossible: the round trip to europe-southwest1 from Zaragoza alone consumes a significant part of it, and a network outage would stop the production line. Inference has to be local, and there is no arguing about it.
Architecture: Google Distributed Cloud (or simply hardware with a GPU and an inference runtime) at the plant running the model; the images and results are uploaded asynchronously to Cloud Storage; training happens in Vertex AI (module 5) on the historical data; new models are deployed to the plant from the model registry.
It is the canonical edge pattern: inference below, training above. And note that the part justifying the platform is not the latency in itself, but that the line cannot stop if the link goes down.
Solution 2 — Writing and deploying a conformance policy
Location in the repository. Both files go in config-flota/cluster/politicas/, because they are cluster policies and not policies for a specific namespace:
alpinashop-infra/config-flota/cluster/politicas/ ├── solo-artifact-registry.yaml # already exists ├── etiquetas-obligatorias.yaml # already exists ├── no-escalada-privilegios.yaml # new └── limites-contenedor.yaml # new
Policy 1 — no privilege escalation:
# config-flota/cluster/politicas/no-escalada-privilegios.yaml
apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sPSPAllowPrivilegeEscalationContainer
metadata:
name: no-escalada-privilegios
spec:
enforcementAction: dryrun
match:
kinds:
- apiGroups: [""]
kinds: ["Pod"]
excludedNamespaces:
- kube-system
- gke-system
- gmp-system
- gatekeeper-system
- config-management-systemPolicy 2 — CPU and memory limits:
# config-flota/cluster/politicas/limites-contenedor.yaml
apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sContainerLimits
metadata:
name: limites-obligatorios
spec:
enforcementAction: dryrun
match:
kinds:
- apiGroups: [""]
kinds: ["Pod"]
namespaces: ["tienda", "pagos"]
parameters:
cpu: "2"
memory: "4Gi"Note that K8sContainerLimits does two things at once, and it is worth knowing: it requires the limits to exist and it also requires them not to exceed the stated values. The values 2 vCPU and 4Gi are the ceiling per container; they are chosen by measuring real consumption and leaving headroom, not by eye.
Safe deployment procedure, step by step:
- Branch and pull request in
alpinashop-infra. Never straight tomain: review is part of the control (06-02). - Local validation before pushing, so you do not discover a syntax error in production:
kubectl apply --dry-run=server -f config-flota/cluster/politicas/ - Merge to
mainafter review. Config Sync picks it up within 30 seconds. - Verify the synchronisation, which is the step people skip:
gcloud beta container fleet config-management status nomos status # Config Sync specific tool, more detailed - Wait between 7 and 14 days in
dryrun, covering at least one full deployment cycle plus the nightly and weekly batch processes. A week that does not include the monthly close can hide a breach. - Review violations:
kubectl get constraints -o json | jq -r ' .items[] | select(.status.totalViolations > 0) | "\\(.metadata.name): \\(.status.totalViolations) violations"' - Fix the source, not the policy. If a pod legitimately needs more than 2 vCPUs, raise the parameter with written justification; if it does not need it, fix the pod.
- Switch to
denyin a pull request separate from the one that introduced the policy, so the change of mode is visible in the history and can be reverted on its own.
What to check before moving to deny — Marta's list:
- Zero violations for at least 7 consecutive days.
- The weekly and monthly
CronJobruns have executed at least once in that period. - The system namespaces are excluded and this has been verified, not assumed.
- There is a documented exception procedure: how somebody requests an exemption, who approves it and when it expires.
- The rollback procedure is written down:
git revertthe commit and wait 30 seconds. Knowing that going back costs a minute is what lets you move forward calmly.
Solution 3 — Rebutting a vendor's proposal
Marta's response to management.
What is true about the proposal. Everything technical. GKE Enterprise genuinely unifies cluster management, Cloud Service Mesh provides automatic mTLS without touching the code, Config Sync and Policy Controller enforce homogeneous policies, and observability is centralised. The integrator is not exaggerating any capability. The problem is not that it is false; it is that it answers questions we have not asked.
What is irrelevant to us, point by point.
- "You will unify management": unify the management of one cluster in the cloud and one ERP server that does not even run containers. There is no fleet to manage; there are two different systems that only need to exchange two data flows.
- "End-to-end mTLS between the shop and the cloud": an IPsec Cloud VPN tunnel already encrypts that traffic, and it costs a few euros a month. The mesh's mTLS solves encryption between services, and we have one service.
- "Homogeneous policies": valuable from several clusters upwards. With one, Policy Controller does contribute — and we are going to use it — but that does not require the Enterprise licence or a deployment in the shop.
- "Ready for multicloud growth": there is no business plan contemplating another cloud. We would be paying today for an option nobody has asked for.
- It also introduces a new risk the proposal does not mention: Google hardware in the back room of a shop that today runs on one server and a UPS. That is one more thing that can break, need updating and fall out of support, in a place with no technical staff.
Questions for the integrator. These are the ones that reveal whether the proposal was made for us or is a template:
- Which containerised workloads are we going to run in the Sabadell shop? (Real answer: none. The ERP is a desktop application with MySQL.)
- How many services call one another in our current architecture? (Real answer: practically one. The mesh has nothing to mesh.)
- Does the €900/month include support, hardware updates for the shop and the operating hours, or only the licence?
- What happens the day the catalogue moves to Cloud Run as per DA-001? Much of this proposal stops applying, because there will no longer be a cluster to manage.
- Who operates this? We are one infrastructure person. Is a 24×7 managed service included, and at what price?
- What is the cost and the timescale of getting out of this architecture if it does not fit in two years?
Counter-proposal.
| Real need | Proposed solution | Approximate cost |
|---|---|---|
| Private, encrypted connectivity with Sabadell | HA Cloud VPN, two tunnels with BGP | Of the order of tens of €/month |
| Synchronise stock every 10 min | Workflows + Cloud Scheduler | Pennies a month |
| Push orders to the ERP tolerating outages | Pub/Sub pedidos-nuevos with retries and DLQ |
Already exists |
| Policies on the cluster | Policy Controller, while there is still a cluster | Included in the base tier |
| Joint observability | We already have Cloud Monitoring and Logging (06-04, 06-06) | Already exists |
Total cost of the order of tens of euros a month, against €18,000 plus €900/month. The difference is not saved: it is invested in what we actually lack — reliability, SLOs and disaster recovery (07-06) — which is where the real risk sits today.
The sentence for management. We are not rejecting the proposal because it is expensive, but because it solves the problem of a company we are not. If we open five more shops and the ERP moves to containers, we will look at it again — and that is why DA-004 states in writing when it gets reviewed.
Conclusion
You have grasped the full map of hybrid and multicloud, and — more importantly — you know how to place yourself on it.
You can distinguish hybrid, multicloud, edge and that fourth category nobody names, accidental multicloud, which is almost everybody's real situation. You know the five reasons for not being on one cloud alone in their honest version: the prior investment that expires by itself, the latency whose strong argument is not the milliseconds but offline operation, the regulation with its three levels — residency, sovereignty and physical control — that almost nobody distinguishes, the vendor lock-in that has degrees and whose correct management is knowing which one you are at, and the continuity that is almost always solved with another region rather than another cloud. And you know the hidden costs of each model, with the most expensive one underlined: the human one.
You know what Anthos is and what it has become: GKE Enterprise for unified cluster management, Google Distributed Cloud for Google software outside Google, GKE Multi-Cloud for clusters on AWS and Azure, and Cloud Service Mesh for the mesh. You will recognise the old name in documentation and exams and you will know how to translate it.
You have mastered the central concept, the fleet, and its least obvious and most powerful consequence: namespace sameness, which turns tienda into "the shop team" across every cluster at once. You know its components one by one, with Config Sync — pull-based GitOps, with the decisive advantage of not requiring inbound access to the remote sites — and Policy Controller — managed Gatekeeper, with its template library and the indispensable dryrun before deny — as the two that add value even with a single cluster.
You have done it in practice: registering alpinashop-cluster in a fleet, enabling Config Sync against alpinashop-infra, writing two real policies — only images from your own Artifact Registry and mandatory labels with validated values — measuring their violations in dryrun, fixing the source and only then denying. With the number one mistake flagged: always exclude the system namespaces.
You understand the service mesh without the hype: what it genuinely solves — automatic mTLS with per-service identity and uniform telemetry without touching the code — and what complexity it adds — duplicated containers, one more network hop, six new objects and harder debugging — with the honest threshold of around twenty services and several teams below which it does not pay off. And you know Binary Authorization as signature-based admission control, which will return in 07-04 and which, unlike the mesh, does apply to Cloud Run.
You know how it is billed — per fleet vCPU, as a GKE edition — and why the price is not what decides it: what decides it is that you pay for capabilities that are not going to be used and for learning hours you do not have. You know the lighter alternatives, from Cloud VPN and Pub/Sub as a bridge through to Config Connector, with the historical note about Cloud Run for Anthos. And you take away the table of when yes and when no with its summary question: do I have workloads running outside GCP that need the same management? If you only need to talk to an outside system, you do not need Anthos: you need networking.
And you have DA-004 written down: AlpinaShop does not adopt GKE Enterprise. It connects to Sabadell over HA Cloud VPN, synchronises the stock with Workflows every ten minutes, pushes the orders through Pub/Sub tolerating outages, keeps the till working offline, adopts Policy Controller in dryrun because it contributes with a single cluster, and puts in writing the three conditions that would force the decision to be reviewed.
What that decision implies remains outstanding on two fronts. One is the network: the tunnel to Sabadell, the Cloud Router, the addressing that must not overlap and everything that has to be designed so that two networks born separately can talk to each other without surprises — that is 07-03.
The other is older. This lesson has decided not to build a bigger container platform; the next one does exactly the opposite and builds a smaller one. Because if AlpinaShop has a single web service, with irregular traffic, already containerised and stateless, the question pending since 02-07 has an answer that can no longer be postponed. DA-001 is fulfilled in the next lesson: the catalogue is moving to Cloud Run.
Google Cloud Platform (GCP) Course
Module 1: Introduction to Google Cloud Platform
- What is Google Cloud Platform?
- Setting Up Your GCP Account
- A Tour of the GCP Console
- Projects, Resource Hierarchy and Billing
- Regions, Zones and the Shared Responsibility Model
- Cloud Shell and the gcloud CLI
Module 2: Core GCP Services
- Compute Engine: Virtual Machines on Google Cloud
- Cloud Storage: Object Storage
- Cloud SQL: Managed Relational Databases
- App Engine: Platform as a Service
- Google Kubernetes Engine (GKE)
- NoSQL Databases: Firestore, Bigtable and Spanner
- How to Choose the Right Compute Service
Module 3: Networking and Security
- VPC Networks
- Cloud Load Balancing
- Cloud CDN
- Identity and Access Management (IAM)
- Cloud Armor
- Secrets and Encryption: Secret Manager and Cloud KMS
- Cloud DNS, TLS Certificates and Publishing Services Securely
Module 4: Data and Analytics
- BigQuery: The Analytical Data Warehouse
- Cloud Dataflow: Batch and Streaming Data Processing
- Cloud Dataproc: Managed Spark and Hadoop
- Cloud Pub/Sub: Asynchronous Messaging
- Cloud Data Fusion: Code-Free Data Integration
- Orchestrating Pipelines with Cloud Composer and Workflows
- Data Governance and Dashboards with Dataplex and Looker Studio
Module 5: Machine Learning and AI
- Vertex AI: The Machine Learning Platform on GCP
- AutoML: Custom Models Without Writing Code
- TensorFlow on GCP: Training and Serving Models
- Natural Language API
- Vision API
- Generative AI on Vertex AI: Gemini Models and Embeddings
- MLOps: From Model to Product with Vertex AI Pipelines
Module 6: DevOps and Monitoring
- Cloud Build: Continuous Integration on GCP
- Cloud Source Repositories and Source Code Management
- Cloud Functions: Serverless Functions
- Cloud Monitoring (formerly Stackdriver): Metrics, Dashboards and Alerts
- Cloud Deployment Manager and Native Infrastructure as Code
- Cloud Logging and Cloud Trace: Logs, Traces and Diagnostics
- Terraform on GCP: Infrastructure as Code in Practice
Module 7: Advanced GCP Topics
- Hybrid and Multicloud with Anthos
- Serverless Computing with Cloud Run
- Advanced Networking: Shared VPC, Peering and Hybrid Connectivity
- Security Best Practices
- Cost Management and Optimization
- Reliability: SLOs, High Availability and Disaster Recovery
- Governance at Scale: Organization, Policies and Auditing
