We closed the previous lesson with a conclusion: the investment in App Engine does not compound outside App Engine, but the investment in containers is useful everywhere. Kubernetes is the de facto standard for orchestrating containers, and Google Kubernetes Engine is Google's managed implementation — significantly, the house where Kubernetes was born, derived from the internal Borg system.

Kubernetes has a reputation for complexity, and it is partly deserved. But that complexity answers a real problem: when you have dozens of containers spread over several machines, someone has to decide where each one runs, restart the ones that fail, distribute traffic, update without interrupting the service and grow when a peak arrives. If Kubernetes does not do it, you do it by hand.

In this lesson you will learn Kubernetes from scratch with just enough to work, create the alpinashop-cluster cluster, publish the Flask catalogue image in Artifact Registry and deploy alpinashop-web with complete manifests, autoscaling, configuration, secrets and interruption-free updates together with their rollback.

Contents

  1. Kubernetes in ten minutes: the objects that matter
  2. Nodes and control plane
  3. What GKE adds
  4. Autopilot versus Standard
  5. Creating alpinashop-cluster and connecting kubectl
  6. Building and publishing the image in Artifact Registry
  7. The alpinashop-web Deployment
  8. The Service: exposing the application
  9. Scaling: manual, HPA and cluster autoscaling
  10. ConfigMaps and Secrets
  11. Interruption-free updates and rollback
  12. Basic cluster observability

  1. Kubernetes in ten minutes: the objects that matter

Kubernetes is a declarative system: you describe the desired state in YAML files and the system works continuously to make reality match that description. If you ask for three replicas and one dies, Kubernetes creates another. You do not give orders; you declare goals.

The essential objects:

Object What it is Analogy
Container A process packaged with all its dependencies The application and its environment, in a box
Pod The smallest unit Kubernetes deploys: one or more containers sharing network and storage A disposable "logical server"
ReplicaSet Guarantees that N identical pods exist The one that counts and replenishes
Deployment Manages ReplicaSets and orchestrates updates What you actually write
Service A stable name and IP that distribute traffic across pods The internal load balancer
Ingress HTTP(S) routing from outside towards several Services The reverse proxy
Namespace A logical partition of the cluster A folder with permissions and quotas
ConfigMap Non-sensitive configuration A configuration file
Secret Sensitive data (with caveats, section 10) A sealed envelope, not an armoured one
Node A machine (VM) that runs pods The physical server

Three ideas to internalise before writing any YAML:

  • Pods are ephemeral and disposable. They are born, they die and they are recreated with another name and another IP. Never connect to a pod by its IP: that is what the Service is for. It is the same lesson as the MIG instances in 02-01, taken to the extreme.
  • You almost never create pods directly. You create a Deployment, which creates a ReplicaSet, which creates the pods. Each layer adds a guarantee.
  • Everything is identified through labels. A Service does not know its pods by name: it selects the ones carrying a particular label. If the selector does not match the pods' labels, the Service exists but sends traffic to nobody. It is the number one beginner's mistake.
graph TD
    subgraph "Control plane (managed by Google)"
        API[API Server]
        SCHED[Scheduler]
        CM[Controller Manager]
        ETCD[(etcd)]
    end

    subgraph "Nodes (Compute Engine VMs)"
        subgraph "Node 1"
            P1[Pod alpinashop-web]
            P2[Pod alpinashop-web]
        end
        subgraph "Node 2"
            P3[Pod alpinashop-web]
        end
    end

    DEP[Deployment<br/>alpinashop-web] --> RS[ReplicaSet]
    RS --> P1
    RS --> P2
    RS --> P3

    SVC[Service LoadBalancer<br/>alpinashop-web] --> P1
    SVC --> P2
    SVC --> P3

    KUBECTL[kubectl] --> API
    API --> SCHED
    API --> CM
    API --> ETCD
    INTERNET((Internet)) --> SVC

  1. Nodes and control plane

A Kubernetes cluster has two halves:

The control plane is the brain. It contains the API Server (the only way in: kubectl and everything else talks to it), the Scheduler (which decides what node each pod goes to), the Controller Managers (the loops that compare actual state with desired state and act) and etcd (the database that holds the cluster's entire state).

The nodes are the machines that run the pods. In GKE they are Compute Engine VMs — the same ones from lesson 02-01 — each with the kubelet (the agent that talks to the control plane and starts containers), a container runtime and kube-proxy (which implements the Services' networking).

Setting up and maintaining a control plane yourself is considerable work: high availability of etcd, certificates, coordinated updates, backups. That is exactly the part GKE takes away from you.

  1. What GKE adds

On top of vanilla Kubernetes, GKE contributes:

  • A managed control plane with an SLA. Google deploys it, replicates it, patches it and updates it. In regional mode, replicated across several zones.
  • Automatic updates of the control plane and the nodes, with release channels (rapid, regular, stable) so you can choose how much novelty you want.
  • Automatic node repair: a node that stops responding is recreated.
  • Node autoscaling: if no more pods fit, nodes are added; if there are spare ones, they are removed.
  • Integration with Google's network: LoadBalancer-type Services create native Google Cloud load balancers, and pods get IPs from the VPC (03-01).
  • Integration with IAM and Workload Identity: pods authenticate to Google Cloud APIs with an identity of their own, without keys.
  • Cloud Logging and Cloud Monitoring built in from the very first moment.
  • Autopilot: a mode where you do not even see the nodes.

  1. Autopilot versus Standard

It is the first decision when creating a cluster, and it shapes day-to-day work.

Autopilot Standard
Who manages the nodes? Google, entirely You: size, number, image, pools
Do you see the VMs? No Yes, in Compute Engine
Billing By the CPU, memory and disk requested by your pods By the node VMs, used or not
Node scaling Automatic and invisible Autoscaler configurable per pool
Security Hardened by default (no privileges, no host access) Configurable, more permissive
DaemonSets and host access Limited Allowed
GPUs and special hardware Supported with restrictions Full control
Operational overhead Minimal Medium-high
When to choose it The general case, small teams, standard applications Specific hardware needs, privileged agents, fine-grained cost tuning at scale

The billing difference is the key to understanding it. In Standard you pay for the nodes: if you have three e2-standard-4 VMs and your pods use 15 %, you pay 100 %. In Autopilot you pay what your pods request: if a pod asks for 250 mCPU and 512 MiB, that is what is billed. This has two direct consequences: the requests in your manifests stop being a recommendation and become your bill, and there is no incentive to "fill up" nodes.

AlpinaShop chooses Autopilot. Marta is the only person responsible for infrastructure in a 40-person company; she has no time to size node pools or tune the binpacking. The application is a standard web container with no special requirements. Autopilot removes a whole category of work in exchange for a per-unit surcharge that, at this scale, is irrelevant next to Marta's hours.

  1. Creating alpinashop-cluster and connecting kubectl

gcloud services enable container.googleapis.com artifactregistry.googleapis.com

gcloud container clusters create-auto alpinashop-cluster \
  --project=alpinashop-prod \
  --region=europe-west1 \
  --release-channel=regular \
  --labels=entorno=prod,equipo=plataforma,centro-coste=tienda,aplicacion=catalogo

Notes on the command:

  • create-auto creates an Autopilot cluster. For Standard it would be create with all the node pool configuration.
  • --region (not --zone): Autopilot clusters are always regional, with the control plane replicated across zones. It is the right choice for production and consistent with the reasoning in 01-05.
  • --release-channel=regular: stable versions with automatic updates. stable is more conservative and rapid gives early access to new versions.

Creation takes between 5 and 10 minutes. Afterwards you have to tell kubectl which cluster to talk to:

# Install the authentication plugin if it is missing (Cloud Shell already has it)
gcloud components install gke-gcloud-auth-plugin

# Get credentials: writes the configuration into ~/.kube/config
gcloud container clusters get-credentials alpinashop-cluster \
  --region=europe-west1 --project=alpinashop-prod

# Check the connection
kubectl cluster-info
kubectl get nodes
kubectl get namespaces

On Autopilot, kubectl get nodes may return an empty list at first: the nodes appear when there are pods to run. It is disconcerting the first time and it is exactly the expected behaviour.

We create a namespace of our own, instead of working in default:

kubectl create namespace tienda
kubectl config set-context --current --namespace=tienda

Namespaces let you separate environments and teams within a cluster, with their own resource quotas and RBAC permissions. Using default for everything is a bad habit that you pay for when the cluster grows.

  1. Building and publishing the image in Artifact Registry

Kubernetes runs container images, so the Flask catalogue has to be packaged. Artifact Registry is Google Cloud's artefact registry — the successor to Container Registry, which has been retired; if you find documentation using gcr.io, it is out of date.

gcloud artifacts repositories create alpinashop \
  --repository-format=docker \
  --location=europe-west1 \
  --description="AlpinaShop container images"

# Configure Docker to authenticate against this registry
gcloud auth configure-docker europe-west1-docker.pkg.dev

The catalogue's Dockerfile:

# Slim base image: smaller than the full one, with less attack surface
FROM python:3.12-slim

# Stops Python writing .pyc files and forces unbuffered output,
# essential for the logs to reach Cloud Logging in real time.
ENV PYTHONDONTWRITEBYTECODE=1 \
    PYTHONUNBUFFERED=1

WORKDIR /app

# Copy ONLY requirements.txt first: if it does not change, Docker reuses
# the dependency layer in subsequent builds. It is the most
# profitable cache optimisation in a Dockerfile.
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt

# Now the code, which changes with every commit
COPY . .

# Unprivileged user: Autopilot REJECTS containers that
# try to run as root.
RUN useradd --create-home --uid 1000 alpina && chown -R alpina:alpina /app
USER 1000

EXPOSE 8080

# 2 workers and 4 threads: the app spends a lot of time waiting for the database,
# so the threads make good use of that dead time.
CMD ["gunicorn", "--bind", "0.0.0.0:8080", "--workers", "2", \
     "--threads", "4", "--timeout", "60", "main:app"]

Building and publishing. There are two routes:

# Option A: build locally with Docker
IMAGE="europe-west1-docker.pkg.dev/alpinashop-prod/alpinashop/catalogo:v1"
docker build -t "$IMAGE" .
docker push "$IMAGE"

# Option B: build in the cloud with Cloud Build (no local Docker required)
gcloud builds submit --tag "$IMAGE" .

Option B is especially convenient from Cloud Shell and is the seed of the continuous integration pipeline we will set up in 06-01.

Check what has been published:

gcloud artifacts docker images list \
  europe-west1-docker.pkg.dev/alpinashop-prod/alpinashop

# Vulnerability scanning (if enabled on the repository)
gcloud artifacts docker images scan "$IMAGE"

Two good tagging practices: never use :latest in production — you will not know which version is running and you will not be able to roll back — and tag with the commit hash (catalogo:a3f9c1d), which makes it traceable which code is in each pod. It is the same idea of immutable names that we applied to Storage objects in 02-02 and to instance templates in 02-01.

  1. The alpinashop-web Deployment

Now the main manifest. Every block is commented:

# deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: alpinashop-web
  namespace: tienda
  labels:
    app: alpinashop-web
    entorno: prod
spec:
  # Desired number of pods. The HPA (section 9) will override it later.
  replicas: 3

  # The Deployment governs the pods matching this selector.
  # It MUST match the labels of the template below.
  selector:
    matchLabels:
      app: alpinashop-web

  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxSurge: 1          # up to 1 extra pod during the update
      maxUnavailable: 0    # never drop below the desired number: zero interruption

  template:
    metadata:
      labels:
        app: alpinashop-web    # label the Service will use
        entorno: prod
    spec:
      # Kubernetes service account bound through Workload Identity
      # to a Google Cloud service account (section 10).
      serviceAccountName: sa-catalogo

      containers:
        - name: catalogo
          image: europe-west1-docker.pkg.dev/alpinashop-prod/alpinashop/catalogo:v1

          ports:
            - name: http
              containerPort: 8080

          # requests: what the scheduler reserves. On Autopilot, WHAT YOU PAY.
          # limits: the ceiling. If memory is exceeded, the container dies (OOMKilled).
          resources:
            requests:
              cpu: "250m"          # 0.25 of a vCPU
              memory: "512Mi"
              ephemeral-storage: "1Gi"
            limits:
              cpu: "500m"
              memory: "512Mi"      # same as requests: avoids memory surprises

          env:
            - name: BUCKET_CATALOGO
              value: "alpinashop-catalogo"
            - name: ENTORNO
              valueFrom:
                configMapKeyRef:
                  name: config-catalogo
                  key: entorno
            - name: INSTANCIA_SQL
              valueFrom:
                configMapKeyRef:
                  name: config-catalogo
                  key: instancia_sql
            - name: DB_USER
              valueFrom:
                configMapKeyRef:
                  name: config-catalogo
                  key: db_user
            - name: DB_PASS
              valueFrom:
                secretKeyRef:
                  name: secreto-catalogo
                  key: db_pass

          # Is the container alive? If this fails, Kubernetes RESTARTS it.
          livenessProbe:
            httpGet:
              path: /salud
              port: 8080
            initialDelaySeconds: 15
            periodSeconds: 20
            failureThreshold: 3

          # Is it ready to receive traffic? If this fails, traffic is WITHDRAWN
          # but it is NOT restarted. It is the probe that avoids serving errors
          # during start-up or a momentary overload.
          readinessProbe:
            httpGet:
              path: /salud
              port: 8080
            initialDelaySeconds: 5
            periodSeconds: 5
            failureThreshold: 2

          # Hardening required by Autopilot
          securityContext:
            runAsNonRoot: true
            runAsUser: 1000
            allowPrivilegeEscalation: false
            capabilities:
              drop: ["ALL"]

      # Spreads the pods across zones: if europe-west1-b goes down, the shop carries on.
      topologySpreadConstraints:
        - maxSkew: 1
          topologyKey: topology.kubernetes.io/zone
          whenUnsatisfiable: ScheduleAnyway
          labelSelector:
            matchLabels:
              app: alpinashop-web

The points people get wrong most often:

  • selector.matchLabels must match template.metadata.labels. Otherwise the Deployment does not recognise its own pods.
  • requests versus limits. requests is what the scheduler reserves and, on Autopilot, what you are billed for. limits is the ceiling: exceeding the memory limit kills the container with OOMKilled. Making the request and limit memory equal avoids surprises.
  • Liveness and readiness are not the same thing. The first restarts; the second only withdraws traffic. A badly configured liveness probe — for example, pointing at a route that queries the database — causes cascading restarts when the database is slow, making the problem worse. Keep /salud lightweight and free of external dependencies.
  • maxUnavailable: 0 guarantees that capacity never drops during a deployment.

Applying and checking:

kubectl apply -f deployment.yaml

kubectl get deployments
kubectl get pods -o wide          # shows which node and zone each pod lands in
kubectl describe pod <pod-name>
kubectl logs -f deployment/alpinashop-web

  1. The Service: exposing the application

Pods have ephemeral, changing IPs. A Service offers a stable DNS name and a virtual IP that distribute traffic across the pods matching its selector.

Service type Reach Use
ClusterIP (default) Inside the cluster only Communication between internal services
NodePort A port on every node Rarely used directly
LoadBalancer External public IP Exposing a service to the internet
ExternalName DNS alias to an external host Integration with outside services
# service.yaml
apiVersion: v1
kind: Service
metadata:
  name: alpinashop-web
  namespace: tienda
  annotations:
    # Native load balancing by pod IP: traffic goes straight to the pod,
    # with no extra hop through the node. Less latency and better health checking.
    cloud.google.com/neg: '{"ingress": true}'
spec:
  type: LoadBalancer
  selector:
    app: alpinashop-web      # MUST match the pods' labels
  ports:
    - name: http
      protocol: TCP
      port: 80               # port exposed to the outside
      targetPort: 8080       # container port
kubectl apply -f service.yaml

# The external IP takes 1-2 minutes to appear
kubectl get service alpinashop-web --watch

IP=$(kubectl get service alpinashop-web -o jsonpath='{.status.loadBalancer.ingress[0].ip}')
curl -s "http://$IP/"

A LoadBalancer-type Service in GKE automatically creates a Google Cloud network load balancer. It is the fastest way to expose something, but for a real shop you will want HTTPS, your own domain, path-based routing and protection against attacks. That is achieved with an Ingress or a Gateway in front of a global HTTP(S) load balancer, together with Cloud Armor and managed certificates: the content of lessons 03-02, 03-05 and 03-07. Here we deliberately stop at the L4 LoadBalancer.

Inside the cluster, any pod can call this service by its DNS name:

http://alpinashop-web.tienda.svc.cluster.local
http://alpinashop-web            # short form, within the same namespace

  1. Scaling: manual, HPA and cluster autoscaling

Kubernetes scales at two independent levels: the number of pods and the number of nodes.

Manual pod scaling:

kubectl scale deployment alpinashop-web --replicas=5
kubectl get pods -w

HorizontalPodAutoscaler (HPA): adjusts the replicas automatically according to a metric.

# hpa.yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: alpinashop-web
  namespace: tienda
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: alpinashop-web

  minReplicas: 3
  maxReplicas: 20

  metrics:
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: 60    # % over the CPU REQUEST, not over the limit
    - type: Resource
      resource:
        name: memory
        target:
          type: Utilization
          averageUtilization: 75

  behavior:
    scaleUp:
      stabilizationWindowSeconds: 30   # reacts quickly to peaks
      policies:
        - type: Percent
          value: 100                   # at most, double every 30 s
          periodSeconds: 30
    scaleDown:
      stabilizationWindowSeconds: 300  # comes down slowly: avoids oscillation
      policies:
        - type: Pods
          value: 1
          periodSeconds: 60            # remove at most 1 pod per minute
kubectl apply -f hpa.yaml
kubectl get hpa alpinashop-web --watch
kubectl describe hpa alpinashop-web    # shows why it scaled or did not

The behavior block is what separates an HPA that works from one that oscillates. The asymmetry is deliberate: up fast and down slow. A traffic peak requires immediate capacity; withdrawing capacity in a hurry causes flapping, with pods being created and destroyed non-stop. The 300-second window on the way down is a very reasonable recommendation as a starting point.

Watch out for one detail: the HPA's percentage is calculated over the request, not over the limit. With requests.cpu: 250m and a 60 % target, the HPA scales when average usage goes above 150 mCPU per pod. A badly set request throws the whole autoscaling out.

Cluster autoscaling. On Autopilot it is automatic and invisible: if the new pods do not fit, Google adds capacity; if there is spare, it removes it. There is nothing to configure, which is precisely its value proposition. On Standard it has to be configured per pool:

gcloud container clusters update alpinashop-cluster-std \
  --enable-autoscaling --min-nodes=1 --max-nodes=10 \
  --node-pool=pool-principal --region=europe-west1

A quick load test to see the HPA in action:

kubectl run generador-carga --rm -it --image=busybox:1.36 --restart=Never -- \
  /bin/sh -c "while true; do wget -q -O- http://alpinashop-web.tienda/; done"

In another terminal, kubectl get hpa --watch will show CPU usage rising and, after a few seconds, the number of replicas increasing.

  1. ConfigMaps and Secrets

ConfigMap for non-sensitive configuration:

# configmap.yaml
apiVersion: v1
kind: ConfigMap
metadata:
  name: config-catalogo
  namespace: tienda
data:
  entorno: "produccion"
  instancia_sql: "alpinashop-prod:europe-west1:alpinashop-pedidos"
  db_user: "app_catalogo"
  db_name: "tienda"
  productos_por_pagina: "24"
kubectl apply -f configmap.yaml
kubectl get configmap config-catalogo -o yaml

Secret for sensitive data:

kubectl create secret generic secreto-catalogo \
  --from-literal=db_pass='<password>' \
  --namespace=tienda

And here comes the most important warning in this section: Kubernetes Secrets are base64-encoded, not encrypted. Anyone with permission to read secrets in that namespace sees them in the clear:

kubectl get secret secreto-catalogo -o jsonpath='{.data.db_pass}' | base64 -d

Base64 is not encryption, it is an encoding. A Kubernetes Secret protects against the password appearing in a kubectl describe or in a manifest in the repository, and nothing more.

The correct approach in production is Secret Manager (lesson 03-06), integrated with GKE through the Secret Manager CSI driver add-on, which mounts the secrets as files fetched at run time, with centralised rotation and auditing. The rule of thumb: Kubernetes Secrets are fine for development and for low-impact data; real production credentials live in Secret Manager.

Workload Identity deserves a paragraph of its own because it solves the problem at the root. It binds a Kubernetes service account to a Google Cloud one, so that pods obtain Google Cloud credentials automatically, with no key files:

PROJECT=alpinashop-prod

# 1. Google Cloud service account
gcloud iam service-accounts create sa-catalogo-gke \
  --display-name="Catalogue on GKE"

gcloud projects add-iam-policy-binding $PROJECT \
  --member="serviceAccount:sa-catalogo-gke@$PROJECT.iam.gserviceaccount.com" \
  --role="roles/cloudsql.client"

gcloud storage buckets add-iam-policy-binding gs://alpinashop-catalogo \
  --member="serviceAccount:sa-catalogo-gke@$PROJECT.iam.gserviceaccount.com" \
  --role="roles/storage.objectUser"

# 2. Kubernetes service account
kubectl create serviceaccount sa-catalogo --namespace=tienda

# 3. Bind the two together
gcloud iam service-accounts add-iam-policy-binding \
  "sa-catalogo-gke@$PROJECT.iam.gserviceaccount.com" \
  --role="roles/iam.workloadIdentityUser" \
  --member="serviceAccount:$PROJECT.svc.id.goog[tienda/sa-catalogo]"

kubectl annotate serviceaccount sa-catalogo --namespace=tienda \
  iam.gke.io/gcp-service-account="sa-catalogo-gke@$PROJECT.iam.gserviceaccount.com"

With this, the same storage.Client() and the same Cloud SQL connector from the previous lessons work inside the pod with no credentials at all. It is the natural continuation of what you already saw in Compute Engine and App Engine: identity comes from the platform, not from a file.

  1. Interruption-free updates and rollback

With strategy: RollingUpdate and maxUnavailable: 0, updating the image replaces the pods one at a time without reducing capacity:

# Publish the new version
gcloud builds submit \
  --tag europe-west1-docker.pkg.dev/alpinashop-prod/alpinashop/catalogo:v2 .

# Update the Deployment
kubectl set image deployment/alpinashop-web \
  catalogo=europe-west1-docker.pkg.dev/alpinashop-prod/alpinashop/catalogo:v2

# Follow the progress
kubectl rollout status deployment/alpinashop-web

# Revision history
kubectl rollout history deployment/alpinashop-web

The process, step by step: a pod with v2 is created, the system waits for its readinessProbe to give the go-ahead, a v1 pod is withdrawn, and it repeats. Because maxUnavailable: 0, there are always at least three ready pods. Here you see why the readinessProbe is indispensable: without it, Kubernetes would accept the new version as soon as the container started, before gunicorn was accepting requests, and some clients would get errors.

If the new version fails:

# Go back to the previous revision
kubectl rollout undo deployment/alpinashop-web

# Go back to a specific revision
kubectl rollout undo deployment/alpinashop-web --to-revision=3

# Pause an update in progress (manual canary)
kubectl rollout pause deployment/alpinashop-web
kubectl rollout resume deployment/alpinashop-web

To protect the service during cluster maintenance operations too — node updates, rescaling — define a PodDisruptionBudget:

apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: alpinashop-web
  namespace: tienda
spec:
  minAvailable: 2
  selector:
    matchLabels:
      app: alpinashop-web

This tells Kubernetes never to leave fewer than 2 pods available when it decides to move pods. Without a PDB, a node update could drain an entire node and leave the shop with less capacity than it needs at the worst moment.

  1. Basic cluster observability

GKE sends metrics and logs to Cloud Monitoring and Cloud Logging with no additional configuration. The everyday diagnosis commands:

# General state
kubectl get all -n tienda
kubectl get events -n tienda --sort-by=.metadata.creationTimestamp

# Actual consumption of pods and nodes
kubectl top pods -n tienda
kubectl top nodes

# Logs
kubectl logs -f deployment/alpinashop-web
kubectl logs deployment/alpinashop-web --previous   # logs of the container that died

# Diagnosing a pod that does not start
kubectl describe pod <pod>          # the Events section almost always says why

# Open a shell inside the container
kubectl exec -it <pod> -- /bin/sh

# Test the service without exposing it
kubectl port-forward service/alpinashop-web 8080:80

The pod states you will see most often and what they mean:

State Usual cause
Pending There are no resources to schedule it; on Autopilot, capacity is being added
ImagePullBackOff The image does not exist or permission on Artifact Registry is missing
CrashLoopBackOff The container starts and dies in a loop. Look at kubectl logs --previous
OOMKilled It exceeded limits.memory. Raise the limit or fix the leak
Running but not ready The readinessProbe is failing; check the path and the port

In Cloud Monitoring there are predefined GKE dashboards with usage by cluster, namespace and workload, and alerts can be defined on container restarts, unavailable pods or latency. The full treatment is in 06-04 and 06-06.

Common Mistakes and Tips

  • A Service selector that does not match the pods' labels. The Service exists, gives no error and sends traffic to nobody. Check with kubectl get endpoints alpinashop-web.
  • Not defining requests and limits. On Autopilot default values are applied that are probably not yours; on Standard, a pod without limits can starve its neighbours.
  • Pointing the liveness probe at a route that queries the database. If the database is slow, Kubernetes restarts every pod and makes the incident worse.
  • Using the :latest tag. You do not know which version is running and you cannot roll back reliably.
  • Storing real credentials in Kubernetes Secrets. Base64 is not encryption. Use Secret Manager.
  • Running as root. Autopilot rejects it, and on Standard it is an unnecessary risk.
  • Scaling down too fast in the HPA. It causes oscillation. Use a wide stabilisation window.
  • Always working in the default namespace. Separate by environment and team from the start.
  • Forgetting the PodDisruptionBudget. Cluster maintenance can leave you without capacity.
  • Tip: kubectl describe before any other command when something does not work; the Events section usually gives the answer.
  • Tip: tag images with the commit hash. Complete traceability between code and pod.
  • Tip: use kubectl apply -f on files versioned in Git, never kubectl edit in production. What is not in Git does not exist.
  • Tip: kubectl port-forward lets you test an internal service without exposing it to the internet.

Exercises

Exercise 1: cluster, image and first deployment

  1. Create an Autopilot cluster alpinashop-cluster-ej in europe-west1 and connect kubectl.
  2. Create an Artifact Registry repository in europe-west1 and publish an image of a minimal Flask application that shows the pod name (the HOSTNAME variable) and responds on /salud.
  3. Write the Deployment with 2 replicas, reasonable requests and limits, liveness and readiness probes, and a securityContext without root.
  4. Write the LoadBalancer-type Service and check with curl that the distribution across pods works.
  5. Check with kubectl get endpoints that the Service has pods associated with it.

Exercise 2: scaling and interruption-free update

  1. Scale manually to 4 replicas and check which zones the pods land in.
  2. Create an HPA from 2 to 8 replicas with a 60 % CPU target and asymmetric behaviour.
  3. Generate load and observe the scaling.
  4. Publish a v2 of the image with a visible change and update the Deployment without interruption, checking with curl in a loop that there is not a single error.
  5. Roll back to the previous version and check the revision history.

Exercise 3: configuration, secrets and diagnosis

  1. Create a ConfigMap with the environment and the number of products per page, and a Secret with a fictitious password.
  2. Inject both into the Deployment as environment variables and verify their value inside the pod.
  3. Demonstrate that the Secret is not encrypted.
  4. Deliberately cause a CrashLoopBackOff (for example, with a wrong start-up command) and diagnose it step by step.
  5. Explain what you would use in production instead of a Kubernetes Secret and why.

Solutions

Solution 1

gcloud container clusters create-auto alpinashop-cluster-ej \
  --region=europe-west1 --release-channel=regular

gcloud container clusters get-credentials alpinashop-cluster-ej --region=europe-west1
kubectl create namespace tienda
kubectl config set-context --current --namespace=tienda

gcloud artifacts repositories create alpinashop \
  --repository-format=docker --location=europe-west1
# main.py
import os
from flask import Flask

app = Flask(__name__)


@app.route("/")
def home():
    return f"<h1>AlpinaShop</h1><p>Pod: {os.environ.get('HOSTNAME', 'local')}</p>"


@app.route("/salud")
def health():
    return "ok", 200
FROM python:3.12-slim
ENV PYTHONUNBUFFERED=1
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY . .
RUN useradd --create-home --uid 1000 alpina && chown -R alpina:alpina /app
USER 1000
EXPOSE 8080
CMD ["gunicorn", "--bind", "0.0.0.0:8080", "--workers", "2", "main:app"]
IMG="europe-west1-docker.pkg.dev/$(gcloud config get-value project)/alpinashop/catalogo:v1"
gcloud builds submit --tag "$IMG" .
# deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: alpinashop-web
  namespace: tienda
spec:
  replicas: 2
  selector:
    matchLabels:
      app: alpinashop-web
  template:
    metadata:
      labels:
        app: alpinashop-web
    spec:
      containers:
        - name: catalogo
          image: europe-west1-docker.pkg.dev/PROJECT/alpinashop/catalogo:v1
          ports:
            - containerPort: 8080
          resources:
            requests: { cpu: "250m", memory: "512Mi" }
            limits:   { cpu: "500m", memory: "512Mi" }
          livenessProbe:
            httpGet: { path: /salud, port: 8080 }
            initialDelaySeconds: 15
            periodSeconds: 20
          readinessProbe:
            httpGet: { path: /salud, port: 8080 }
            initialDelaySeconds: 5
            periodSeconds: 5
          securityContext:
            runAsNonRoot: true
            runAsUser: 1000
            allowPrivilegeEscalation: false
            capabilities: { drop: ["ALL"] }
kubectl apply -f deployment.yaml -f service.yaml

IP=$(kubectl get svc alpinashop-web -o jsonpath='{.status.loadBalancer.ingress[0].ip}')
for i in $(seq 1 10); do curl -s "http://$IP/" | grep -o "Pod: [a-z0-9-]*"; done

# 5. Check that the Service has endpoints
kubectl get endpoints alpinashop-web

If kubectl get endpoints returns <none>, the Service's selector does not match the pods' labels, or no pod is passing the readinessProbe. It is the first check to make when a Service does not respond.

Solution 2

# 1. Manual scaling and distribution across zones
kubectl scale deployment alpinashop-web --replicas=4
kubectl get pods -o custom-columns=\
NAME:.metadata.name,NODE:.spec.nodeName,STATUS:.status.phase
# 2. hpa.yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: alpinashop-web
  namespace: tienda
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: alpinashop-web
  minReplicas: 2
  maxReplicas: 8
  metrics:
    - type: Resource
      resource:
        name: cpu
        target: { type: Utilization, averageUtilization: 60 }
  behavior:
    scaleUp:
      stabilizationWindowSeconds: 30
    scaleDown:
      stabilizationWindowSeconds: 300
kubectl apply -f hpa.yaml

# 3. Generate load and observe
kubectl run carga --rm -it --image=busybox:1.36 --restart=Never -- \
  /bin/sh -c "while true; do wget -q -O- http://alpinashop-web.tienda/ >/dev/null; done"
# in another terminal:
kubectl get hpa alpinashop-web --watch

# 4. Interruption-free update, checking that there are no errors
gcloud builds submit --tag "${IMG%:v1}:v2" .

while true; do
  curl -s -o /dev/null -w "%{http_code} " "http://$IP/"
  sleep 0.3
done &

kubectl set image deployment/alpinashop-web catalogo="${IMG%:v1}:v2"
kubectl rollout status deployment/alpinashop-web

# 5. Rollback
kubectl rollout undo deployment/alpinashop-web
kubectl rollout history deployment/alpinashop-web

During the update, the curl loop must show nothing but 200 codes. If 502s or 503s appeared, the cause is almost always a badly configured readinessProbe or a maxUnavailable greater than zero.

Solution 3

# 1. ConfigMap and Secret
kubectl create configmap config-catalogo \
  --from-literal=entorno=produccion \
  --from-literal=productos_por_pagina=24

kubectl create secret generic secreto-catalogo \
  --from-literal=db_pass='TestPassword123'
# 2. Injection into the Deployment
          env:
            - name: ENTORNO
              valueFrom:
                configMapKeyRef: { name: config-catalogo, key: entorno }
            - name: PRODUCTOS_POR_PAGINA
              valueFrom:
                configMapKeyRef: { name: config-catalogo, key: productos_por_pagina }
            - name: DB_PASS
              valueFrom:
                secretKeyRef: { name: secreto-catalogo, key: db_pass }
kubectl apply -f deployment.yaml
kubectl exec -it deployment/alpinashop-web -- env | grep -E "ENTORNO|PRODUCTOS|DB_PASS"

# 3. The Secret is not encrypted
kubectl get secret secreto-catalogo -o jsonpath='{.data.db_pass}' | base64 -d; echo

# 4. Cause and diagnose a CrashLoopBackOff
kubectl set image deployment/alpinashop-web catalogo=python:3.12-slim
kubectl get pods
kubectl describe pod <pod>              # Events: Back-off restarting failed container
kubectl logs <pod> --previous           # output of the container that died
kubectl rollout undo deployment/alpinashop-web

Correct diagnosis always follows this order: kubectl get pods to see the state, kubectl describe pod to read the Events (which tell you whether the problem is the image, the resources or the probes) and kubectl logs --previous to see what the container printed just before dying. In this case, the python:3.12-slim image has no useful start-up command and terminates immediately, with Kubernetes retrying with exponential backoff.

  1. In production I would use Secret Manager with GKE's CSI add-on, or Workload Identity to remove the credential outright. The reasons: Kubernetes Secrets are only base64-encoded and readable by anyone with read permissions in the namespace; they have no automatic rotation; they leave no audit trail of who accessed them; and if they are versioned in Git, the credential stays in the history forever. Secret Manager encrypts at rest, controls access with granular IAM, logs every access and allows rotation without redeploying. It is studied in 03-06.

Cleaning up:

kubectl delete namespace tienda
gcloud container clusters delete alpinashop-cluster-ej --region=europe-west1 --quiet

A forgotten cluster is one of the most expensive resources you can leave running. Delete it when you finish the exercises.

Conclusion

You have travelled through Kubernetes from the concepts to a real application in production. You know that it is a declarative system where you describe the desired state and the system converges towards it, and you know the objects that matter: the pod as an ephemeral unit, the ReplicaSet that counts and replenishes, the Deployment that orchestrates updates, the Service that gives a stable identity to a changing set of pods, and the namespaces that partition the cluster. You have understood the separation between control plane and nodes, and why Google managing the former is the core of GKE's value, along with automatic repair, node autoscaling, updates by channel and integration with the network and with IAM.

You have compared Autopilot and Standard, and you have seen that the essential difference is what gets billed: the nodes in Standard, what your pods request in Autopilot. AlpinaShop chose Autopilot because Marta is a single person and the application has no special hardware requirements. You have created the regional alpinashop-cluster, connected kubectl, written a Dockerfile with dependency caching and an unprivileged user, and published the image in Artifact Registry with versioned tags and never :latest. You have written the complete alpinashop-web Deployment with requests and limits, clearly differentiated liveness and readiness probes, a security context and distribution across zones; and the LoadBalancer-type Service, knowing that HTTPS, your own domain and advanced routing will arrive in module 3.

You have scaled by hand, configured a HorizontalPodAutoscaler with asymmetric behaviour — fast up, slow down — and checked that the percentage is calculated over the request. You have managed configuration with ConfigMaps and discovered that Kubernetes Secrets are only base64-encoded, with Secret Manager and Workload Identity as the correct answer. And you have carried out an update without a single 500 error and its rollback in one command, additionally protected with a PodDisruptionBudget.

AlpinaShop now has its compute solved at four different levels and its transactional data in Cloud SQL. But not everything fits in a relational database: the shopping cart that changes with every click, the user sessions, the telemetry of which products people look at. In 02-06, NoSQL Databases: Firestore, Bigtable and Spanner, we will see why other data models exist — document, wide column, distributed relational — we will understand consistency and the CAP theorem without dogmatism, store AlpinaShop's cart in Firestore and its click telemetry in Bigtable, place Spanner, Memorystore and BigQuery on a comparative map, and close with a decision tree and the reasoned allocation of which AlpinaShop data lives where.

Google Cloud Platform (GCP) Course

Module 1: Introduction to Google Cloud Platform

Module 2: Core GCP Services

Module 3: Networking and Security

Module 4: Data and Analytics

Module 5: Machine Learning and AI

Module 6: DevOps and Monitoring

Module 7: Advanced GCP Topics

Module 8: Final Project

© Copyright 2026. All rights reserved