We closed the previous lesson with a conclusion: the investment in App Engine does not compound outside App Engine, but the investment in containers is useful everywhere. Kubernetes is the de facto standard for orchestrating containers, and Google Kubernetes Engine is Google's managed implementation — significantly, the house where Kubernetes was born, derived from the internal Borg system.
Kubernetes has a reputation for complexity, and it is partly deserved. But that complexity answers a real problem: when you have dozens of containers spread over several machines, someone has to decide where each one runs, restart the ones that fail, distribute traffic, update without interrupting the service and grow when a peak arrives. If Kubernetes does not do it, you do it by hand.
In this lesson you will learn Kubernetes from scratch with just enough to work, create the alpinashop-cluster cluster, publish the Flask catalogue image in Artifact Registry and deploy alpinashop-web with complete manifests, autoscaling, configuration, secrets and interruption-free updates together with their rollback.
Contents
- Kubernetes in ten minutes: the objects that matter
- Nodes and control plane
- What GKE adds
- Autopilot versus Standard
- Creating
alpinashop-clusterand connectingkubectl - Building and publishing the image in Artifact Registry
- The
alpinashop-webDeployment - The Service: exposing the application
- Scaling: manual, HPA and cluster autoscaling
- ConfigMaps and Secrets
- Interruption-free updates and rollback
- Basic cluster observability
- Kubernetes in ten minutes: the objects that matter
Kubernetes is a declarative system: you describe the desired state in YAML files and the system works continuously to make reality match that description. If you ask for three replicas and one dies, Kubernetes creates another. You do not give orders; you declare goals.
The essential objects:
| Object | What it is | Analogy |
|---|---|---|
| Container | A process packaged with all its dependencies | The application and its environment, in a box |
| Pod | The smallest unit Kubernetes deploys: one or more containers sharing network and storage | A disposable "logical server" |
| ReplicaSet | Guarantees that N identical pods exist | The one that counts and replenishes |
| Deployment | Manages ReplicaSets and orchestrates updates | What you actually write |
| Service | A stable name and IP that distribute traffic across pods | The internal load balancer |
| Ingress | HTTP(S) routing from outside towards several Services | The reverse proxy |
| Namespace | A logical partition of the cluster | A folder with permissions and quotas |
| ConfigMap | Non-sensitive configuration | A configuration file |
| Secret | Sensitive data (with caveats, section 10) | A sealed envelope, not an armoured one |
| Node | A machine (VM) that runs pods | The physical server |
Three ideas to internalise before writing any YAML:
- Pods are ephemeral and disposable. They are born, they die and they are recreated with another name and another IP. Never connect to a pod by its IP: that is what the Service is for. It is the same lesson as the MIG instances in 02-01, taken to the extreme.
- You almost never create pods directly. You create a Deployment, which creates a ReplicaSet, which creates the pods. Each layer adds a guarantee.
- Everything is identified through labels. A Service does not know its pods by name: it selects the ones carrying a particular label. If the selector does not match the pods' labels, the Service exists but sends traffic to nobody. It is the number one beginner's mistake.
graph TD
subgraph "Control plane (managed by Google)"
API[API Server]
SCHED[Scheduler]
CM[Controller Manager]
ETCD[(etcd)]
end
subgraph "Nodes (Compute Engine VMs)"
subgraph "Node 1"
P1[Pod alpinashop-web]
P2[Pod alpinashop-web]
end
subgraph "Node 2"
P3[Pod alpinashop-web]
end
end
DEP[Deployment<br/>alpinashop-web] --> RS[ReplicaSet]
RS --> P1
RS --> P2
RS --> P3
SVC[Service LoadBalancer<br/>alpinashop-web] --> P1
SVC --> P2
SVC --> P3
KUBECTL[kubectl] --> API
API --> SCHED
API --> CM
API --> ETCD
INTERNET((Internet)) --> SVC
- Nodes and control plane
A Kubernetes cluster has two halves:
The control plane is the brain. It contains the API Server (the only way in: kubectl and everything else talks to it), the Scheduler (which decides what node each pod goes to), the Controller Managers (the loops that compare actual state with desired state and act) and etcd (the database that holds the cluster's entire state).
The nodes are the machines that run the pods. In GKE they are Compute Engine VMs — the same ones from lesson 02-01 — each with the kubelet (the agent that talks to the control plane and starts containers), a container runtime and kube-proxy (which implements the Services' networking).
Setting up and maintaining a control plane yourself is considerable work: high availability of etcd, certificates, coordinated updates, backups. That is exactly the part GKE takes away from you.
- What GKE adds
On top of vanilla Kubernetes, GKE contributes:
- A managed control plane with an SLA. Google deploys it, replicates it, patches it and updates it. In regional mode, replicated across several zones.
- Automatic updates of the control plane and the nodes, with release channels (
rapid,regular,stable) so you can choose how much novelty you want. - Automatic node repair: a node that stops responding is recreated.
- Node autoscaling: if no more pods fit, nodes are added; if there are spare ones, they are removed.
- Integration with Google's network: LoadBalancer-type Services create native Google Cloud load balancers, and pods get IPs from the VPC (03-01).
- Integration with IAM and Workload Identity: pods authenticate to Google Cloud APIs with an identity of their own, without keys.
- Cloud Logging and Cloud Monitoring built in from the very first moment.
- Autopilot: a mode where you do not even see the nodes.
- Autopilot versus Standard
It is the first decision when creating a cluster, and it shapes day-to-day work.
| Autopilot | Standard | |
|---|---|---|
| Who manages the nodes? | Google, entirely | You: size, number, image, pools |
| Do you see the VMs? | No | Yes, in Compute Engine |
| Billing | By the CPU, memory and disk requested by your pods | By the node VMs, used or not |
| Node scaling | Automatic and invisible | Autoscaler configurable per pool |
| Security | Hardened by default (no privileges, no host access) | Configurable, more permissive |
| DaemonSets and host access | Limited | Allowed |
| GPUs and special hardware | Supported with restrictions | Full control |
| Operational overhead | Minimal | Medium-high |
| When to choose it | The general case, small teams, standard applications | Specific hardware needs, privileged agents, fine-grained cost tuning at scale |
The billing difference is the key to understanding it. In Standard you pay for the nodes: if you have three e2-standard-4 VMs and your pods use 15 %, you pay 100 %. In Autopilot you pay what your pods request: if a pod asks for 250 mCPU and 512 MiB, that is what is billed. This has two direct consequences: the requests in your manifests stop being a recommendation and become your bill, and there is no incentive to "fill up" nodes.
AlpinaShop chooses Autopilot. Marta is the only person responsible for infrastructure in a 40-person company; she has no time to size node pools or tune the binpacking. The application is a standard web container with no special requirements. Autopilot removes a whole category of work in exchange for a per-unit surcharge that, at this scale, is irrelevant next to Marta's hours.
- Creating
alpinashop-cluster and connecting kubectl
alpinashop-cluster and connecting kubectlgcloud services enable container.googleapis.com artifactregistry.googleapis.com
gcloud container clusters create-auto alpinashop-cluster \
--project=alpinashop-prod \
--region=europe-west1 \
--release-channel=regular \
--labels=entorno=prod,equipo=plataforma,centro-coste=tienda,aplicacion=catalogoNotes on the command:
create-autocreates an Autopilot cluster. For Standard it would becreatewith all the node pool configuration.--region(not--zone): Autopilot clusters are always regional, with the control plane replicated across zones. It is the right choice for production and consistent with the reasoning in 01-05.--release-channel=regular: stable versions with automatic updates.stableis more conservative andrapidgives early access to new versions.
Creation takes between 5 and 10 minutes. Afterwards you have to tell kubectl which cluster to talk to:
# Install the authentication plugin if it is missing (Cloud Shell already has it)
gcloud components install gke-gcloud-auth-plugin
# Get credentials: writes the configuration into ~/.kube/config
gcloud container clusters get-credentials alpinashop-cluster \
--region=europe-west1 --project=alpinashop-prod
# Check the connection
kubectl cluster-info
kubectl get nodes
kubectl get namespacesOn Autopilot, kubectl get nodes may return an empty list at first: the nodes appear when there are pods to run. It is disconcerting the first time and it is exactly the expected behaviour.
We create a namespace of our own, instead of working in default:
Namespaces let you separate environments and teams within a cluster, with their own resource quotas and RBAC permissions. Using default for everything is a bad habit that you pay for when the cluster grows.
- Building and publishing the image in Artifact Registry
Kubernetes runs container images, so the Flask catalogue has to be packaged. Artifact Registry is Google Cloud's artefact registry — the successor to Container Registry, which has been retired; if you find documentation using gcr.io, it is out of date.
gcloud artifacts repositories create alpinashop \
--repository-format=docker \
--location=europe-west1 \
--description="AlpinaShop container images"
# Configure Docker to authenticate against this registry
gcloud auth configure-docker europe-west1-docker.pkg.devThe catalogue's Dockerfile:
# Slim base image: smaller than the full one, with less attack surface
FROM python:3.12-slim
# Stops Python writing .pyc files and forces unbuffered output,
# essential for the logs to reach Cloud Logging in real time.
ENV PYTHONDONTWRITEBYTECODE=1 \
PYTHONUNBUFFERED=1
WORKDIR /app
# Copy ONLY requirements.txt first: if it does not change, Docker reuses
# the dependency layer in subsequent builds. It is the most
# profitable cache optimisation in a Dockerfile.
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
# Now the code, which changes with every commit
COPY . .
# Unprivileged user: Autopilot REJECTS containers that
# try to run as root.
RUN useradd --create-home --uid 1000 alpina && chown -R alpina:alpina /app
USER 1000
EXPOSE 8080
# 2 workers and 4 threads: the app spends a lot of time waiting for the database,
# so the threads make good use of that dead time.
CMD ["gunicorn", "--bind", "0.0.0.0:8080", "--workers", "2", \
"--threads", "4", "--timeout", "60", "main:app"]Building and publishing. There are two routes:
# Option A: build locally with Docker
IMAGE="europe-west1-docker.pkg.dev/alpinashop-prod/alpinashop/catalogo:v1"
docker build -t "$IMAGE" .
docker push "$IMAGE"
# Option B: build in the cloud with Cloud Build (no local Docker required)
gcloud builds submit --tag "$IMAGE" .Option B is especially convenient from Cloud Shell and is the seed of the continuous integration pipeline we will set up in 06-01.
Check what has been published:
gcloud artifacts docker images list \
europe-west1-docker.pkg.dev/alpinashop-prod/alpinashop
# Vulnerability scanning (if enabled on the repository)
gcloud artifacts docker images scan "$IMAGE"Two good tagging practices: never use :latest in production — you will not know which version is running and you will not be able to roll back — and tag with the commit hash (catalogo:a3f9c1d), which makes it traceable which code is in each pod. It is the same idea of immutable names that we applied to Storage objects in 02-02 and to instance templates in 02-01.
- The
alpinashop-web Deployment
alpinashop-web DeploymentNow the main manifest. Every block is commented:
# deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: alpinashop-web
namespace: tienda
labels:
app: alpinashop-web
entorno: prod
spec:
# Desired number of pods. The HPA (section 9) will override it later.
replicas: 3
# The Deployment governs the pods matching this selector.
# It MUST match the labels of the template below.
selector:
matchLabels:
app: alpinashop-web
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 1 # up to 1 extra pod during the update
maxUnavailable: 0 # never drop below the desired number: zero interruption
template:
metadata:
labels:
app: alpinashop-web # label the Service will use
entorno: prod
spec:
# Kubernetes service account bound through Workload Identity
# to a Google Cloud service account (section 10).
serviceAccountName: sa-catalogo
containers:
- name: catalogo
image: europe-west1-docker.pkg.dev/alpinashop-prod/alpinashop/catalogo:v1
ports:
- name: http
containerPort: 8080
# requests: what the scheduler reserves. On Autopilot, WHAT YOU PAY.
# limits: the ceiling. If memory is exceeded, the container dies (OOMKilled).
resources:
requests:
cpu: "250m" # 0.25 of a vCPU
memory: "512Mi"
ephemeral-storage: "1Gi"
limits:
cpu: "500m"
memory: "512Mi" # same as requests: avoids memory surprises
env:
- name: BUCKET_CATALOGO
value: "alpinashop-catalogo"
- name: ENTORNO
valueFrom:
configMapKeyRef:
name: config-catalogo
key: entorno
- name: INSTANCIA_SQL
valueFrom:
configMapKeyRef:
name: config-catalogo
key: instancia_sql
- name: DB_USER
valueFrom:
configMapKeyRef:
name: config-catalogo
key: db_user
- name: DB_PASS
valueFrom:
secretKeyRef:
name: secreto-catalogo
key: db_pass
# Is the container alive? If this fails, Kubernetes RESTARTS it.
livenessProbe:
httpGet:
path: /salud
port: 8080
initialDelaySeconds: 15
periodSeconds: 20
failureThreshold: 3
# Is it ready to receive traffic? If this fails, traffic is WITHDRAWN
# but it is NOT restarted. It is the probe that avoids serving errors
# during start-up or a momentary overload.
readinessProbe:
httpGet:
path: /salud
port: 8080
initialDelaySeconds: 5
periodSeconds: 5
failureThreshold: 2
# Hardening required by Autopilot
securityContext:
runAsNonRoot: true
runAsUser: 1000
allowPrivilegeEscalation: false
capabilities:
drop: ["ALL"]
# Spreads the pods across zones: if europe-west1-b goes down, the shop carries on.
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: ScheduleAnyway
labelSelector:
matchLabels:
app: alpinashop-webThe points people get wrong most often:
selector.matchLabelsmust matchtemplate.metadata.labels. Otherwise the Deployment does not recognise its own pods.requestsversuslimits.requestsis what the scheduler reserves and, on Autopilot, what you are billed for.limitsis the ceiling: exceeding the memory limit kills the container withOOMKilled. Making the request and limit memory equal avoids surprises.- Liveness and readiness are not the same thing. The first restarts; the second only withdraws traffic. A badly configured liveness probe — for example, pointing at a route that queries the database — causes cascading restarts when the database is slow, making the problem worse. Keep
/saludlightweight and free of external dependencies. maxUnavailable: 0guarantees that capacity never drops during a deployment.
Applying and checking:
kubectl apply -f deployment.yaml
kubectl get deployments
kubectl get pods -o wide # shows which node and zone each pod lands in
kubectl describe pod <pod-name>
kubectl logs -f deployment/alpinashop-web
- The Service: exposing the application
Pods have ephemeral, changing IPs. A Service offers a stable DNS name and a virtual IP that distribute traffic across the pods matching its selector.
| Service type | Reach | Use |
|---|---|---|
ClusterIP (default) |
Inside the cluster only | Communication between internal services |
NodePort |
A port on every node | Rarely used directly |
LoadBalancer |
External public IP | Exposing a service to the internet |
ExternalName |
DNS alias to an external host | Integration with outside services |
# service.yaml
apiVersion: v1
kind: Service
metadata:
name: alpinashop-web
namespace: tienda
annotations:
# Native load balancing by pod IP: traffic goes straight to the pod,
# with no extra hop through the node. Less latency and better health checking.
cloud.google.com/neg: '{"ingress": true}'
spec:
type: LoadBalancer
selector:
app: alpinashop-web # MUST match the pods' labels
ports:
- name: http
protocol: TCP
port: 80 # port exposed to the outside
targetPort: 8080 # container portkubectl apply -f service.yaml
# The external IP takes 1-2 minutes to appear
kubectl get service alpinashop-web --watch
IP=$(kubectl get service alpinashop-web -o jsonpath='{.status.loadBalancer.ingress[0].ip}')
curl -s "http://$IP/"A LoadBalancer-type Service in GKE automatically creates a Google Cloud network load balancer. It is the fastest way to expose something, but for a real shop you will want HTTPS, your own domain, path-based routing and protection against attacks. That is achieved with an Ingress or a Gateway in front of a global HTTP(S) load balancer, together with Cloud Armor and managed certificates: the content of lessons 03-02, 03-05 and 03-07. Here we deliberately stop at the L4 LoadBalancer.
Inside the cluster, any pod can call this service by its DNS name:
http://alpinashop-web.tienda.svc.cluster.local http://alpinashop-web # short form, within the same namespace
- Scaling: manual, HPA and cluster autoscaling
Kubernetes scales at two independent levels: the number of pods and the number of nodes.
Manual pod scaling:
HorizontalPodAutoscaler (HPA): adjusts the replicas automatically according to a metric.
# hpa.yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: alpinashop-web
namespace: tienda
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: alpinashop-web
minReplicas: 3
maxReplicas: 20
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 60 # % over the CPU REQUEST, not over the limit
- type: Resource
resource:
name: memory
target:
type: Utilization
averageUtilization: 75
behavior:
scaleUp:
stabilizationWindowSeconds: 30 # reacts quickly to peaks
policies:
- type: Percent
value: 100 # at most, double every 30 s
periodSeconds: 30
scaleDown:
stabilizationWindowSeconds: 300 # comes down slowly: avoids oscillation
policies:
- type: Pods
value: 1
periodSeconds: 60 # remove at most 1 pod per minutekubectl apply -f hpa.yaml
kubectl get hpa alpinashop-web --watch
kubectl describe hpa alpinashop-web # shows why it scaled or did notThe behavior block is what separates an HPA that works from one that oscillates. The asymmetry is deliberate: up fast and down slow. A traffic peak requires immediate capacity; withdrawing capacity in a hurry causes flapping, with pods being created and destroyed non-stop. The 300-second window on the way down is a very reasonable recommendation as a starting point.
Watch out for one detail: the HPA's percentage is calculated over the request, not over the limit. With requests.cpu: 250m and a 60 % target, the HPA scales when average usage goes above 150 mCPU per pod. A badly set request throws the whole autoscaling out.
Cluster autoscaling. On Autopilot it is automatic and invisible: if the new pods do not fit, Google adds capacity; if there is spare, it removes it. There is nothing to configure, which is precisely its value proposition. On Standard it has to be configured per pool:
gcloud container clusters update alpinashop-cluster-std \
--enable-autoscaling --min-nodes=1 --max-nodes=10 \
--node-pool=pool-principal --region=europe-west1A quick load test to see the HPA in action:
kubectl run generador-carga --rm -it --image=busybox:1.36 --restart=Never -- \
/bin/sh -c "while true; do wget -q -O- http://alpinashop-web.tienda/; done"In another terminal, kubectl get hpa --watch will show CPU usage rising and, after a few seconds, the number of replicas increasing.
- ConfigMaps and Secrets
ConfigMap for non-sensitive configuration:
# configmap.yaml
apiVersion: v1
kind: ConfigMap
metadata:
name: config-catalogo
namespace: tienda
data:
entorno: "produccion"
instancia_sql: "alpinashop-prod:europe-west1:alpinashop-pedidos"
db_user: "app_catalogo"
db_name: "tienda"
productos_por_pagina: "24"Secret for sensitive data:
kubectl create secret generic secreto-catalogo \
--from-literal=db_pass='<password>' \
--namespace=tiendaAnd here comes the most important warning in this section: Kubernetes Secrets are base64-encoded, not encrypted. Anyone with permission to read secrets in that namespace sees them in the clear:
Base64 is not encryption, it is an encoding. A Kubernetes Secret protects against the password appearing in a kubectl describe or in a manifest in the repository, and nothing more.
The correct approach in production is Secret Manager (lesson 03-06), integrated with GKE through the Secret Manager CSI driver add-on, which mounts the secrets as files fetched at run time, with centralised rotation and auditing. The rule of thumb: Kubernetes Secrets are fine for development and for low-impact data; real production credentials live in Secret Manager.
Workload Identity deserves a paragraph of its own because it solves the problem at the root. It binds a Kubernetes service account to a Google Cloud one, so that pods obtain Google Cloud credentials automatically, with no key files:
PROJECT=alpinashop-prod
# 1. Google Cloud service account
gcloud iam service-accounts create sa-catalogo-gke \
--display-name="Catalogue on GKE"
gcloud projects add-iam-policy-binding $PROJECT \
--member="serviceAccount:sa-catalogo-gke@$PROJECT.iam.gserviceaccount.com" \
--role="roles/cloudsql.client"
gcloud storage buckets add-iam-policy-binding gs://alpinashop-catalogo \
--member="serviceAccount:sa-catalogo-gke@$PROJECT.iam.gserviceaccount.com" \
--role="roles/storage.objectUser"
# 2. Kubernetes service account
kubectl create serviceaccount sa-catalogo --namespace=tienda
# 3. Bind the two together
gcloud iam service-accounts add-iam-policy-binding \
"sa-catalogo-gke@$PROJECT.iam.gserviceaccount.com" \
--role="roles/iam.workloadIdentityUser" \
--member="serviceAccount:$PROJECT.svc.id.goog[tienda/sa-catalogo]"
kubectl annotate serviceaccount sa-catalogo --namespace=tienda \
iam.gke.io/gcp-service-account="sa-catalogo-gke@$PROJECT.iam.gserviceaccount.com"With this, the same storage.Client() and the same Cloud SQL connector from the previous lessons work inside the pod with no credentials at all. It is the natural continuation of what you already saw in Compute Engine and App Engine: identity comes from the platform, not from a file.
- Interruption-free updates and rollback
With strategy: RollingUpdate and maxUnavailable: 0, updating the image replaces the pods one at a time without reducing capacity:
# Publish the new version
gcloud builds submit \
--tag europe-west1-docker.pkg.dev/alpinashop-prod/alpinashop/catalogo:v2 .
# Update the Deployment
kubectl set image deployment/alpinashop-web \
catalogo=europe-west1-docker.pkg.dev/alpinashop-prod/alpinashop/catalogo:v2
# Follow the progress
kubectl rollout status deployment/alpinashop-web
# Revision history
kubectl rollout history deployment/alpinashop-webThe process, step by step: a pod with v2 is created, the system waits for its readinessProbe to give the go-ahead, a v1 pod is withdrawn, and it repeats. Because maxUnavailable: 0, there are always at least three ready pods. Here you see why the readinessProbe is indispensable: without it, Kubernetes would accept the new version as soon as the container started, before gunicorn was accepting requests, and some clients would get errors.
If the new version fails:
# Go back to the previous revision
kubectl rollout undo deployment/alpinashop-web
# Go back to a specific revision
kubectl rollout undo deployment/alpinashop-web --to-revision=3
# Pause an update in progress (manual canary)
kubectl rollout pause deployment/alpinashop-web
kubectl rollout resume deployment/alpinashop-webTo protect the service during cluster maintenance operations too — node updates, rescaling — define a PodDisruptionBudget:
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: alpinashop-web
namespace: tienda
spec:
minAvailable: 2
selector:
matchLabels:
app: alpinashop-webThis tells Kubernetes never to leave fewer than 2 pods available when it decides to move pods. Without a PDB, a node update could drain an entire node and leave the shop with less capacity than it needs at the worst moment.
- Basic cluster observability
GKE sends metrics and logs to Cloud Monitoring and Cloud Logging with no additional configuration. The everyday diagnosis commands:
# General state
kubectl get all -n tienda
kubectl get events -n tienda --sort-by=.metadata.creationTimestamp
# Actual consumption of pods and nodes
kubectl top pods -n tienda
kubectl top nodes
# Logs
kubectl logs -f deployment/alpinashop-web
kubectl logs deployment/alpinashop-web --previous # logs of the container that died
# Diagnosing a pod that does not start
kubectl describe pod <pod> # the Events section almost always says why
# Open a shell inside the container
kubectl exec -it <pod> -- /bin/sh
# Test the service without exposing it
kubectl port-forward service/alpinashop-web 8080:80The pod states you will see most often and what they mean:
| State | Usual cause |
|---|---|
Pending |
There are no resources to schedule it; on Autopilot, capacity is being added |
ImagePullBackOff |
The image does not exist or permission on Artifact Registry is missing |
CrashLoopBackOff |
The container starts and dies in a loop. Look at kubectl logs --previous |
OOMKilled |
It exceeded limits.memory. Raise the limit or fix the leak |
Running but not ready |
The readinessProbe is failing; check the path and the port |
In Cloud Monitoring there are predefined GKE dashboards with usage by cluster, namespace and workload, and alerts can be defined on container restarts, unavailable pods or latency. The full treatment is in 06-04 and 06-06.
Common Mistakes and Tips
- A Service selector that does not match the pods' labels. The Service exists, gives no error and sends traffic to nobody. Check with
kubectl get endpoints alpinashop-web. - Not defining
requestsandlimits. On Autopilot default values are applied that are probably not yours; on Standard, a pod without limits can starve its neighbours. - Pointing the liveness probe at a route that queries the database. If the database is slow, Kubernetes restarts every pod and makes the incident worse.
- Using the
:latesttag. You do not know which version is running and you cannot roll back reliably. - Storing real credentials in Kubernetes Secrets. Base64 is not encryption. Use Secret Manager.
- Running as root. Autopilot rejects it, and on Standard it is an unnecessary risk.
- Scaling down too fast in the HPA. It causes oscillation. Use a wide stabilisation window.
- Always working in the
defaultnamespace. Separate by environment and team from the start. - Forgetting the PodDisruptionBudget. Cluster maintenance can leave you without capacity.
- Tip:
kubectl describebefore any other command when something does not work; the Events section usually gives the answer. - Tip: tag images with the commit hash. Complete traceability between code and pod.
- Tip: use
kubectl apply -fon files versioned in Git, neverkubectl editin production. What is not in Git does not exist. - Tip:
kubectl port-forwardlets you test an internal service without exposing it to the internet.
Exercises
Exercise 1: cluster, image and first deployment
- Create an Autopilot cluster
alpinashop-cluster-ejineurope-west1and connectkubectl. - Create an Artifact Registry repository in
europe-west1and publish an image of a minimal Flask application that shows the pod name (theHOSTNAMEvariable) and responds on/salud. - Write the Deployment with 2 replicas, reasonable requests and limits, liveness and readiness probes, and a
securityContextwithout root. - Write the LoadBalancer-type Service and check with
curlthat the distribution across pods works. - Check with
kubectl get endpointsthat the Service has pods associated with it.
Exercise 2: scaling and interruption-free update
- Scale manually to 4 replicas and check which zones the pods land in.
- Create an HPA from 2 to 8 replicas with a 60 % CPU target and asymmetric behaviour.
- Generate load and observe the scaling.
- Publish a
v2of the image with a visible change and update the Deployment without interruption, checking withcurlin a loop that there is not a single error. - Roll back to the previous version and check the revision history.
Exercise 3: configuration, secrets and diagnosis
- Create a ConfigMap with the environment and the number of products per page, and a Secret with a fictitious password.
- Inject both into the Deployment as environment variables and verify their value inside the pod.
- Demonstrate that the Secret is not encrypted.
- Deliberately cause a
CrashLoopBackOff(for example, with a wrong start-up command) and diagnose it step by step. - Explain what you would use in production instead of a Kubernetes Secret and why.
Solutions
Solution 1
gcloud container clusters create-auto alpinashop-cluster-ej \
--region=europe-west1 --release-channel=regular
gcloud container clusters get-credentials alpinashop-cluster-ej --region=europe-west1
kubectl create namespace tienda
kubectl config set-context --current --namespace=tienda
gcloud artifacts repositories create alpinashop \
--repository-format=docker --location=europe-west1# main.py
import os
from flask import Flask
app = Flask(__name__)
@app.route("/")
def home():
return f"<h1>AlpinaShop</h1><p>Pod: {os.environ.get('HOSTNAME', 'local')}</p>"
@app.route("/salud")
def health():
return "ok", 200FROM python:3.12-slim
ENV PYTHONUNBUFFERED=1
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY . .
RUN useradd --create-home --uid 1000 alpina && chown -R alpina:alpina /app
USER 1000
EXPOSE 8080
CMD ["gunicorn", "--bind", "0.0.0.0:8080", "--workers", "2", "main:app"]IMG="europe-west1-docker.pkg.dev/$(gcloud config get-value project)/alpinashop/catalogo:v1"
gcloud builds submit --tag "$IMG" .# deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: alpinashop-web
namespace: tienda
spec:
replicas: 2
selector:
matchLabels:
app: alpinashop-web
template:
metadata:
labels:
app: alpinashop-web
spec:
containers:
- name: catalogo
image: europe-west1-docker.pkg.dev/PROJECT/alpinashop/catalogo:v1
ports:
- containerPort: 8080
resources:
requests: { cpu: "250m", memory: "512Mi" }
limits: { cpu: "500m", memory: "512Mi" }
livenessProbe:
httpGet: { path: /salud, port: 8080 }
initialDelaySeconds: 15
periodSeconds: 20
readinessProbe:
httpGet: { path: /salud, port: 8080 }
initialDelaySeconds: 5
periodSeconds: 5
securityContext:
runAsNonRoot: true
runAsUser: 1000
allowPrivilegeEscalation: false
capabilities: { drop: ["ALL"] }kubectl apply -f deployment.yaml -f service.yaml
IP=$(kubectl get svc alpinashop-web -o jsonpath='{.status.loadBalancer.ingress[0].ip}')
for i in $(seq 1 10); do curl -s "http://$IP/" | grep -o "Pod: [a-z0-9-]*"; done
# 5. Check that the Service has endpoints
kubectl get endpoints alpinashop-webIf kubectl get endpoints returns <none>, the Service's selector does not match the pods' labels, or no pod is passing the readinessProbe. It is the first check to make when a Service does not respond.
Solution 2
# 1. Manual scaling and distribution across zones
kubectl scale deployment alpinashop-web --replicas=4
kubectl get pods -o custom-columns=\
NAME:.metadata.name,NODE:.spec.nodeName,STATUS:.status.phase# 2. hpa.yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: alpinashop-web
namespace: tienda
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: alpinashop-web
minReplicas: 2
maxReplicas: 8
metrics:
- type: Resource
resource:
name: cpu
target: { type: Utilization, averageUtilization: 60 }
behavior:
scaleUp:
stabilizationWindowSeconds: 30
scaleDown:
stabilizationWindowSeconds: 300kubectl apply -f hpa.yaml
# 3. Generate load and observe
kubectl run carga --rm -it --image=busybox:1.36 --restart=Never -- \
/bin/sh -c "while true; do wget -q -O- http://alpinashop-web.tienda/ >/dev/null; done"
# in another terminal:
kubectl get hpa alpinashop-web --watch
# 4. Interruption-free update, checking that there are no errors
gcloud builds submit --tag "${IMG%:v1}:v2" .
while true; do
curl -s -o /dev/null -w "%{http_code} " "http://$IP/"
sleep 0.3
done &
kubectl set image deployment/alpinashop-web catalogo="${IMG%:v1}:v2"
kubectl rollout status deployment/alpinashop-web
# 5. Rollback
kubectl rollout undo deployment/alpinashop-web
kubectl rollout history deployment/alpinashop-webDuring the update, the curl loop must show nothing but 200 codes. If 502s or 503s appeared, the cause is almost always a badly configured readinessProbe or a maxUnavailable greater than zero.
Solution 3
# 1. ConfigMap and Secret
kubectl create configmap config-catalogo \
--from-literal=entorno=produccion \
--from-literal=productos_por_pagina=24
kubectl create secret generic secreto-catalogo \
--from-literal=db_pass='TestPassword123'# 2. Injection into the Deployment
env:
- name: ENTORNO
valueFrom:
configMapKeyRef: { name: config-catalogo, key: entorno }
- name: PRODUCTOS_POR_PAGINA
valueFrom:
configMapKeyRef: { name: config-catalogo, key: productos_por_pagina }
- name: DB_PASS
valueFrom:
secretKeyRef: { name: secreto-catalogo, key: db_pass }kubectl apply -f deployment.yaml
kubectl exec -it deployment/alpinashop-web -- env | grep -E "ENTORNO|PRODUCTOS|DB_PASS"
# 3. The Secret is not encrypted
kubectl get secret secreto-catalogo -o jsonpath='{.data.db_pass}' | base64 -d; echo
# 4. Cause and diagnose a CrashLoopBackOff
kubectl set image deployment/alpinashop-web catalogo=python:3.12-slim
kubectl get pods
kubectl describe pod <pod> # Events: Back-off restarting failed container
kubectl logs <pod> --previous # output of the container that died
kubectl rollout undo deployment/alpinashop-webCorrect diagnosis always follows this order: kubectl get pods to see the state, kubectl describe pod to read the Events (which tell you whether the problem is the image, the resources or the probes) and kubectl logs --previous to see what the container printed just before dying. In this case, the python:3.12-slim image has no useful start-up command and terminates immediately, with Kubernetes retrying with exponential backoff.
- In production I would use Secret Manager with GKE's CSI add-on, or Workload Identity to remove the credential outright. The reasons: Kubernetes Secrets are only base64-encoded and readable by anyone with read permissions in the namespace; they have no automatic rotation; they leave no audit trail of who accessed them; and if they are versioned in Git, the credential stays in the history forever. Secret Manager encrypts at rest, controls access with granular IAM, logs every access and allows rotation without redeploying. It is studied in 03-06.
Cleaning up:
kubectl delete namespace tienda
gcloud container clusters delete alpinashop-cluster-ej --region=europe-west1 --quietA forgotten cluster is one of the most expensive resources you can leave running. Delete it when you finish the exercises.
Conclusion
You have travelled through Kubernetes from the concepts to a real application in production. You know that it is a declarative system where you describe the desired state and the system converges towards it, and you know the objects that matter: the pod as an ephemeral unit, the ReplicaSet that counts and replenishes, the Deployment that orchestrates updates, the Service that gives a stable identity to a changing set of pods, and the namespaces that partition the cluster. You have understood the separation between control plane and nodes, and why Google managing the former is the core of GKE's value, along with automatic repair, node autoscaling, updates by channel and integration with the network and with IAM.
You have compared Autopilot and Standard, and you have seen that the essential difference is what gets billed: the nodes in Standard, what your pods request in Autopilot. AlpinaShop chose Autopilot because Marta is a single person and the application has no special hardware requirements. You have created the regional alpinashop-cluster, connected kubectl, written a Dockerfile with dependency caching and an unprivileged user, and published the image in Artifact Registry with versioned tags and never :latest. You have written the complete alpinashop-web Deployment with requests and limits, clearly differentiated liveness and readiness probes, a security context and distribution across zones; and the LoadBalancer-type Service, knowing that HTTPS, your own domain and advanced routing will arrive in module 3.
You have scaled by hand, configured a HorizontalPodAutoscaler with asymmetric behaviour — fast up, slow down — and checked that the percentage is calculated over the request. You have managed configuration with ConfigMaps and discovered that Kubernetes Secrets are only base64-encoded, with Secret Manager and Workload Identity as the correct answer. And you have carried out an update without a single 500 error and its rollback in one command, additionally protected with a PodDisruptionBudget.
AlpinaShop now has its compute solved at four different levels and its transactional data in Cloud SQL. But not everything fits in a relational database: the shopping cart that changes with every click, the user sessions, the telemetry of which products people look at. In 02-06, NoSQL Databases: Firestore, Bigtable and Spanner, we will see why other data models exist — document, wide column, distributed relational — we will understand consistency and the CAP theorem without dogmatism, store AlpinaShop's cart in Firestore and its click telemetry in Bigtable, place Spanner, Memorystore and BigQuery on a comparative map, and close with a decision tree and the reasoned allocation of which AlpinaShop data lives where.
Google Cloud Platform (GCP) Course
Module 1: Introduction to Google Cloud Platform
- What is Google Cloud Platform?
- Setting Up Your GCP Account
- A Tour of the GCP Console
- Projects, Resource Hierarchy and Billing
- Regions, Zones and the Shared Responsibility Model
- Cloud Shell and the gcloud CLI
Module 2: Core GCP Services
- Compute Engine: Virtual Machines on Google Cloud
- Cloud Storage: Object Storage
- Cloud SQL: Managed Relational Databases
- App Engine: Platform as a Service
- Google Kubernetes Engine (GKE)
- NoSQL Databases: Firestore, Bigtable and Spanner
- How to Choose the Right Compute Service
Module 3: Networking and Security
- VPC Networks
- Cloud Load Balancing
- Cloud CDN
- Identity and Access Management (IAM)
- Cloud Armor
- Secrets and Encryption: Secret Manager and Cloud KMS
- Cloud DNS, TLS Certificates and Publishing Services Securely
Module 4: Data and Analytics
- BigQuery: The Analytical Data Warehouse
- Cloud Dataflow: Batch and Streaming Data Processing
- Cloud Dataproc: Managed Spark and Hadoop
- Cloud Pub/Sub: Asynchronous Messaging
- Cloud Data Fusion: Code-Free Data Integration
- Orchestrating Pipelines with Cloud Composer and Workflows
- Data Governance and Dashboards with Dataplex and Looker Studio
Module 5: Machine Learning and AI
- Vertex AI: The Machine Learning Platform on GCP
- AutoML: Custom Models Without Writing Code
- TensorFlow on GCP: Training and Serving Models
- Natural Language API
- Vision API
- Generative AI on Vertex AI: Gemini Models and Embeddings
- MLOps: From Model to Product with Vertex AI Pipelines
Module 6: DevOps and Monitoring
- Cloud Build: Continuous Integration on GCP
- Cloud Source Repositories and Source Code Management
- Cloud Functions: Serverless Functions
- Cloud Monitoring (formerly Stackdriver): Metrics, Dashboards and Alerts
- Cloud Deployment Manager and Native Infrastructure as Code
- Cloud Logging and Cloud Trace: Logs, Traces and Diagnostics
- Terraform on GCP: Infrastructure as Code in Practice
Module 7: Advanced GCP Topics
- Hybrid and Multicloud with Anthos
- Serverless Computing with Cloud Run
- Advanced Networking: Shared VPC, Peering and Hybrid Connectivity
- Security Best Practices
- Cost Management and Optimization
- Reliability: SLOs, High Availability and Disaster Recovery
- Governance at Scale: Organization, Policies and Auditing
