We have spent three lessons writing storageClassName: rutasnorte-fast without ever having created that class, and it worked because static matching only compared strings of text. The trick is over. The StorageClass is the object that turns that string into something with real consequences: the cluster creates the exact volume at the moment someone asks for it, with the disk type, the reclaim policy and the capabilities the class defines. It is the leap from static to dynamic provisioning, and it changes daily operations completely: nobody creates a PersistentVolume by hand again. In this lesson you will take the StorageClass apart field by field — with special attention to volumeBindingMode, responsible for one of the most expensive and hardest-to-diagnose failures in the cloud —, understand the default class and the difference between storageClassName: "" and omitting the field, design the two Rutas Norte classes and watch a PersistentVolume appear on its own, without writing it.
Contents
- From static to dynamic provisioning
- Anatomy of the StorageClass
provisioner: who creates the volumeparameters: the details of the diskreclaimPolicyandallowVolumeExpansionvolumeBindingMode: the field that averts a disaster- The default class and its traps
storageClassName: ""versus omitting the field- The classes shipped by minikube and the managed clouds
- The Rutas Norte storage classes
- Hands-on: a PVC that creates its own volume
- Migrating a PVC from one class to another
- From static to dynamic provisioning
Let us review the flow you have practised so far and compare it with the one you are about to build:
flowchart LR
subgraph EST["STATIC (05-02, 05-03)"]
direction TB
A1["A human creates the PV<br/>in advance"] --> A4{"Is there a compatible PV<br/>when the PVC arrives?"}
A4 -->|"Yes"| A5["Bound"]
A4 -->|"No"| A6["Pending<br/>until somebody acts"]
end
subgraph DIN["DYNAMIC (this lesson)"]
direction TB
B1["The team creates the PVC<br/>with storageClassName"] --> B2["The provisioner<br/>creates the REAL volume"]
B2 --> B3["The PV is created<br/>automatically"] --> B4["Bound, in seconds"]
end
The four problems with static provisioning we listed in 05-02 disappear at a stroke:
| Static problem | How dynamic solves it |
|---|---|
| Manual work on the critical path | The application team self-serves: it creates the PVC and the volume appears |
| Waste from size fitting | The volume is created at exactly the size requested |
| Orphaned volumes | With reclaimPolicy: Delete, the volume is destroyed when the PVC is deleted |
| Blind topology | volumeBindingMode: WaitForFirstConsumer creates the volume where the pod fits |
The change of mindset is this: the platform stops offering specific volumes and starts offering a service catalogue. "We have fast disk with retention and cheap standard disk; ask for the one you need and the cluster creates it for you."
- Anatomy of the StorageClass
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: rutasnorte-fast # the name is what is written in the PVC
labels: { app.kubernetes.io/part-of: rutas-norte }
annotations: { storageclass.kubernetes.io/is-default-class: "false" }
provisioner: ebs.csi.aws.com # WHO creates the volume
parameters: # HOW it creates it (provisioner-specific)
type: gp3
iops: "6000"
throughput: "250"
encrypted: "true"
fsType: ext4
reclaimPolicy: Retain # what the created PVs inherit
allowVolumeExpansion: true # can it be grown later?
volumeBindingMode: WaitForFirstConsumer
mountOptions: ["noatime"]
allowedTopologies: # (optional) restrict zones
- matchLabelExpressions:
- { key: topology.kubernetes.io/zone, values: ["eu-west-1a", "eu-west-1b"] }Two structural characteristics before getting into the fields:
- It is a cluster resource (no namespace), like the PersistentVolume. It is managed by the platform team.
- It is practically immutable. Apart from a handful of fields (
allowVolumeExpansion, the annotations), a StorageClass cannot be modified once created: it has to be deleted and created again. And take note, deleting a StorageClass does not affect the PVs already created, which carry on working with the configuration they had.
provisioner: who creates the volume
provisioner: who creates the volumeThe provisioner identifies the component that will serve requests for that class. It is the only genuinely mandatory part, along with the name.
| Provisioner | Who implements it | Where it is used |
|---|---|---|
k8s.io/minikube-hostpath |
minikube | Your practice cluster |
rancher.io/local-path |
local-path-provisioner | kind, k3s, development clusters |
ebs.csi.aws.com |
AWS EBS CSI driver | EKS |
disk.csi.azure.com |
Azure Disk CSI driver | AKS |
pd.csi.storage.gke.io |
Google PD CSI driver | GKE |
efs.csi.aws.com |
AWS EFS CSI driver | EKS, when ReadWriteMany is needed |
kubernetes.io/no-provisioner |
Nobody | Classes for local volumes, created by hand |
Two special cases worth recognising. kubernetes.io/no-provisioner is used for classes that provision nothing, which sounds absurd until you see what it is for: grouping local volumes created by hand and taking advantage of volumeBindingMode: WaitForFirstConsumer so that scheduling respects the disk's location. And provisioners with the kubernetes.io/ prefix (such as kubernetes.io/aws-ebs) are the old in-tree drivers, removed from the Kubernetes code: on a 1.30+ cluster they do not work, and if you inherit manifests with those names they have to be migrated to the equivalent CSI driver.
To find out which CSI drivers are actually installed on your cluster: kubectl get csidrivers.
parameters: the details of the disk
parameters: the details of the diskparameters is a map that is opaque to Kubernetes: it is passed as-is to the provisioner, which is what interprets it. That is why the valid keys depend entirely on the driver, and a misspelled key does not throw an error when the class is created, but when the first PVC is created.
Real examples by provider:
# AWS EBS: general-purpose SSD with custom IOPS and throughput
parameters:
type: gp3 # gp2, gp3, io1, io2, st1, sc1
iops: "6000"
throughput: "250" # MiB/s
encrypted: "true"
kmsKeyId: "arn:aws:kms:eu-west-1:111122223333:key/abcd-1234"
fsType: ext4On Google Cloud the equivalent keys are type (from pd-standard to pd-extreme) and replication-type: regional-pd, which replicates the disk across two zones; on Azure they are skuName (Standard_LRS, StandardSSD_LRS, Premium_LRS, UltraSSD_LRS) and cachingmode.
One parameter that must not be missing from any Rutas Norte class is encryption at rest. bookings-postgres stores customers' names, ID numbers, phone numbers and email addresses; that disk must be encrypted, just as the Secrets in etcd already are since 03-02. On AWS it is encrypted: "true"; on Azure and Google encryption at rest is on by default and what you configure is the key.
reclaimPolicy and allowVolumeExpansion
reclaimPolicy and allowVolumeExpansionreclaimPolicy
The PVs created by this class inherit this policy. The values are those of 05-02: Retain or Delete. And the default value, if you do not set it, is Delete.
This deserves a warning in capitals: the default classes of every managed cloud use Delete. That is, on a freshly created cluster, if your team deletes a database's PVC — or the namespace containing it — the disk is destroyed. Always check this first thing when you arrive at a new cluster:
kubectl get sc -o custom-columns=\
NAME:.metadata.name,POLICY:.reclaimPolicy,EXPANSION:.allowVolumeExpansion,\
MODE:.volumeBindingModeAs the class is almost immutable, you cannot change its policy; what you do is create your own class with Retain for the data that matters. That is exactly what we will do with rutasnorte-fast.
allowVolumeExpansion
It declares whether the PVCs of this class can be grown after they are created. It is one of the few fields that can be modified on the fly, with kubectl patch sc rutasnorte-fast -p '{"allowVolumeExpansion": true}'.
Set it to true on every class intended for data that grows. The cost is zero and the alternative — migrating the data to a bigger volume with the database stopped — is an operation for the small hours. The mechanics of expansion are detailed in 05-05, including the fact that it cannot be shrunk.
volumeBindingMode: the field that averts a disaster
volumeBindingMode: the field that averts a disasterIt is the least understood field of the StorageClass and the one that causes the most incidents in the cloud. It has two values:
| Mode | When the volume is created and bound | Consequence |
|---|---|---|
Immediate |
As soon as the PVC is created, without knowing where the pod will go | The volume may be born where the pod does not fit |
WaitForFirstConsumer |
When the first pod using the PVC appears | The volume is born where the scheduler has decided to put the pod |
The specific case that explains it
Rutas Norte has its production cluster spread across three availability zones: eu-west-1a, eu-west-1b and eu-west-1c. With volumeBindingMode: Immediate:
sequenceDiagram
participant Dev as Team
participant API as apiserver
participant Prov as CSI provisioner
participant Sched as Scheduler
Dev->>API: creates a 200Gi PVC
API->>Prov: there is a pending PVC
Prov->>Prov: creates the disk in eu-west-1c<br/>(zone chosen BLINDLY)
Prov->>API: PV with nodeAffinity zone=eu-west-1c
Note over API: PVC Bound. Everything looks fine.
Dev->>API: creates the bookings-postgres pod
API->>Sched: schedule this pod
Sched-->>API: Pod Pending: the volume demands eu-west-1c<br/>and there is no free CPU/memory there
The pod stays Pending indefinitely with an event along these lines:
0/6 nodes are available: 4 node(s) had volume node affinity conflict,
2 Insufficient cpu. preemption: 0/6 nodes are available.And there is no comfortable fix: the disk already exists in the wrong zone and it does not move. You have to delete the PVC, delete the disk and start again, trusting to luck.
With WaitForFirstConsumer the order is reversed: the scheduler decides first, taking into account CPU, memory, affinities, taints and tolerations (06-05), and then the provisioner creates the disk in that node's zone. The conflict is impossible by construction.
The side effect to recognise so as not to panic: with WaitForFirstConsumer, a freshly created PVC stays Pending on purpose until a pod that uses it exists, and kubectl describe pvc says so with the event WaitForFirstConsumer: waiting for first consumer to be created before binding.
That is not an error. It is the mode working. As soon as you create the Deployment, the PVC goes to Bound in seconds.
Rutas Norte rule: WaitForFirstConsumer on every class, no exceptions. The only scenario where Immediate is preferable is a single-zone cluster with uniform network storage, and even there it adds nothing.
- The default class and its traps
A cluster can have one StorageClass marked as the default. When a PVC omits the storageClassName field, an admission controller injects the name of that class into it.
It is marked with the annotation storageclass.kubernetes.io/is-default-class: "true" in the class's metadata, and in the output of kubectl get sc it appears flagged as standard (default). Changing it is a couple of kubectl patch commands:
# remove the mark from the current one
kubectl patch sc standard -p \
'{"metadata":{"annotations":{"storageclass.kubernetes.io/is-default-class":"false"}}}'
# put it on the new one
kubectl patch sc rutasnorte-standard -p \
'{"metadata":{"annotations":{"storageclass.kubernetes.io/is-default-class":"true"}}}'If there are two default classes
Nothing stops you marking two, and it is not a configuration error you can spot at a glance. The behaviour is as follows: the admission controller picks the most recently created one and ignores the rest. In earlier versions it would go as far as rejecting the PVC. In any case it is a misconfiguration, and its symptom is one of the worst: the PVCs work, but some end up on fast, expensive disks and others on slow ones, with no apparent pattern.
kubectl get sc -o json | jq -r '.items[] |
select(.metadata.annotations["storageclass.kubernetes.io/is-default-class"]=="true") |
.metadata.name' # if it returns more than one line, fix it todayIf there is no default class at all
Then a PVC that omits storageClassName receives no class and stays Pending forever, looking for a static PV that probably does not exist, with the event ProvisioningFailed: no persistent volumes available for this claim and no storage class is set.
storageClassName: "" versus omitting the field
storageClassName: "" versus omitting the fieldThis distinction is worth a section of its own because it is subtle, it is constantly forgotten and it produces baffling failures.
| In the PVC | What the admission controller does | Which PV it can match |
|---|---|---|
| Field omitted | It injects the default class | With PVs of that class, or one is provisioned |
storageClassName: "" |
It touches nothing: the PVC is left classless | Only with classless PVs (static ones) |
storageClassName: my-class |
It touches nothing | With PVs of my-class, or one is provisioned |
The empty string is the explicit way of saying: "disable dynamic provisioning for this PVC; I want one specific static PV". It is what you have to write when you are restoring a pre-existing volume and you do not want the cluster to create a new, empty one alongside. The classic mistake and its symptom: on a cluster with a default class, someone creates a classless static PV to restore some data, writes a PVC omitting storageClassName, and instead of binding to the PV with the data, the cluster provisions a new, empty volume. The PVC ends up Bound, the application starts with no errors and the database is empty. It is a silent failure, and the protection is twofold: storageClassName: "" and volumeName pointing at the specific PV.
# Restore: I want THAT volume, not a new one
spec:
storageClassName: "" # no class: no dynamic provisioning
volumeName: pv-bookings-postgres-10gi # and moreover, exactly this one
accessModes: ["ReadWriteOnce"]
resources: { requests: { storage: 10Gi } }
- The classes shipped by minikube and the managed clouds
Your practice cluster already has a class, supplied by the storage-provisioner addon you enabled in 01-04:
Name: standard IsDefaultClass: Yes
Provisioner: k8s.io/minikube-hostpath
Parameters: <none> AllowVolumeExpansion: <unset>
ReclaimPolicy: Delete VolumeBindingMode: Immediatek8s.io/minikube-hostpath creates directories under /tmp/hostpath-provisioner/<namespace>/<pvc-name> inside the VM. Its limitations are what you would expect from a single-node cluster: it supports neither expansion nor snapshots, and its policy is Delete. It is perfectly good for learning the flow, but not for practising what is in 05-05; there we will install the hostpath CSI driver, which does support them.
An indicative comparison table of the typical classes you will come across:
| Environment | Usual class | Provisioner | Backing | Expansion | Modes | Policy |
|---|---|---|---|---|---|---|
| minikube | standard |
k8s.io/minikube-hostpath |
VM directory | No | All (one node) | Delete |
| kind / k3s | local-path |
rancher.io/local-path |
Node directory | No | RWO | Delete |
| EKS (AWS) | gp2 / gp3 |
ebs.csi.aws.com |
EBS | Yes | RWO | Delete |
| EKS with files | efs-sc |
efs.csi.aws.com |
EFS | N/A | RWX | Delete |
| GKE (Google) | standard-rwo, premium-rwo |
pd.csi.storage.gke.io |
Persistent Disk | Yes | RWO | Delete |
| AKS (Azure) | managed-csi, managed-csi-premium |
disk.csi.azure.com |
Managed Disk | Yes | RWO | Delete |
| AKS with files | azurefile-csi |
file.csi.azure.com |
Azure Files | Yes | RWX | Delete |
Prices and performance vary and there is no point memorising them; what is worth retaining is the pattern: every cloud offers at least one standard block disk, one premium disk and a more expensive shared filesystem for ReadWriteMany, and the pre-installed classes come with Delete.
- The Rutas Norte storage classes
With all of the above, here is the design of the platform's catalogue: two classes, each with a clear intent.
rutasnorte-fast: for the data that cannot be lost
# k8s/base/storageclass-fast.yaml
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: rutasnorte-fast
labels: { app.kubernetes.io/part-of: rutas-norte }
annotations:
storageclass.kubernetes.io/is-default-class: "false"
rutasnorte.example/description: "Encrypted SSD for databases. Retains the volume when the PVC is deleted."
provisioner: ebs.csi.aws.com # in pre and pro; in minikube see below
parameters:
type: gp3
iops: "6000"
throughput: "250"
encrypted: "true" # personal customer data: mandatory
fsType: ext4
reclaimPolicy: Retain # <-- the central decision
allowVolumeExpansion: true # the DB will grow every peak season
volumeBindingMode: WaitForFirstConsumer
mountOptions: ["noatime"]Why Retain on the database class. It is the most important decision in the catalogue and it is worth being able to defend:
- The
bookings-postgresPVC contains the name, ID number, phone number and email address of every customer and every booking sold. Losing it is not a technical incident: it is a business outage and a personal data breach. - The mechanisms that can delete a PVC without anyone intending it are many and everyday: a mistaken
kubectl delete namespace(02-06), akubectl delete -f k8s/over the whole directory, a GitOps tool that syncs and "prunes" resources no longer in Git (10-05), an environment clean-up script. - With
Delete, any of those accidents destroys the disk irreversibly. WithRetain, it leaves a PV inReleasedthat is rescued in two minutes with thekubectl patchof 05-02. - The cost of
Retainis real but small: orphaned volumes that have to be cleaned up by hand and that keep being billed. It is offset by a monthly review of PVs inReleased.
The general rule that follows: Retain for business data, Delete for rebuildable data.
rutasnorte-standard: for everything else
# k8s/base/storageclass-standard.yaml
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: rutasnorte-standard
labels: { app.kubernetes.io/part-of: rutas-norte }
annotations:
storageclass.kubernetes.io/is-default-class: "true"
rutasnorte.example/description: "Standard disk for rebuildable data. Deleted along with the PVC."
provisioner: ebs.csi.aws.com
parameters: { type: gp3, encrypted: "true", fsType: ext4 }
reclaimPolicy: Delete
allowVolumeExpansion: true
volumeBindingMode: WaitForFirstConsumerIt is the default class, and that choice is deliberate too: if someone creates a PVC without thinking, let them land on the cheap, deletable class, not the expensive, retained one. The database class has to be asked for explicitly, which forces a conscious decision.
rutasnorte-fast |
rutasnorte-standard |
|
|---|---|---|
| Intended use | bookings-postgres, backup volume (05-06) |
Working spaces, attachments, reports |
| Reclaim policy | Retain |
Delete |
| Performance | SSD with provisioned IOPS | Standard SSD |
| Expansion | Yes | Yes |
| Default | No | Yes |
| Encryption | Yes | Yes |
In the practice cluster
In minikube there is no ebs.csi.aws.com. To practise, the same classes are created with the local provisioner. Note that the name does not change: it is the only thing the application manifests see, and that is why the bookings-postgres PVC is identical in minikube and in production.
# k8s/environments/dev/storageclasses-minikube.yaml
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: rutasnorte-fast
labels: { app.kubernetes.io/part-of: rutas-norte }
provisioner: k8s.io/minikube-hostpath
reclaimPolicy: Retain # the contract: it is preserved
allowVolumeExpansion: false # the minikube provisioner does not support it
volumeBindingMode: Immediate # minikube has a single node
# --- plus a second rutasnorte-standard class, identical except for:
# reclaimPolicy: Delete and the is-default-class: "true" annotationThat the dev configuration differs from production's in parameters and provisioner is normal and correct: what must stay identical is the contract, that is, the class name and its reclaim policy.
- Hands-on: a PVC that creates its own volume
We apply the classes and check the effect. First, the default class in its place:
kubectl apply -f k8s/environments/dev/storageclasses-minikube.yaml
kubectl patch sc standard -p \
'{"metadata":{"annotations":{"storageclass.kubernetes.io/is-default-class":"false"}}}'
kubectl get scNAME PROVISIONER RECLAIMPOLICY VOLUMEBINDINGMODE AGE
rutasnorte-standard (default) k8s.io/minikube-hostpath Delete Immediate 5s
rutasnorte-fast k8s.io/minikube-hostpath Retain Immediate 5s
standard k8s.io/minikube-hostpath Delete Immediate 12dNow, and this is the important part: we delete the static PersistentVolume we created by hand in 05-02, we check with kubectl get pv that none are left (No resources found) and we create a new PVC with no PV waiting.
# k8s/base/report-attachments-pvc.yaml
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: report-attachments
namespace: rutas-norte-dev
labels: { app: occupancy-reports, app.kubernetes.io/part-of: rutas-norte, environment: dev }
spec:
accessModes: ["ReadWriteOnce"]
resources: { requests: { storage: 3Gi } }
# no storageClassName: it will use the default class (rutasnorte-standard)NAME STATUS VOLUME CAPACITY STORAGECLASS
persistentvolumeclaim/report-attachments Bound pvc-7f3e1a2b-9c4d-... 3Gi rutasnorte-standard
NAME RECLAIM POLICY STATUS CLAIM
persistentvolume/pvc-7f3e1a2b-9c4d-... Delete Bound rutas-norte-dev/report-attachmentsNote the four details that sum up the lesson:
- A PersistentVolume you did not write has appeared. The provisioner created it on seeing the PVC.
- Its name is
pvc-<pvc-uid>: if a PV name starts withpvc-, it is dynamic. CAPACITY: 3Gi, exactly what was asked for, andRECLAIM POLICY: Delete, inherited from the default class.
And the check that the database PVC, by asking for the rutasnorte-fast class, gets a volume with Retain:
kubectl apply -f k8s/base/bookings-postgres-pvc.yaml
kubectl apply -f k8s/base/bookings-postgres-deployment.yaml
kubectl get pv -o custom-columns=\
NAME:.metadata.name,CLASS:.spec.storageClassName,\
POLICY:.spec.persistentVolumeReclaimPolicy,CLAIM:.spec.claimRef.nameNAME CLASS POLICY CLAIM
pvc-7f3e1a2b-9c4d-... rutasnorte-standard Delete report-attachments
pvc-2b8c9d0e-1f2a-... rutasnorte-fast Retain bookings-postgres-dataTwo volumes, two different policies, neither one written by hand. That is the catalogue working.
- Migrating a PVC from one class to another
The last operational piece, and the short answer is blunt: you cannot change the class of an existing PVC. The storageClassName field is immutable once bound, and rightly so: the class determines the physical disk, and a disk does not change type because you edit a YAML file. A kubectl patch pvc ... -p '{"spec":{"storageClassName":"rutasnorte-fast"}}' crashes into:
The PersistentVolumeClaim "bookings-postgres-data" is invalid:
spec: Forbidden: spec is immutable after creation except resources.requests
and volumeAttributesClassName for bound claimsThe message also tells you in advance the only thing that can be changed: resources.requests, that is, the expansion of 05-05. The real migration procedure involves copying the data, and there are three variants depending on what you can afford:
a) With downtime (the standard for a database)
# 1. Backup BEFORE anything else (05-06)
kubectl exec -n rutas-norte-dev deploy/bookings-postgres -- \
pg_dump -U rutasnorte -d bookings -Fc -f /tmp/pre-migration.dump
# 2. Stop writes: with no pods, there are no changes to the volume
kubectl scale deploy bookings-postgres -n rutas-norte-dev --replicas=0
# 3. Create the target PVC bookings-postgres-data-v2 with the new class
kubectl apply -f k8s/base/bookings-postgres-pvc-v2.yaml# 4. A Job that mounts BOTH volumes and copies
apiVersion: batch/v1
kind: Job
metadata:
name: migrate-postgres-volume
namespace: rutas-norte-dev
labels: { app: bookings-postgres, app.kubernetes.io/part-of: rutas-norte, environment: dev }
spec:
backoffLimit: 2
template:
metadata:
labels: { app: bookings-postgres, environment: dev }
spec:
restartPolicy: Never
containers:
- name: copy
image: busybox:1.36
# -a preserves permissions, owners and timestamps: essential
command: ["sh", "-c", "cp -a /source/. /target/ && ls -la /target"]
volumeMounts:
- { name: source, mountPath: /source, readOnly: true }
- { name: target, mountPath: /target }
resources:
requests: { cpu: 200m, memory: 256Mi }
limits: { cpu: "1", memory: 512Mi }
volumes: # BOTH PVCs at once
- name: source
persistentVolumeClaim: { claimName: bookings-postgres-data }
- name: target
persistentVolumeClaim: { claimName: bookings-postgres-data-v2 }kubectl apply -f migrate-postgres-volume.yaml
kubectl wait --for=condition=complete job/migrate-postgres-volume -n rutas-norte-dev --timeout=600s
# 5. Point the Deployment at the new PVC and start it
kubectl patch deploy bookings-postgres -n rutas-norte-dev --type=json -p='[
{"op":"replace",
"path":"/spec/template/spec/volumes/0/persistentVolumeClaim/claimName",
"value":"bookings-postgres-data-v2"}]'
kubectl scale deploy bookings-postgres -n rutas-norte-dev --replicas=1
# 6. VERIFY before deleting anything
kubectl exec -n rutas-norte-dev deploy/bookings-postgres -- \
psql -U rutasnorte -d bookings -c "SELECT count(*) FROM bookings;"Step 6 is not optional. Do not delete the source PVC until you have verified the data at the target, and even then leave it a few days. Remember that if its class is Retain, deleting it leaves the PV in a recoverable Released state.
b) With the logical dump as an intermediary
For a database it is usually cleaner: pg_dump from the source, create the new empty PVC, start PostgreSQL on it and restore with pg_restore. It is slower but it validates the data along the way (a binary copy would carry corruption across; a logical dump does not restore if it is corrupt). It is detailed in 05-06.
c) With no downtime
It requires application-level replication: bring up a PostgreSQL replica on the new volume, sync it and promote it. It is a database procedure, not a Kubernetes one, and it is delegated to an operator (06-07).
Common Mistakes and Tips
| Mistake | Symptom | Fix |
|---|---|---|
| Leaving the cloud's default class on a data volume | The PVC is deleted and the disk disappears | Your own class with reclaimPolicy: Retain |
Using volumeBindingMode: Immediate across several zones |
volume node affinity conflict, pod Pending forever |
WaitForFirstConsumer always |
Panicking at WaitForFirstConsumer |
People think the PVC is broken | It is expected: it binds when the pod is created |
| Two classes marked as default | Volumes on unpredictable classes | Leave only one; check with jq |
| No default class at all | Classless PVC Pending forever |
Mark one, or name the class in every PVC |
Confusing "" with omitting the field |
An empty volume is provisioned instead of using the PV with the data | "" and volumeName on every restore |
| Trying to change a PVC's class | spec is immutable after creation |
Copy the data to a new PVC |
Forgetting allowVolumeExpansion |
The DB disk cannot be grown without downtime | Set it to true; it is modifiable on the fly |
Using kubernetes.io/* provisioners |
The PVC is never provisioned | The in-tree ones are removed; use CSI |
| Not encrypting the database disk | Personal data in the clear on the storage | encrypted: "true" or its equivalent |
Tips:
- Document every class with an annotation (
rutasnorte.example/description). Whoever picks the class is usually a developer who does not know what lies underneath, and that one line saves them asking. - Audit the catalogue on day one on any cluster with a
kubectl get sc -o custom-columns=...showingprovisioner,reclaimPolicy,allowVolumeExpansion,volumeBindingModeand the default-class annotation all at once. - Few classes, with names that state the intent.
rutasnorte-fastandrutasnorte-standardare understandable;sc-gp3-6000iops-encryptedrequires knowing AWS and ties the name to the provider. - Class names are the contract between environments. Keep the same names in dev, pre and pro even if underneath there are different provisioners: that way the application manifests are identical.
Exercises
Exercise 1: build the Rutas Norte catalogue
On your minikube:
- Create
rutasnorte-fastandrutasnorte-standardwith the local provisioner, leavingrutasnorte-standardas the default class and removing the mark fromstandard. - Verify with a single command that there is only one default class.
- Create a 1 GiB PVC without
storageClassNameand check which class it ends up on and with what policy. - Create another 1 GiB PVC with
storageClassName: rutasnorte-fastand compare the reclaim policy of both PVs. - Delete both PVCs and explain why one of the PVs disappears and the other does not.
Exercise 2: the zone incident
A colleague from the platform team has created this class for the production cluster, which has nodes in eu-west-1a, eu-west-1b and eu-west-1c:
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata: { name: rutasnorte-fast }
provisioner: ebs.csi.aws.com
parameters: { type: gp3 }
volumeBindingMode: ImmediateThe next day, the bookings-postgres pod has been Pending for 40 minutes with this event:
Explain exactly what has happened, why the PVC is Bound while the pod does not start, how the incident is resolved today and what three changes have to be made to the class so that it does not happen again.
Exercise 3: the PVC that bound to the wrong volume
The team has to restore the bookings-postgres data in pre-production. An administrator has created by hand a PV called pv-restore-pre (20 GiB, Retain, without storageClassName) pointing at the volume with the recovered data. A developer applies this PVC:
apiVersion: v1
kind: PersistentVolumeClaim
metadata: { name: bookings-postgres-data, namespace: rutas-norte-pre }
spec:
accessModes: ["ReadWriteOnce"]
resources: { requests: { storage: 20Gi } }The PVC goes Bound in seconds, PostgreSQL starts with no errors… and the database is empty. Explain what has happened and write the correct PVC.
Solutions
Exercise 1
# 1
kubectl apply -f k8s/environments/dev/storageclasses-minikube.yaml
kubectl patch sc standard -p \
'{"metadata":{"annotations":{"storageclass.kubernetes.io/is-default-class":"false"}}}'
# 2
kubectl get sc -o json | jq -r \
'.items[] | select(.metadata.annotations["storageclass.kubernetes.io/is-default-class"]=="true") | .metadata.name'
# -> rutasnorte-standard (a single line: correct)
# 3 and 4: two 1Gi PVCs, one WITHOUT storageClassName (test-default)
# and another with storageClassName: rutasnorte-fast (test-fast)
kubectl apply -f pvc-class-tests.yaml
kubectl get pv -o custom-columns=\
NAME:.metadata.name,CLASS:.spec.storageClassName,\
POLICY:.spec.persistentVolumeReclaimPolicy,CLAIM:.spec.claimRef.nameNAME CLASS POLICY CLAIM
pvc-a1b2c3d4-... rutasnorte-standard Delete test-default
pvc-e5f6a7b8-... rutasnorte-fast Retain test-fastThe classless PVC received rutasnorte-standard by injection from the admission controller (check it with kubectl get pvc test-default -o yaml: the field appears written into the object, even though you did not put it there). And on deleting both PVCs (step 5), kubectl get pv only returns pvc-e5f6a7b8-... 1Gi Retain Released rutas-norte-dev/test-fast.
Only one is left. The rutasnorte-standard one had Delete and was destroyed with its PVC; the rutasnorte-fast one inherited Retain from its class and was left in Released with the data inside. It is exactly the protection we were after for the database. Clean-up: kubectl delete pv pvc-e5f6a7b8-....
Exercise 2
What has happened. The class uses volumeBindingMode: Immediate (and moreover that is the default value, so simply omitting it is enough to fall into the trap). When the PVC was created, the EBS provisioner created the disk immediately, picking a zone with no information whatsoever about where the pod would end up — let us say eu-west-1c. The resulting PV carries a nodeAffinity demanding topology.kubernetes.io/zone=eu-west-1c.
Why the PVC is Bound and the pod does not start. They are two independent decisions: the PVC-PV binding is done by the PersistentVolume controller and only looks at capacity, modes and class; the pod's placement is done by the scheduler, which on top of the volume's affinity evaluates CPU, memory and taints. The PVC is perfectly bound; what does not exist is a node satisfying both the disk's zone and the pod's resources. The 6 node(s) had volume node affinity conflict are the nodes in the other two zones, and the 3 Insufficient memory are those in eu-west-1c, which do work zone-wise but are full — bookings-postgres asks for 2 GiB with Guaranteed QoS (03-05).
How it is resolved today. You have to choose between two routes. Option A is to make room in eu-west-1c — locate the nodes with kubectl get nodes -L topology.kubernetes.io/zone, look at their Allocated resources with kubectl describe node, and free up load or add a node there. Option B is to rebuild the volume: delete the Deployment and the PVC in rutas-norte-pro, correct the class and apply again.
Option B is only acceptable because the volume has no data yet. If it did, the only way out would be a restore from backup (05-06), because an EBS disk does not change zone.
The three changes to the class:
provisioner: ebs.csi.aws.com
parameters: { type: gp3, encrypted: "true" } # 4th: personal data (advisable)
reclaimPolicy: Retain # 1: do not destroy the data on deleting the PVC
allowVolumeExpansion: true # 2: be able to grow without downtime
volumeBindingMode: WaitForFirstConsumer # 3: THE fix for the incidentvolumeBindingMode: WaitForFirstConsumer, which reverses the order: schedule the pod first, then create the disk in its zone. It eliminates the conflict by construction.reclaimPolicy: Retain, because the missing field defaulted toDeleteand this is the production database.allowVolumeExpansion: true, so as not to have to repeat this manoeuvre the day the 200 GiB fall short.
Remember that the StorageClass cannot be edited: it has to be deleted and recreated. The PVs that already exist keep the configuration they were born with.
Exercise 3
What has happened. The PVC omits storageClassName. The admission controller therefore injected the cluster's default class (rutasnorte-standard). From then on, the PV pv-restore-pre was discarded as a candidate — it is classless, and the class must match exactly — and in its place the provisioner created a new, empty, 20 GiB volume. The PVC bound to that new volume in seconds, PostgreSQL found an empty directory, ran initdb and started up happily: a brand-new, empty database, without a single error in the logs.
It is the most dangerous failure of the lesson precisely because nothing fails. And there is collateral damage: the PV with the data is still Available, but if nobody notices and the application starts operating on the empty database, the restore window is lost.
The correct PVC carries double protection:
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: bookings-postgres-data
namespace: rutas-norte-pre
labels: { app: bookings-postgres, app.kubernetes.io/part-of: rutas-norte, environment: pre }
spec:
accessModes: ["ReadWriteOnce"]
volumeMode: Filesystem
# 1: EMPTY string -> no class -> no dynamic provisioning
storageClassName: ""
# 2: and moreover, exactly THIS volume and no other
volumeName: pv-restore-pre
resources: { requests: { storage: 20Gi } }And the safe restore procedure, so as not to repeat the incident:
# 1. Pre-assign the PV to the PVC (claimRef with no uid): nobody else can take it
kubectl patch pv pv-restore-pre -p '{"spec":{"claimRef":{
"apiVersion":"v1","kind":"PersistentVolumeClaim",
"namespace":"rutas-norte-pre","name":"bookings-postgres-data"}}}'
# 2. Create the PVC and VERIFY which volume it has bound to
kubectl apply -f pvc-restore.yaml
kubectl get pvc bookings-postgres-data -n rutas-norte-pre \
-o jsonpath='{.spec.volumeName}'; echo # it must say pv-restore-pre
# 3. Only then, start the application and count the bookings
kubectl scale deploy bookings-postgres -n rutas-norte-pre --replicas=1Step 2 — checking volumeName before starting the application — is the verification that would have avoided the incident.
Conclusion
You have made the leap that changes a cluster's daily operations: from static provisioning, where a human creates each PersistentVolume in advance with its waste, its waiting and its orphans, to dynamic, where the application team creates a PVC and the exact volume appears in seconds. You have watched a PV you did not write come into being, with the pvc-<uid> name that gives away its origin, at precisely the size requested and with the policy inherited from its class.
You know the StorageClass field by field. The provisioner, which identifies the driver — CSI in production, k8s.io/minikube-hostpath in practice, and never the in-tree kubernetes.io/* ones, now removed. The parameters, opaque to Kubernetes and specific to each driver, where the disk type, the IOPS and, mandatorily at Rutas Norte, encryption at rest go. The reclaimPolicy the PVs inherit, with the warning you must check on day one on any new cluster: the clouds' default classes come with Delete, so deleting a PVC deletes the disk. The allowVolumeExpansion, which costs nothing and saves an all-nighter of migration. And, above all, the volumeBindingMode: Immediate creates the disk blindly and may leave it in a zone where the pod does not fit, producing a volume node affinity conflict that is irreparable without restoring from backup; WaitForFirstConsumer reverses the order — schedule the pod first, then create the disk where that pod is — and eliminates the problem by construction, at the price of a deliberate Pending that no longer alarms you.
You have mastered the default class: its storageclass.kubernetes.io/is-default-class annotation, the mess produced by having two and the eternal Pending of having none. And you tell apart precisely the two silences: omitting storageClassName lets the default class be injected into you, while storageClassName: "" means "no class, no dynamic provisioning, I want a static PV". That two-quote difference is what separates a correct restore from an empty database that starts without a single error, as you saw in the third exercise.
And you have designed the platform's catalogue: rutasnorte-fast with Retain for bookings-postgres, because the disk holds personal customer data and there are too many everyday ways of deleting a PVC by accident; and rutasnorte-standard with Delete as the default class, so that a slip always lands on the cheap and deletable and the protected class has to be asked for deliberately. The names stay identical in dev, pre and pro even though the provisioner and the parameters change underneath: the class name is the contract between environments. And you know that a PVC's class is immutable, so migrating means copying the data, with downtime and with verification before deleting anything.
What remains is to open the black box. We have said "the provisioner creates the volume" without explaining who it is, how it finds out there is a pending PVC, how it attaches the disk to the node and how it mounts it in the pod. And there are two capabilities the StorageClass announces but that we have not used yet: the expansion enabled by allowVolumeExpansion and the snapshots, indispensable before touching the schema of a production database. All that is the machinery of the Container Storage Interface, and it is the next lesson: Dynamic Provisioning, Expansion and Snapshots.
Kubernetes Course
Module 1: Introduction to Kubernetes
- What Is Kubernetes?
- Kubernetes Architecture
- Key Concepts and Terminology
- Setting Up a Kubernetes Cluster
- The Kubernetes CLI: kubectl
- Objects, YAML Manifests and the Declarative Model
- The Course Project: the Rutas Norte Platform
Module 2: Core Kubernetes Components
- Pods
- ReplicaSets
- Deployments
- Updates, Rollbacks and Deployment Strategies
- Services
- Namespaces
- Labels, Selectors and Annotations
Module 3: Configuration and Secret Management
- ConfigMaps
- Secrets
- Environment Variables
- Resource Quotas and Limits
- LimitRanges and Quality of Service (QoS) Classes
- ServiceAccounts and API Access from Pods
Module 4: Networking in Kubernetes
- Cluster Networking
- Service Types
- Internal DNS and Service Discovery
- Ingress Controllers
- TLS and Certificate Management with cert-manager
- Network Policies
Module 5: Storage in Kubernetes
- Volumes
- Persistent Volumes
- Persistent Volume Claims
- Storage Classes
- Dynamic Provisioning, Expansion and Snapshots
- Backup and Restore of Persistent Data
Module 6: Advanced Kubernetes Concepts
- StatefulSets
- DaemonSets
- Jobs and CronJobs
- Init Containers, Sidecars and Multi-Container Patterns
- Scheduling: Affinity, Taints and Tolerations
- Custom Resource Definitions (CRDs)
- Operators and the Controller Pattern
Module 7: Monitoring and Logging
- Health Checks and Probes
- Metrics Server and kubectl top
- Monitoring with Prometheus
- Visualization and Alerting with Grafana and Alertmanager
- Centralized Logging with Elasticsearch, Fluentd and Kibana (EFK)
- Application Debugging and Cluster Events
Module 8: Kubernetes Security
- Role-Based Access Control (RBAC)
- Security Contexts and Container Hardening
- Pod Security Policies and Pod Security Standards
- Network Security
- Image Security
- Auditing, Scanning and Vulnerability Management
Module 9: Scaling and Performance
- Horizontal Pod Autoscaling
- Vertical Pod Autoscaling
- Cluster Autoscaling
- Event-Driven and Custom-Metric Scaling with KEDA
- High Availability: PodDisruptionBudgets and Topology
- Performance Tuning
Module 10: Kubernetes Ecosystem and Tooling
- Minikube and Local Environments with kind
- Kubeadm
- Helm
- Kustomize
- GitOps with Argo CD and Flux
- Managed Kubernetes: EKS, AKS and GKE
Module 11: Case Studies and Real-World Applications
- Deploying a Web Application
- Running Stateful Applications
- CI/CD with Kubernetes
- Deployment Strategies: Blue-Green and Canary
- Multi-Cluster Management
- Production Operations: Incidents, Runbooks and Costs
