The previous lesson ended with an unsolved problem: bookings-api and web-store are running in rutas-norte-dev, but as bare pods. If the node goes down, if the kubelet evicts them for lack of memory or if somebody deletes them, they vanish and nobody brings them back to life. We need something that watches permanently and acts. That something is a controller, and the first one we are going to take apart is the ReplicaSet: the object whose sole mission is to guarantee that at all times a specific number of pods matching a set of labels exists. In this lesson you will see what a controller is exactly and how the ReplicaSet controller embodies the reconciliation loop you studied in the architecture, you will write your first ReplicaSet manifest for web-store, you will verify self-healing by deleting pods on purpose, you will understand how cascading deletion relies on ownerReferences, you will discover the adoption mechanism for orphaned pods and why an overly broad selector is a time bomb, and you will learn to scale with kubectl scale. And you will finish knowing why, despite all of this, you will almost never write a ReplicaSet by hand.
Contents
- What a controller is
- The ReplicaSet controller and the reconciliation loop
- Anatomy of the manifest:
replicas,selectorandtemplate - The golden rule: the
templatemust match theselector - Self-healing live with
web-store ownerReferencesand cascading deletion- Adoption of orphaned pods and the danger of broad selectors
- Manual scaling with
kubectl scale - ReplicaSet versus ReplicationController
- Why a ReplicaSet is almost never created by hand
- What a controller is
In Kubernetes Architecture we defined the reconciliation loop as the heart of the system. A controller is simply a program that runs that loop for one type of object:
flowchart LR
A["Observe<br/>the desired state<br/>(spec)"] --> B["Observe<br/>the actual state<br/>(the world)"]
B --> C{"Do they match?"}
C -->|Yes| D["Do nothing.<br/>Update status"]
C -->|No| E["Act to bring<br/>the actual closer<br/>to the desired"]
E --> A
D --> A
Traits shared by every Kubernetes controller that are worth internalising:
- They never finish. They are infinite loops, not one-shot scripts.
- They are idempotent. Running the loop a thousand times with the state already correct changes nothing.
- They only talk to the kube-apiserver. They never contact nodes or containers directly. They request changes from the API and the kubelet makes them real.
- They work by level, not by event. They do not react to "a pod has been deleted"; they constantly compare how many there are against how many there should be. That is why they are robust: if the controller was down for ten minutes, on coming back it corrects the accumulated difference with no need to replay history.
Most of them live inside the kube-controller-manager: the Deployment one, the ReplicaSet one, the Job one, the Namespace one, the endpoints one, the ServiceAccount one... All with the same anatomy and different responsibilities.
- The ReplicaSet controller and the reconciliation loop
Let's make the generic loop concrete for the case at hand. The ReplicaSet controller runs, non-stop, for every ReplicaSet in the cluster:
- It reads
spec.replicas: how many pods must exist (desired state). - It reads
spec.selector: which labels identify its pods. - It asks the API how many pods in that namespace match the selector and are not terminating (actual state).
- It compares:
- If pods are missing, it creates as many as are needed using
spec.templateas the mould. - If there are too many, it picks victims and deletes them.
- If they match, it does nothing and updates the
status.
- If pods are missing, it creates as many as are needed using
flowchart TD
RS["ReplicaSet web-store<br/>spec.replicas = 3<br/>spec.selector: app=web-store"]
RS --> Q{"Pods with app=web-store<br/>found: 2"}
Q -->|2 < 3| CREATE["Create 1 pod<br/>from spec.template"]
CREATE --> API["kube-apiserver<br/>persists the new pod"]
API --> SCH["kube-scheduler<br/>assigns it a node"]
SCH --> KUB["kubelet<br/>starts the containers"]
KUB --> Q
One decisive nuance, which explains everything that follows: the ReplicaSet does not keep a list of the pods it has created. It has no memory. On every turn of the loop it asks again, "which pods are there with these labels?". Its relationship with the pods is purely by label. This has two enormous consequences:
- If you delete one of its pods, on the next turn it sees one is missing and creates a new one. Self-healing.
- If a pod with those same labels appears out of nowhere, created by somebody else, the ReplicaSet counts it as its own. Adoption. And if that makes one too many, it will delete somebody.
How it picks who to delete when there are too many: the controller prioritises removing Pending pods over Running ones, those that have been ready for less time, those with more restarts and those on nodes with more replicas. In other words, it always sacrifices the least valuable.
- Anatomy of the manifest:
replicas, selector and template
replicas, selector and templateLet's turn web-store into something that looks after itself. Create the file in the project repository:
# k8s/base/web-store-replicaset.yaml
apiVersion: apps/v1
kind: ReplicaSet
metadata:
name: web-store
namespace: rutas-norte-dev
labels:
app: web-store
app.kubernetes.io/part-of: rutas-norte
environment: dev
spec:
replicas: 3
selector:
matchLabels:
app: web-store
environment: dev
template:
metadata:
labels:
app: web-store
app.kubernetes.io/part-of: rutas-norte
environment: dev
spec:
containers:
- name: nginx
image: nginx:1.27-alpine
ports:
- name: http
containerPort: 80
resources:
requests:
cpu: "50m"
memory: "64Mi"
limits:
cpu: "200m"
memory: "128Mi"A field-by-field read, paying attention to what changes compared with the bare pod:
apiVersion: apps/v1. Here there is a group. The Pod is in the core group (v1), but ReplicaSet, Deployment, StatefulSet and DaemonSet live in theappsgroup. If you write plainv1, the API will reject the manifest withno matches for kind "ReplicaSet" in version "v1".metadata(the top-level one): it identifies the ReplicaSet, not the pods. Its labels are there so you can find the ReplicaSet itself withkubectl get rs -l ....spec.replicas: 3: the desired state. It is a number, not a range. Automatic scaling by load arrives in module 9.spec.selector.matchLabels: which pods are mine. Here we demand two labels: a pod must haveapp=web-storeandenvironment=devto count. The selector also accepts richer expressions (matchExpressions), which we will see in Labels, Selectors and Annotations. It is a mandatory and immutable field.spec.template: the pod template. It is exactly a pod manifest withoutapiVersionorkind, with its ownmetadata(the labels of the child pods) and itsspec(the containers). It is the mould the controller uses to manufacture replicas.
Notice the Russian-doll structure, the point where everybody gets stuck at first:
ReplicaSet
├── metadata <- identifies the ReplicaSet
└── spec
├── replicas <- how many pods
├── selector <- which pods are mine
└── template
├── metadata <- labels of EACH POD created
└── spec <- containers of EACH POD createdBefore applying it, remove the bare web-store pod so it does not interfere with the experiment (in section 7 you will see exactly what would happen if we did not):
kubectl delete pod web-store --ignore-not-found
kubectl apply -f k8s/base/web-store-replicaset.yaml
kubectl get rs,podspod "web-store" deleted
replicaset.apps/web-store created
NAME DESIRED CURRENT READY AGE
replicaset.apps/web-store 3 3 3 8s
NAME READY STATUS RESTARTS AGE
pod/bookings-api 1/1 Running 0 51m
pod/web-store-4kx7d 1/1 Running 0 8s
pod/web-store-9wq2m 1/1 Running 0 8s
pod/web-store-pv6cl 1/1 Running 0 8sTwo observations about that output:
- The pods are called
web-store-<random suffix>. The controller generates the name from the ReplicaSet's name plus five random characters, because there cannot be three objects with the same name in a namespace. - The ReplicaSet columns:
DESIREDisspec.replicas;CURRENT, how many pods exist;READY, how many are ready to take traffic. When all three match, the loop is in equilibrium.
- The golden rule: the
template must match the selector
template must match the selectorThis is the constraint that causes the most rejected manifests:
The labels in
spec.template.metadata.labelsmust satisfyspec.selector. If they do not, the API rejects the object.
The reason is pure logic: if the ReplicaSet manufactured pods that its own selector does not recognise, it would count zero pods of its own, create three more, fail to recognise those either, create another three… an infinite loop that would fill the cluster. Kubernetes prevents this at the root by validating the manifest.
Check it by triggering the error on purpose:
sed 's/app: web-store$/app: web-store-wrong/' k8s/base/web-store-replicaset.yaml \
| kubectl apply --dry-run=server -f -The ReplicaSet "web-store" is invalid: spec.template.metadata.labels:
Invalid value: map[string]string{"app":"web-store-wrong", ...}:
`selector` does not match template `labels`Note the asymmetry, which is valid and very useful: the template may have more labels than the selector. In our manifest the selector asks for app and environment, but the pods also carry app.kubernetes.io/part-of. That is allowed; what is forbidden is missing any of the labels the selector demands.
As a project rule: keep the selector as small and stable as possible (one or two labels that unambiguously identify the component and the environment) and add the rest of the metadata only in the template.
- Self-healing live with
web-store
web-storeThe time has come to check what module 1 left pending. Open two terminals.
In the first one, start watching:
In the second one, kill a pod behind its back (use one of the real names from your cluster):
In the first terminal you will see something like this:
NAME READY STATUS RESTARTS AGE
web-store-4kx7d 1/1 Running 0 4m
web-store-9wq2m 1/1 Running 0 4m
web-store-pv6cl 1/1 Running 0 4m
web-store-9wq2m 1/1 Terminating 0 4m
web-store-t8m4r 0/1 Pending 0 0s
web-store-t8m4r 0/1 ContainerCreating 0 0s
web-store-9wq2m 0/1 Terminating 0 4m
web-store-t8m4r 1/1 Running 0 2sUnder two seconds. And something very revealing: the replacement pod (web-store-t8m4r) appears in Pending before the deletion of the previous one has finished. The controller does not wait for the old one to disappear entirely; as soon as the API marks the pod as terminating, it stops counting it and acts.
The ReplicaSet's events tell the same story in writing:
Events:
Type Reason Age From Message
---- ------ ---- ---- -------
Normal SuccessfulCreate 4m replicaset-controller Created pod: web-store-4kx7d
Normal SuccessfulCreate 4m replicaset-controller Created pod: web-store-9wq2m
Normal SuccessfulCreate 4m replicaset-controller Created pod: web-store-pv6cl
Normal SuccessfulCreate 9s replicaset-controller Created pod: web-store-t8m4rCompare that with what happened in module 1 when we deleted the bare pod: nothing. Silence. Here, by contrast, there is a replicaset-controller signing off every creation. This is, in one line, the difference between a container and a platform. For Rutas Norte it means that March's four-and-a-half-hour night-time outage would have fixed itself in two seconds, in the small hours, without anybody noticing.
An even more forceful experiment: delete them all at once.
pod "web-store-4kx7d" deleted
pod "web-store-pv6cl" deleted
pod "web-store-t8m4r" deleted
NAME READY STATUS RESTARTS AGE
web-store-2jf9x 1/1 Running 0 3s
web-store-hs4bd 1/1 Running 0 3s
web-store-zq7nm 1/1 Running 0 3sThree new pods, with new names and new IPs. And there the course's next problem peeks out: if the IPs change every time, which address does web-store connect to in order to talk to bookings-api? That is the job of Services.
ownerReferences and cascading deletion
ownerReferences and cascading deletionHow does Kubernetes know that those pods "belong" to the ReplicaSet, if the relationship is by labels? Because, when creating them, the controller stamps a reference to their owner on every pod:
[
{
"apiVersion": "apps/v1",
"kind": "ReplicaSet",
"name": "web-store",
"uid": "6c1f8f0e-2b7a-4a91-9a3d-1d8f2c5e77b1",
"controller": true,
"blockOwnerDeletion": true
}
]Field by field:
| Field | Meaning |
|---|---|
kind / name / uid |
Who the owner is. The uid matters: if you delete the ReplicaSet and create another with the same name, it is a different object |
controller: true |
This owner is the controller of the pod. There can be only one |
blockOwnerDeletion: true |
The owner is not considered deleted until this child disappears |
The garbage collector, another kube-controller-manager controller, works on top of those references: it walks the objects, and when it finds one whose owner no longer exists, it deletes it. This is cascading deletion, and it is why kubectl delete rs web-store takes the three pods with it.
Kubernetes offers three propagation policies:
| Policy | Behaviour | How to request it |
|---|---|---|
Background (default) |
Deletes the owner immediately; the collector deletes the children afterwards | kubectl delete rs web-store |
Foreground |
Marks the owner as "being deleted", deletes the children first and the owner last | --cascade=foreground |
Orphan |
Deletes only the owner and leaves the children alive, stripping their ownerReferences |
--cascade=orphan |
The third is a very useful emergency resource, and it also sets the scene for the next section:
replicaset.apps "web-store" deleted
NAME READY STATUS RESTARTS AGE
pod/web-store-2jf9x 1/1 Running 0 6m
pod/web-store-hs4bd 1/1 Running 0 6m
pod/web-store-zq7nm 1/1 Running 0 6mThe ReplicaSet no longer exists, but the three pods keep serving traffic. They are now orphans: nobody is watching them. If you delete one, it does not come back.
- Adoption of orphaned pods and the danger of broad selectors
We have three orphaned pods with the labels app=web-store, environment=dev. Create the ReplicaSet again:
replicaset.apps/web-store created
NAME DESIRED CURRENT READY AGE
replicaset.apps/web-store 3 3 3 4s
NAME READY STATUS RESTARTS AGE
pod/web-store-2jf9x 1/1 Running 0 8m
pod/web-store-hs4bd 1/1 Running 0 8m
pod/web-store-zq7nm 1/1 Running 0 8mLook closely at the ages: 8 minutes. The ReplicaSet has just been born and has not created a single pod. It has adopted the three orphans because they matched its selector and had no owner. On adopting them, it has written its ownerReferences into them:
The exact rules of adoption:
- The pod must be in the same namespace.
- Its labels must match the selector.
- It must not already have an
ownerReferenceswithcontroller: true. A pod with an owner is not stolen: the ReplicaSet ignores it.
The danger: overly broad selectors
Here is the trap. Suppose somebody, in a hurry, defines the web-store ReplicaSet with a lazy selector:
# WRONG: selector far too broad
spec:
replicas: 3
selector:
matchLabels:
app.kubernetes.io/part-of: rutas-norte # every component carries this label!That selector matches every pod on the platform: web-store, bookings-api, redis-cache, the worker... The consequences, in a chain:
- The ReplicaSet counts every Rutas Norte pod in the namespace. Say 5.
- Its desired state is 3. There are 2 too many.
- It picks two victims and deletes them. It could perfectly well delete
bookings-apiand the worker. - Those components' controllers recreate them, the ReplicaSet again sees too many and deletes them once more. A war of controllers.
It is a real and fairly common incident in young clusters, and the symptom —pods deleting themselves with no explanation— is baffling until you understand the mechanism. Let's simulate it safely to see it with our own eyes:
# A bare pod that happens to carry the selector's labels
kubectl run intruder --image=nginx:1.27-alpine \
--labels="app=web-store,environment=dev,app.kubernetes.io/part-of=rutas-norte"
sleep 5
kubectl get pods -l app=web-storepod/intruder created
NAME READY STATUS RESTARTS AGE
intruder 1/1 Terminating 0 5s
web-store-2jf9x 1/1 Running 0 12m
web-store-hs4bd 1/1 Running 0 12m
web-store-zq7nm 1/1 Running 0 12mThe ReplicaSet has adopted intruder, counted 4 where there should have been 3 and executed it on the spot (it was the youngest, the preferred victim). Had it been an important pod instead of a test nginx, it would have gone just the same.
The project rules that follow from this:
- The selector must identify the component unambiguously:
app: <component>plusenvironment: <environment>, neverpart-ofalone. - No component shares a selector with another.
- Never create bare pods with the labels of a governed component. To debug, use different labels or ephemeral pods with
--rm.
- Manual scaling with
kubectl scale
kubectl scaleChanging the number of replicas means changing spec.replicas. There are three ways.
Imperative, fast (for incidents):
Conditional, very useful in scripts to avoid trampling on somebody else's changes:
If there are not exactly 5 replicas at that moment, the command fails instead of applying the change.
Declarative, the correct one according to the project conventions: edit replicas: 5 in the file, commit and apply.
| Approach | Advantage | Problem |
|---|---|---|
kubectl scale |
Instant, ideal during an incident | The cluster stops matching Git; the next apply reverts the change |
| Editing the manifest | Traceable, reviewable, reproducible | Slower |
The rule you already know from module 1 still stands: if a change is not in Git, it does not exist. Scale with kubectl scale to put out a fire, but move the change into the file straight afterwards.
Go back to 3 before moving on:
And watch what happens when scaling down:
NAME READY STATUS RESTARTS AGE
web-store-2jf9x 1/1 Running 0 16m
web-store-hs4bd 1/1 Running 0 16m
web-store-k9dpq 1/1 Terminating 0 40s
web-store-w2sxv 1/1 Terminating 0 40s
web-store-zq7nm 1/1 Running 0 16mThe ones that go are the youngest, exactly as we predicted: the controller sacrifices whatever is least established.
- ReplicaSet versus ReplicationController
You will come across old documentation and forum answers talking about ReplicationController. It is the ancestor of the ReplicaSet, from the Kubernetes 1.0 era, and it still exists for compatibility, but you should not use it.
| Aspect | ReplicationController | ReplicaSet |
|---|---|---|
| API group | v1 (core) |
apps/v1 |
| Selector | Simple equality only (app: web-store) |
Equality and sets (matchExpressions: in, notin, exists) |
| Selector field | spec.selector as a flat map |
spec.selector.matchLabels / matchExpressions |
| Used by the Deployment | No | Yes |
| Status | Obsolete in practice | Current |
The relevant functional difference is the set-based selector. Thanks to matchExpressions, a Deployment can issue queries such as "bookings-api pods whose pod-template-hash is one of these two", which is exactly what it needs in order to manage two simultaneous versions during a rolling update. With the ReplicationController's flat selector that was not expressible, and that is why Deployments were built on top of ReplicaSets.
Operational summary: if you see kind: ReplicationController, it is legacy code. Migrate to a Deployment.
- Why a ReplicaSet is almost never created by hand
And here comes the lesson's twist: after all of this, in your professional life you will write very few ReplicaSets. What you will write are Deployments.
The reason is that the ReplicaSet lacks the essential thing for operating a service: it does not know how to change version. It is superb at keeping N copies of a template, but if you modify the image in its template:
kubectl set image rs/web-store nginx=nginx:1.27.1-alpine
kubectl get pods -l app=web-store -o custom-columns=NAME:.metadata.name,IMAGE:.spec.containers[0].imagereplicaset.apps/web-store image updated
NAME IMAGE
web-store-2jf9x nginx:1.27-alpine
web-store-hs4bd nginx:1.27-alpine
web-store-zq7nm nginx:1.27-alpineThe pods still run the old image. The ReplicaSet only uses the template when it creates a pod; the existing ones are left alone because the desired state ("three pods with these labels") is still satisfied. To roll out the new version you would have to delete them by hand, one by one, waiting in between. No pacing control, no verification, no way back.
That is precisely what the Deployment brings: it manages ReplicaSets, not pods. When you change the image, it creates a new ReplicaSet with the new version and shifts replicas from one to the other in a controlled way.
| Need | ReplicaSet | Deployment |
|---|---|---|
| Keep N replicas alive | Yes | Yes (through a ReplicaSet) |
| Scale | Yes | Yes |
| Update the image with no downtime | No | Yes |
| Go back to the previous version | No | Yes (rollout undo) |
| Revision history | No | Yes |
| Pause a rollout halfway | No | Yes |
That is why the project rule is:
At Rutas Norte, ReplicaSets are not created by hand. Deployments are created, and they manage their ReplicaSets.
So what is this lesson for? For understanding the middle layer. When in the next lesson you see two ReplicaSets of the same Deployment coexisting during an update, or when in production you have to work out why there are pods from two versions at once, the explanation will lie in what you have just learned. ReplicaSets do not go away: they become invisible because somebody else manages them.
Leave the cluster ready for the next lesson:
replicaset.apps "web-store" deleted
pod "bookings-api" deleted
No resources found in rutas-norte-dev namespace.Common Mistakes and Tips
- Using
apiVersion: v1for a ReplicaSet. It is inapps/v1. The error isno matches for kind "ReplicaSet" in version "v1". templatelabels that do not match theselector. The API rejects the object withselector does not match template labels. Check that you are looking atspec.template.metadata.labelsand not the top-levelmetadata.labels: they are different places.- Overly broad selectors. This is the most destructive mistake in the lesson: a ReplicaSet can adopt and delete another component's pods. Selector = component + environment, always.
- Creating bare pods with the labels of a governed component. They will be adopted and, most likely, executed within seconds.
- Expecting a change to the
templateto update the existing pods. It does not. It only affects pods created from that moment on. That is the gap the Deployment fills. - Trying to change the
selectorof an existing ReplicaSet. It is immutable. You have to delete and recreate (or, in practice, let the Deployment handle the change by creating a new ReplicaSet). - Confusing
CURRENTwithREADY.CURRENTcounts pods that exist;READY, pods that can serve traffic. A3/3inCURRENTwith0inREADYis a broken rollout. - Tip:
kubectl get rs -o wideadds theCONTAINERS,IMAGESandSELECTORcolumns. It is the fastest way to see at a glance which version each ReplicaSet serves. - Tip: whenever something behaves inexplicably, always look at
kubectl get pod <name> -o jsonpath='{.metadata.ownerReferences}'. Knowing who owns a pod solves half the mysteries.
Exercises
Exercise 1: A bookings-api ReplicaSet and self-healing
Write k8s/base/bookings-api-replicaset.yaml with a ReplicaSet of 2 replicas of bookings-api in rutas-norte-dev, reusing the container from the Pods lesson (the node:20-alpine image with the minimal server) and respecting the project's three labels. The selector must use app and environment.
- Apply it and check that there are 2 pods.
- Delete one and measure how long the replacement takes to appear.
- Find out, with a single command, who owns the new pod.
- Show the ReplicaSet events that document the creation.
Exercise 2: Orphans and adoption
Starting from the previous exercise's ReplicaSet:
- Delete it with the policy that leaves the pods alive and check that they are still there.
- Verify that the pods no longer have an owner.
- Apply the same manifest again and prove, by looking at the pods' ages, that no new pods have been created.
- Explain in three lines why this is a real operational advantage and what would have happened if the new ReplicaSet had asked for 1 replica instead of 2.
Exercise 3: Investigate a dangerous selector
A colleague has deployed this manifest in rutas-norte-dev and ever since "pods delete themselves":
apiVersion: apps/v1
kind: ReplicaSet
metadata:
name: platform-monitor
namespace: rutas-norte-dev
spec:
replicas: 1
selector:
matchLabels:
app.kubernetes.io/part-of: rutas-norte
template:
metadata:
labels:
app.kubernetes.io/part-of: rutas-norte
spec:
containers:
- name: monitor
image: busybox:1.36
command: ["sh", "-c", "while true; do sleep 30; done"]- Explain exactly what is happening and why.
- Predict what would happen if exercise 1's 2
bookings-apipods were running when it is applied. - Propose the corrected manifest.
- State which command you would use, in a real cluster, to identify the culprit behind an unexpected deletion.
Solutions
Solution 1
# k8s/base/bookings-api-replicaset.yaml
apiVersion: apps/v1
kind: ReplicaSet
metadata:
name: bookings-api
namespace: rutas-norte-dev
labels:
app: bookings-api
app.kubernetes.io/part-of: rutas-norte
environment: dev
spec:
replicas: 2
selector:
matchLabels:
app: bookings-api
environment: dev
template:
metadata:
labels:
app: bookings-api
app.kubernetes.io/part-of: rutas-norte
environment: dev
spec:
containers:
- name: api
image: node:20-alpine
command: ["node", "-e"]
args:
- |
const http = require('http');
http.createServer((req, res) => {
res.writeHead(200, {'Content-Type': 'application/json'});
res.end(JSON.stringify({service: 'bookings-api', pod: process.env.HOSTNAME}));
}).listen(3000, () => console.log('bookings-api listening on 3000'));
ports:
- name: http
containerPort: 3000
resources:
requests:
cpu: "100m"
memory: "128Mi"
limits:
cpu: "500m"
memory: "256Mi"kubectl apply -f k8s/base/bookings-api-replicaset.yaml
kubectl get rs bookings-api
kubectl get pods -l app=bookings-apireplicaset.apps/bookings-api created
NAME DESIRED CURRENT READY AGE
bookings-api 2 2 2 12s
NAME READY STATUS RESTARTS AGE
bookings-api-c8n4t 1/1 Running 0 12s
bookings-api-xq7dz 1/1 Running 0 12spod "bookings-api-c8n4t" deleted
NAME READY STATUS RESTARTS AGE
bookings-api-m3kp9 1/1 Running 0 2s
bookings-api-xq7dz 1/1 Running 0 1mThe replacement appears in 1-2 seconds.
kubectl get pod bookings-api-m3kp9 \
-o jsonpath='{.metadata.ownerReferences[0].kind}/{.metadata.ownerReferences[0].name}{"\n"}'
kubectl describe rs bookings-api | tail -5ReplicaSet/bookings-api
Events:
Type Reason Age From Message
Normal SuccessfulCreate 1m replicaset-controller Created pod: bookings-api-c8n4t
Normal SuccessfulCreate 1m replicaset-controller Created pod: bookings-api-xq7dz
Normal SuccessfulCreate 8s replicaset-controller Created pod: bookings-api-m3kp9Solution 2
# 1. Deletion leaving orphans
kubectl delete rs bookings-api --cascade=orphan
kubectl get rs,pods -l app=bookings-apireplicaset.apps "bookings-api" deleted
NAME READY STATUS RESTARTS AGE
pod/bookings-api-m3kp9 1/1 Running 0 4m
pod/bookings-api-xq7dz 1/1 Running 0 5mThe empty output confirms that they no longer have ownerReferences.
# 3. Re-adoption
kubectl apply -f k8s/base/bookings-api-replicaset.yaml
kubectl get rs,pods -l app=bookings-apireplicaset.apps/bookings-api created
NAME DESIRED CURRENT READY AGE
replicaset.apps/bookings-api 2 2 2 3s
NAME READY STATUS RESTARTS AGE
pod/bookings-api-m3kp9 1/1 Running 0 5m
pod/bookings-api-xq7dz 1/1 Running 0 6mThe pods are 5 and 6 minutes old, while the ReplicaSet is 3 seconds old: nothing new has been created, the existing ones have been adopted.
- The operational advantage is that it lets you replace the controller without touching the service: if you need to recreate the ReplicaSet (for instance, because its selector has to change, and that is immutable), you delete it with
--cascade=orphan, apply the new one and the pods keep serving requests without a single second of downtime. Had the new ReplicaSet asked for 1 replica, it would have adopted both orphans, counted 2 where there should have been 1 and deleted one immediately, picking the youngest.
Solution 3
-
The selector
app.kubernetes.io/part-of: rutas-nortematches every pod on the platform, because that label is common to all six components by project convention. Theplatform-monitorReplicaSet adopts everything it finds in the namespace, compares it against itsreplicas: 1and deletes everything in excess. Since the other controllers recreate their pods, you enter a cycle of continuous creation and deletion: the "war of controllers". -
With the 2
bookings-apipods running, applying the manifest would make the monitor adopt them alongside its own pod (3 in total), see 2 too many and delete two of them, most likely thebookings-apiones if they are the youngest. Thebookings-apiReplicaSet would recreate them, the monitor would delete them again, and the API would fill up withSuccessfulDeleteandSuccessfulCreateevents. In all likelihood,bookings-apiwould go down intermittently. -
Corrected manifest: a selector that is exclusive to the component.
apiVersion: apps/v1
kind: ReplicaSet
metadata:
name: platform-monitor
namespace: rutas-norte-dev
labels:
app: platform-monitor
app.kubernetes.io/part-of: rutas-norte
environment: dev
spec:
replicas: 1
selector:
matchLabels:
app: platform-monitor # identifies ONLY this component
environment: dev
template:
metadata:
labels:
app: platform-monitor
app.kubernetes.io/part-of: rutas-norte
environment: dev
spec:
containers:
- name: monitor
image: busybox:1.36
command: ["sh", "-c", "while true; do sleep 30; done"]- To identify the culprit behind an unexpected deletion:
# Which controllers have deleted pods recently in the namespace
kubectl get events -n rutas-norte-dev --sort-by=.lastTimestamp \
--field-selector reason=SuccessfulDelete
# Who currently owns each pod: reveals improper adoptions
kubectl get pods -n rutas-norte-dev \
-o custom-columns=POD:.metadata.name,OWNER:.metadata.ownerReferences[0].name
# Which selectors are declared and whether they overlap
kubectl get rs -n rutas-norte-dev -o wideConclusion
You now have the course's first controller fully taken apart. You know that a controller is an infinite, idempotent, level-driven loop that only talks to the kube-apiserver, and you have seen how the ReplicaSet controller makes it concrete: it reads spec.replicas, counts the pods that match spec.selector and creates or deletes until the numbers add up, using spec.template as the mould. You know the manifest's Russian-doll structure and the non-negotiable rule that the template's labels must satisfy the selector.
You have verified self-healing first-hand: you deleted a web-store pod and it came back in two seconds; you deleted them all and all three came back. That is, literally, the problem that cost Rutas Norte four and a half hours of sales one night in March, solved. You understand that the owner-child relationship is materialised in ownerReferences, that the garbage collector uses it for cascading deletion, and that the three propagation policies —Background, Foreground and Orphan— give you fine control, including the manoeuvre of replacing a controller without interrupting the service. You have seen the friendly face of orphan adoption and its dangerous one: an overly broad selector turns a ReplicaSet into a destroyer of other people's pods. And you know how to scale with kubectl scale, without forgetting that the change must always end up in Git.
But you have also discovered its limit: changing the template's image does not update the existing pods. A ReplicaSet maintains, it does not roll out. It does not know how to change version, or go back, or keep a history. All of that comes from the layer above, which is exactly the next lesson: Deployments. There we will finally turn the bare web-store pod into a 3-replica Deployment, we will create the bookings-api one, and we will watch the very ReplicaSet we have just learned to read appear under the bonnet, managed automatically.
Kubernetes Course
Module 1: Introduction to Kubernetes
- What Is Kubernetes?
- Kubernetes Architecture
- Key Concepts and Terminology
- Setting Up a Kubernetes Cluster
- The Kubernetes CLI: kubectl
- Objects, YAML Manifests and the Declarative Model
- The Course Project: the Rutas Norte Platform
Module 2: Core Kubernetes Components
- Pods
- ReplicaSets
- Deployments
- Updates, Rollbacks and Deployment Strategies
- Services
- Namespaces
- Labels, Selectors and Annotations
Module 3: Configuration and Secret Management
- ConfigMaps
- Secrets
- Environment Variables
- Resource Quotas and Limits
- LimitRanges and Quality of Service (QoS) Classes
- ServiceAccounts and API Access from Pods
Module 4: Networking in Kubernetes
- Cluster Networking
- Service Types
- Internal DNS and Service Discovery
- Ingress Controllers
- TLS and Certificate Management with cert-manager
- Network Policies
Module 5: Storage in Kubernetes
- Volumes
- Persistent Volumes
- Persistent Volume Claims
- Storage Classes
- Dynamic Provisioning, Expansion and Snapshots
- Backup and Restore of Persistent Data
Module 6: Advanced Kubernetes Concepts
- StatefulSets
- DaemonSets
- Jobs and CronJobs
- Init Containers, Sidecars and Multi-Container Patterns
- Scheduling: Affinity, Taints and Tolerations
- Custom Resource Definitions (CRDs)
- Operators and the Controller Pattern
Module 7: Monitoring and Logging
- Health Checks and Probes
- Metrics Server and kubectl top
- Monitoring with Prometheus
- Visualization and Alerting with Grafana and Alertmanager
- Centralized Logging with Elasticsearch, Fluentd and Kibana (EFK)
- Application Debugging and Cluster Events
Module 8: Kubernetes Security
- Role-Based Access Control (RBAC)
- Security Contexts and Container Hardening
- Pod Security Policies and Pod Security Standards
- Network Security
- Image Security
- Auditing, Scanning and Vulnerability Management
Module 9: Scaling and Performance
- Horizontal Pod Autoscaling
- Vertical Pod Autoscaling
- Cluster Autoscaling
- Event-Driven and Custom-Metric Scaling with KEDA
- High Availability: PodDisruptionBudgets and Topology
- Performance Tuning
Module 10: Kubernetes Ecosystem and Tooling
- Minikube and Local Environments with kind
- Kubeadm
- Helm
- Kustomize
- GitOps with Argo CD and Flux
- Managed Kubernetes: EKS, AKS and GKE
Module 11: Case Studies and Real-World Applications
- Deploying a Web Application
- Running Stateful Applications
- CI/CD with Kubernetes
- Deployment Strategies: Blue-Green and Canary
- Multi-Cluster Management
- Production Operations: Incidents, Runbooks and Costs
