The previous lesson ended by pointing out the hidden assumption behind everything we have built so far: we have taken it for granted that, when the HorizontalPodAutoscaler asks for thirty replicas of bookings-api, there is somewhere to put them. There is not. The Rutas Norte cluster has four nodes, and thirty bookings-api pods plus fifteen web-store plus twenty notifications-worker do not fit in them, not by a long way.
That situation has a name and a very concrete appearance: pods in Pending, with a FailedScheduling event saying Insufficient cpu. We already know how to read it from 06-05 and 07-06. What we do not yet know is what to do about it automatically.
This is the third and final level of scaling. The HPA changes the number of pods. The VPA changes the size of each pod. Cluster autoscaling changes the number of nodes. Without it, the other two have a hard ceiling they cannot break through.
We are going to look at the Cluster Autoscaler in detail: how it decides to scale up, the delicate business of removing nodes (which is where all the real problems live), how long a node genuinely takes to become available, the over-provisioning pattern that compensates for those times, and Karpenter as the modern alternative. We will finish with Rutas Norte's concrete capacity plan for the May bank-holiday weekend, with its numbers and its cost.
Contents
- The hard ceiling: when the HPA asks and there is no room
- The exact symptom:
PendingandFailedScheduling - What the Cluster Autoscaler is and where it lives
- How it decides to scale up: simulation and expanders
- Node groups and their relationship with the provider
- Removing nodes: the delicate part
- The reasons a node cannot be drained
- Diagnosis: the status ConfigMap and the events
- The real timings: how long a node genuinely takes
- Over-provisioning with filler pods
- Karpenter: the modern alternative
- How the three autoscalers fit together
- What can be practised on minikube
- The Rutas Norte capacity plan and its cost
- Common Mistakes and Tips
- Exercises
- Conclusion
- The hard ceiling: when the HPA asks and there is no room
Let's set the scene with real numbers. The Rutas Norte production cluster:
| Amount | |
|---|---|
| Nodes | 4 |
| CPU per node | 4 cores |
| Memory per node | 16 GiB |
| Total raw CPU | 16 cores |
| Reserved for kubelet and the system | ~0.5 cores and 1.5 GiB per node |
| Total allocatable CPU | ~14 cores |
| Total allocatable memory | ~58 GiB |
And the idle consumption, with every component at minReplicas:
| Component | Replicas | requests.cpu |
Total CPU |
|---|---|---|---|
bookings-api (api + sidecar) |
4 | 520m | 2.08 |
web-store |
3 | 200m | 0.60 |
notifications-worker |
2 | 320m | 0.64 |
bookings-postgres |
1 | 2000m | 2.00 |
redis-cache |
1 | 300m | 0.30 |
| DaemonSets (Fluentd, Falco, CNI) | 4×3 | ~150m | 1.80 |
| Ingress controller | 2 | 200m | 0.40 |
| Total at idle | 7.82 cores |
That leaves about 6.2 free cores. Now the May bank holiday arrives:
The bookings-api HPA wants 30 replicas.
Current: 4. New: 26. Each one: 520m.
CPU needed: 26 x 520m = 13.52 cores.
CPU available: 6.2 cores.
CPU missing: 7.32 cores.
Replicas that fit: 6.2 / 0.52 = 11.9 -> 11 new replicas.
Replicas that do NOT fit: 15.Fifteen bookings-api pods will be left in Pending. The HPA will have done its job: it asked for 30 replicas and the Deployment created them. The ReplicaSet has them registered. But the scheduler cannot find anywhere to put them, and there they stay.
And here is the worst of it: the HPA does not know. In kubectl get hpa you will see REPLICAS 30. In Grafana, the replicas panel will say 30. The platform will have 15 pods ready and 15 ghosts. And since the HPA computes the CPU average only across the pods that do exist and are ready, it will see them at 90% and will ask for... nothing more, because it is already at maxReplicas.
The result is a platform that falls over under the illusion of being scaled.
- The exact symptom:
Pending and FailedScheduling
Pending and FailedSchedulingLet's learn to recognise it precisely, because it is the trigger for everything else.
NAME READY STATUS RESTARTS AGE
bookings-api-7c9d4f8b6d-2mk8p 2/2 Running 0 4h
bookings-api-7c9d4f8b6d-5xqzn 2/2 Running 0 4h
...
bookings-api-7c9d4f8b6d-qw8vz 0/2 Pending 0 92s
bookings-api-7c9d4f8b6d-rt4nk 0/2 Pending 0 92s
bookings-api-7c9d4f8b6d-sv7mx 0/2 Pending 0 91sPending with 0/2 and no restarts: the pod exists in the API but no node has accepted it. There are no containers running because there is nowhere to run them.
Name: bookings-api-7c9d4f8b6d-qw8vz
Namespace: rutas-norte-pro
Priority: 0
Node: <none>
Status: Pending
...
Events:
Type Reason Age From Message
---- ------ ---- ---- -------
Warning FailedScheduling 95s default-scheduler 0/4 nodes are available:
4 Insufficient cpu. preemption: 0/4 nodes are available:
4 No preemption victims found for incoming pod.
Normal NotTriggerScaleUp 90s cluster-autoscaler pod didn't trigger scale-up:
1 max node group size reachedTwo golden lines:
Node: <none> — confirmation that it is not assigned to any node.
0/4 nodes are available: 4 Insufficient cpu — the scheduler evaluated the four nodes and all four failed on insufficient CPU. The message always has the shape X/Y nodes are available: <reasons>, and the reasons are grouped by cause. The ones you will meet in practice:
| Message | Meaning | Does the CA fix it? |
|---|---|---|
Insufficient cpu |
There is not enough allocatable CPU | Yes |
Insufficient memory |
There is not enough allocatable memory | Yes |
node(s) had untolerated taint {...} |
Taints without a toleration (06-05) | Depends on the node group |
node(s) didn't match Pod's node affinity/selector |
Node affinity not satisfied (06-05) | Only if there is a group that satisfies it |
node(s) didn't match pod topology spread constraints |
Topology spread (09-05) | Sometimes |
node(s) had volume node affinity conflict |
The PV is in another zone (module 5) | No |
pod has unbound immediate PersistentVolumeClaims |
The PVC has not been provisioned | No |
Insufficient nvidia.com/gpu |
Extended resource exhausted | Yes, if there is a GPU group |
The NotTriggerScaleUp line is the Cluster Autoscaler's, and in this example it says max node group size reached: the CA is installed, it has seen the pending pod, and it can do nothing because the node group is already at its maximum size. We will come back to this message in section 8, because it is the main diagnostic tool.
A very useful command to see every pending pod at a glance:
And to see the real occupancy of the nodes:
Allocated resources:
(Total limits may be over 100 percent, i.e., overcommitted.)
Resource Requests Limits
-------- -------- ------
cpu 3410m (97%) 6200m (177%)
memory 9856Mi (67%) 14Gi (98%)Note that the column that matters is Requests, not Limits. The scheduler places pods according to requests. A node at 97% of requests is full for scheduling purposes, even if real consumption is 30%. It is the lesson of module 3, now with capacity consequences.
- What the Cluster Autoscaler is and where it lives
The Cluster Autoscaler (CA) is a component that adjusts the number of nodes in the cluster. Like the VPA, it lives in the kubernetes/autoscaler repository and is not part of the Kubernetes core.
How it works, in one sentence: it watches the pods that cannot be scheduled and adds nodes; it watches the underutilised nodes and removes them.
Where it runs
It runs as a Deployment inside the cluster itself, normally in kube-system, with a ServiceAccount that holds broad permissions (read pods and nodes, create events, evict pods) and cloud-provider credentials to manipulate the node groups.
On managed Kubernetes (10-06) the provider installs and configures it for you: on EKS, AKS and GKE you switch autoscaling on with a toggle and the provider takes care of the rest. That is the most common scenario, and it is why you should understand its behaviour even if you never install it by hand.
What it needs in order to work
| Requirement | Why |
|---|---|
| An API to create and destroy nodes | The CA does not boot machines: it asks the provider to do it |
| Node groups with a minimum and maximum size | That is the unit it operates on |
| Homogeneous nodes inside each group | It simulates using one template node per group |
That every pod declares requests |
Without requests it cannot simulate whether a pod would fit |
That third requirement matters more than it looks. The CA assumes that every node in a group is identical. When it simulates whether a pod would fit on a new node from group A, it uses group A's template. If the group has heterogeneous machines, the simulation lies and the CA makes the wrong decisions. Hence the universal recommendation: one machine type per node group.
And the fourth requirement connects with the whole of module 3: a pod without requests is invisible to the scheduler for capacity purposes, and to the CA as well. Here, once again, requests is the foundation of everything.
- How it decides to scale up: simulation and expanders
The CA loop runs every 10 seconds by default (--scan-interval).
The decision process
- It lists the unschedulable pods. The ones that have been
PendingwithFailedSchedulingfor at least--max-pod-provisioning-time. - It groups equivalent pods. Pods from the same controller with the same requirements are treated as one group, so as not to simulate the same thing twenty times.
- For each node group, it simulates. "If I add a node from group A, how many of these pending pods would fit?" The simulation uses the real scheduler, with all its predicates: resources, taints, affinities, topology spread, host ports.
- It discards the groups that do not help. If adding a node from group B would not allow a single pod to be scheduled (because it has a taint the pod does not tolerate, for example), that group is discarded.
- It chooses among the viable groups using the
expander. - It computes how many nodes are needed and asks the provider to grow the group.
An important detail of step 6: the CA can add several nodes at once. If there are 15 pending pods of 520m each and 6 fit on a 4-core node, it will ask for 3 nodes in one go, not one at a time. This is bounded by --max-nodes-total and by the group's maximum.
The expanders
When several node groups would do the job, the expander decides which one.
| Expander | Criterion | When to use it |
|---|---|---|
random |
Picks at random among the viable ones | The default. Only valid if every group is equivalent |
most-pods |
The group that would let it schedule the most pending pods | When the priority is to unblock fast |
least-waste |
The group that would leave the least idle CPU and memory after placing the pods | The most sensible in general. It optimises the fit |
price |
The cheapest group (requires provider support) | When cost rules and there are varied machine types |
priority |
According to a list of priorities in a ConfigMap | Explicit control: try spot first, then on-demand |
A concrete example of least-waste. There are 3 pending pods needing 500m and 512Mi each (total: 1.5 cores and 1.5 GiB), and two groups available:
| Group | Node type | Allocatable | After placing the 3 pods | Waste |
|---|---|---|---|---|
| A | 2 cores, 8 GiB | 1.8 cores / 7 GiB | 0.3 cores and 5.5 GiB free | CPU: 17%, RAM: 79% |
| B | 8 cores, 32 GiB | 7.7 cores / 30 GiB | 6.2 cores and 28.5 GiB free | CPU: 81%, RAM: 95% |
least-waste picks group A: it fits better and you do not pay for 6 idle cores. most-pods would pick B (many more future pods would fit). price would pick A if it is cheaper.
Configuring the expander:
# Fragment of the cluster-autoscaler Deployment
spec:
containers:
- name: cluster-autoscaler
image: registry.k8s.io/autoscaling/cluster-autoscaler:v1.30.0
command:
- ./cluster-autoscaler
- --cloud-provider=aws
- --namespace=kube-system
- --expander=least-waste
- --scan-interval=10s
- --balance-similar-node-groups # Spreads across equivalent groups (zones)
- --skip-nodes-with-local-storage=false
- --skip-nodes-with-system-pods=false
- --node-group-auto-discovery=asg:tag=k8s.io/cluster-autoscaler/enabled,k8s.io/cluster-autoscaler/rutas-norteThe --balance-similar-node-groups option deserves a mention because it solves a real availability problem: if you have one node group per availability zone, this flag makes the CA spread new nodes evenly across zones instead of piling them into one. Without it, you can end up with twelve nodes in zone A and none in zone B, and a zone outage leaves you with no platform. It connects directly with what we will see in 09-05.
The priority expander: try the cheap machines first
A very profitable pattern in the cloud. Spot (interruptible) instances cost a fraction of the price, but the provider can reclaim them with two minutes' notice.
apiVersion: v1
kind: ConfigMap
metadata:
name: cluster-autoscaler-priority-expander
namespace: kube-system
data:
priorities: |-
100:
- rutas-norte-spot-.* # First choice: spot, much cheaper
50:
- rutas-norte-ondemand-.* # Fallback: on-demand, always availableThe CA tries the priority-100 group first. If the provider has no spot capacity available, it falls back to the priority-50 one. For notifications-worker, which tolerates interruptions without trouble, it is the obvious choice. For bookings-postgres, never.
- Node groups and their relationship with the provider
The CA does not create machines: it manipulates the provider's node groups. Every provider has its own name for the same thing:
| Provider | Name of the concept |
|---|---|
| AWS | Auto Scaling Group (ASG) |
| Azure | Virtual Machine Scale Set (VMSS) / Agent Pool |
| Google Cloud | Managed Instance Group (MIG) / Node Pool |
| On-premise / kubeadm | Requires a custom provider or Cluster API |
A node group has a minimum size, a maximum and a current one. The CA only changes the current one, within those bounds. Everything else (the machine type, the image, the initial taints and labels) is defined by the group.
The Rutas Norte groups:
rutas-norte-general-a min=1 max=8 4 cores / 16 GiB zone A
rutas-norte-general-b min=1 max=8 4 cores / 16 GiB zone B
rutas-norte-general-c min=1 max=8 4 cores / 16 GiB zone C
rutas-norte-data min=2 max=3 8 cores / 32 GiB taint: role=data:NoScheduleThree design decisions worth understanding:
One group per zone. Necessary so that --balance-similar-node-groups can spread across zones and so that zone affinity works. With a single multi-zone group, the CA cannot guarantee which zone the new node will appear in, and that breaks the topology spread of 09-05.
A dedicated data group with a taint. bookings-postgres and redis-cache have tolerations for role=data (06-05) and node affinity towards that group. That way the databases live on big machines with fast disks and do not share a node with the avalanche of bookings-api pods. The taint stops the application pods from landing there.
The data group has max=3, which is very low. That is deliberate: we do not want the CA booting big, expensive machines by mistake. bookings-postgres does not scale horizontally (09-01), so that group almost never needs to grow.
Labels and taints on the new nodes
When the CA simulates whether a pod would fit on a node from group A, it uses a template. If the real node, on boot, receives labels or taints the template did not know about, the simulation lies.
A typical and very frustrating case: a pod with nodeSelector: disk=ssd is Pending. The CA simulates with the group template, which does not carry that label, concludes that the pod would not fit even on a new node, and does not scale up. But the real nodes in the group do receive that label on boot, through a start-up script.
The fix is to declare the labels and taints in the provider's own node-group definition (on AWS, with k8s.io/cluster-autoscaler/node-template/label/... tags on the ASG), so that the CA knows about them before booting the node.
- Removing nodes: the delicate part
Scaling up is easy: there are pending pods, you add nodes. Scaling down is where all the problems are, because removing a node means evicting the pods it holds and trusting that they will reappear somewhere else without breaking anything.
The algorithm
Every 10 seconds, for each node:
- Is it below the utilisation threshold? By default,
--scale-down-utilization-threshold=0.5: the sum of its pods'requestsis less than 50% of its allocatable capacity. - Has it been that way long enough?
--scale-down-unneeded-time=10m: ten consecutive minutes below the threshold. - Can it be drained? It simulates whether all its pods would fit on other existing nodes. If they do not fit, it is left alone.
- Is there any pod that prevents eviction? The list in section 7.
- If everything passes: it cordons the node, evicts its pods respecting the PDBs, waits, and asks the provider to delete it.
Main parameters:
| Parameter | Default | What it controls |
|---|---|---|
--scale-down-enabled |
true |
Master switch for scale-down |
--scale-down-utilization-threshold |
0.5 |
Underutilisation threshold |
--scale-down-unneeded-time |
10m |
Consecutive time before acting |
--scale-down-delay-after-add |
10m |
Wait after having added a node |
--scale-down-delay-after-delete |
0s |
Wait after having deleted a node |
--scale-down-delay-after-failure |
3m |
Wait after a scale-down failure |
--max-graceful-termination-sec |
600 |
Maximum wait for a pod to terminate |
--max-empty-bulk-delete |
10 |
Empty nodes deletable at once |
--scale-down-delay-after-add is the key parameter for avoiding node flapping. Without it, the CA could add a node, see that the cluster is now underutilised and remove it immediately. Ten minutes of grace break that cycle.
About the utilisation threshold, an important calculation:
Default threshold: 50%.
Node with 4 allocatable cores holding these pods:
- 2 bookings-api pods: 2 x 520m = 1040m
- 1 web-store pod: 200m
- DaemonSets (fluentd, falco, cni): 450m
Total requests: 1690m out of 3500m allocatable = 48.3%
48.3% < 50% -> scale-down candidate.
The CA simulates: do those 4 pods fit on the other nodes?
If yes -> it evicts them and deletes the node.
If no -> it leaves it alone.Note that DaemonSets count towards utilisation but do not prevent deletion (they are ignored in step 4, because they disappear with the node). With many DaemonSets, the baseline utilisation of an empty node is already high, and that makes the CA scale down less than you would expect. In clusters with many agents (logging, security, service mesh, monitoring), it is common to have to lower the threshold to 30-40% for scale-down to work at all.
The special case of empty nodes
A node that only holds DaemonSets is considered empty and is deleted via a fast path, without the full simulation. That is why, after a traffic peak, the nodes that end up completely free disappear fairly quickly, while the ones with a stray pod take far longer.
This explains a phenomenon that puzzles a lot of people: after the May bank holiday, eight nodes go away in twenty minutes and two stay for hours with a single pod each. The real fix is not to fiddle with the CA: it is topology spread and consolidation (Karpenter, section 11).
- The reasons a node cannot be drained
This is the list to know by heart, because it explains 90% of the cases of "I have empty nodes the autoscaler will not remove and I am paying for them".
| # | Reason | Detail | How to fix it |
|---|---|---|---|
| 1 | Pods without a controller | A pod created by hand (with no Deployment, ReplicaSet, Job or StatefulSet) cannot be recreated elsewhere: if you evict it, it is gone for good | Always use a controller. Never kubectl run without --restart in production |
| 2 | Pods with local storage | emptyDir, hostPath or local volumes: the data lives on that node and would be lost |
--skip-nodes-with-local-storage=false if you accept the loss; or migrate to a network PV |
| 3 | A restrictive PodDisruptionBudget | The PDB does not allow evicting that pod without dropping below the minimum (09-05) | Review the PDB; make sure there are enough replicas |
| 4 | System pods in kube-system |
Critical components with no PDB | --skip-nodes-with-system-pods=false (carefully), or give them a PDB |
| 5 | The safe-to-evict: "false" annotation |
An explicit "do not move me" marker | Remove the annotation if it no longer applies |
| 6 | Pods that do not fit on any other node | The simulation fails: the pod is too big or has constraints | Add capacity or relax the constraints |
| 7 | A node with the scale-down-disabled: "true" annotation |
Explicit exclusion of the node | Remove the annotation |
The safe-to-evict annotation
It is the most direct control you have:
apiVersion: v1
kind: Pod
metadata:
annotations:
# This pod must NOT be evicted by the Cluster Autoscaler.
# The node hosting it will never be removed while it is here.
cluster-autoscaler.kubernetes.io/safe-to-evict: "false"When "false" makes sense:
- A long batch job that would lose hours of compute if restarted.
occupancy-reportsis a candidate: if it takes 40 minutes and the CA kills it at minute 35, the night's work is lost. - A database migration in flight.
- A pod with irreproducible local state.
And the opposite value, "true", serves to unblock situations:
metadata:
annotations:
# This pod CAN be evicted even though it uses emptyDir.
# We know the contents of its emptyDir are a regenerable cache.
cluster-autoscaler.kubernetes.io/safe-to-evict: "true"It is the clean way to resolve reason 2 without changing the CA's global flag. In Rutas Norte, web-store uses an emptyDir for the nginx cache: marking it as safe-to-evict: "true" lets the CA remove its nodes with no trouble, because that cache regenerates itself.
The PDB as a blocker (a preview of 09-05)
The most frequent case and the most dangerous. A PDB like this:
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: bookings-postgres
namespace: rutas-norte-pro
spec:
minAvailable: 1
selector:
matchLabels:
app: bookings-postgresWith a single replica of bookings-postgres and minAvailable: 1, that pod can never be evicted: doing so would leave 0 available. The node hosting it is pinned permanently. The CA will try, will fail, and will record it in its events.
In this particular case the block is desirable (we do not want the CA moving the database on its own), but the same pattern applied by mistake to an ordinary service produces zombie nodes that nobody can explain. We develop this fully in 09-05.
- Diagnosis: the status ConfigMap and the events
The CA publishes its internal state in a ConfigMap. It is the first place to look.
apiVersion: v1
kind: ConfigMap
metadata:
name: cluster-autoscaler-status
namespace: kube-system
data:
status: |
Cluster-autoscaler status at 2026-05-01 10:14:32:
Cluster-wide:
Health: Healthy (ready=7 unready=0 notStarted=1 registered=8 longNotStarted=0)
LastProbeTime: 2026-05-01 10:14:31
ScaleUp: InProgress (ready=7 registered=8)
LastProbeTime: 2026-05-01 10:14:31
ScaleDown: NoCandidates (candidates=0)
LastProbeTime: 2026-05-01 10:14:31
NodeGroups:
Name: rutas-norte-general-a
Health: Healthy (ready=3 unready=0 notStarted=1 registered=4
cloudProviderTarget=4 (minSize=1, maxSize=8))
ScaleUp: InProgress (ready=3 registered=4)
ScaleDown: NoCandidates (candidates=0)
Name: rutas-norte-general-b
Health: Healthy (ready=2 unready=0 notStarted=0 registered=2
cloudProviderTarget=2 (minSize=1, maxSize=8))
ScaleUp: NoActivity
ScaleDown: NoCandidates (candidates=0)
Name: rutas-norte-data
Health: Healthy (ready=2 unready=0 notStarted=0 registered=2
cloudProviderTarget=2 (minSize=2, maxSize=3))
ScaleUp: NoActivity
ScaleDown: NoCandidates (candidates=0)How to read it:
| Field | Meaning |
|---|---|
ready |
Nodes operational and accepting pods |
unready |
Nodes registered but not ready (a problem) |
notStarted |
Nodes requested that are still booting. This is the one to watch during a peak |
registered |
Total known to the cluster |
cloudProviderTarget |
How many the CA has asked the provider for |
ScaleUp: InProgress |
A scale-up is under way |
ScaleDown: NoCandidates |
No node meets the scale-down criteria |
In the output above: group A has 3 ready and 1 booting. The CA has already requested the node; all you have to do is wait. This information is exactly what you need during an incident to know whether the CA is working or stuck.
ScaleDown states you will see:
| State | Meaning |
|---|---|
NoCandidates |
No node is below the threshold |
CandidatesPresent |
There are candidates, waiting for the unneeded-time |
InProgress |
Draining a node right now |
The events: why it is NOT scaling up
When the CA decides not to do anything, it says so in an event on the pending pod:
kubectl get events -n rutas-norte-pro --field-selector reason=NotTriggerScaleUp \
--sort-by=.lastTimestampLAST SEEN TYPE REASON OBJECT MESSAGE
23s Normal NotTriggerScaleUp pod/bookings-api-7c9d4f8b6d-qw8vz pod didn't trigger scale-up:
1 max node group size reached,
1 node(s) had untolerated taint {role: data}This message is a gift: it tells you exactly why each node group was discarded. Here: one group is at its maximum and another has a taint the pod does not tolerate.
Common NotTriggerScaleUp messages and what they mean:
| Message | What it really means | Action |
|---|---|---|
max node group size reached |
The group is at its maxSize |
Raise the group's maximum |
node(s) had untolerated taint {...} |
The group's nodes carry a taint the pod does not tolerate | Add a toleration or create another group |
node(s) didn't match Pod's node affinity |
The nodeSelector/affinity does not match the template |
Review the group's labels |
Insufficient cpu (in the CA's message) |
Not even a new node from the group would have enough CPU for that pod | The pod is too big. Reduce its requests or use a larger group |
in backoff after failed scale-up |
A previous attempt failed (no capacity at the provider, quota exhausted) | Look at the CA logs and the provider quota |
pod has unbound immediate PersistentVolumeClaims |
It is a storage problem, not a capacity one | Review the StorageClass (module 5) |
And for blocked scale-down:
I0501 10:22:14.882 scale_down.go: Node ip-10-0-3-47 is not suitable for removal:
pod redis-cache-0 is not replicated and has local storage
I0501 10:22:14.883 scale_down.go: Node ip-10-0-2-19 is not suitable for removal:
pdb-blocked: not enough pod disruption budget to move bookings-postgres-0
I0501 10:22:14.884 scale_down.go: Node ip-10-0-1-88 is not suitable for removal:
cluster-autoscaler.kubernetes.io/safe-to-evict annotation set to falseThe CA logs are the definitive source when the ConfigMap and the events are not enough. Keep them in EFK (07-05): they are indispensable for the postmortem of a capacity incident.
- The real timings: how long a node genuinely takes
Here is the figure that changes the strategy, and the one almost nobody mentions.
Breaking the time down
gantt
title Time from Pending pod to Running pod on a new node
dateFormat s
axisFormat %S s
section Detection
Pod goes to Pending :a1, 0, 5s
The CA detects it (scan) :a2, after a1, 10s
section Provisioning
Request to the provider :b1, after a2, 5s
Machine boot :b2, after b1, 45s
Operating system boot :b3, after b2, 20s
section Joining the cluster
kubelet starts and registers :c1, after b3, 15s
CNI and DaemonSets ready :c2, after c1, 30s
Node becomes Ready :c3, after c2, 5s
section Pod
Scheduling :d1, after c3, 2s
Image pull :d2, after d1, 40s
Container start-up :d3, after d2, 25s
Readiness probe OK :d4, after d3, 10s
Adding it up: between 3 and 4 minutes in the good case. In the bad case (a large uncached image, a saturated region, heavy DaemonSets) it can exceed 6 minutes.
Typical breakdown in a table:
| Phase | Typical time | What makes it worse |
|---|---|---|
| Detecting the pending pod | 10-15 s | A high --scan-interval |
| Request to the provider | 5-10 s | Retries because of quota |
| Machine boot | 30-60 s | Machine type, saturated region |
| Operating-system boot | 15-30 s | An unoptimised image |
| kubelet registration | 10-20 s | Complex configuration, slow bootstrap |
| CNI + DaemonSets ready | 20-60 s | Many heavy DaemonSets |
| Pulling the pod's image | 20-120 s | A large image, a distant registry |
| Container start-up + readiness | 20-40 s | An application that is slow to start |
| Total | 2-6 min |
Why this does not save you from a sudden spike
Let's go back to the May bank holiday. The sale opens at 10:00. The real traffic profile:
10:00:00 The sale opens. Traffic goes from 200 rps to 2400 rps in 40 seconds.
10:00:15 The HPA detects CPU at 280%. It asks for 12 replicas.
10:00:20 The Deployment creates 8 pods. 4 fit on the existing nodes.
4 are left Pending.
10:00:30 The CA detects the pending pods and asks for 2 nodes.
10:00:45 The HPA evaluates again: still saturated. It asks for 24 replicas.
More Pending pods.
10:01:00 The CA asks for 3 more nodes.
10:03:30 The first node arrives. 6 pods are scheduled.
10:04:10 The first pods on the new node become Ready.
10:05:00 The other nodes arrive. The platform reaches capacity.
TOTAL DEGRADATION TIME: almost FIVE MINUTES.Five minutes of degraded latency and errors at the most commercially valuable moment of the year. Thousands of users abandoning their purchase. The Cluster Autoscaler, however well configured, cannot solve a spike that rises in forty seconds, because the physics of booting a virtual machine does not allow it.
Two strategies come out of this, and you need both:
- Over-provisioning: having capacity already booted and waiting (section 10).
- Anticipatory scaling: scaling before the peak using a time-based trigger rather than a reactive one. That is KEDA's cron trigger, which we will see in 09-04.
- Over-provisioning with filler pods
The pattern that solves the problem of the previous section, and which makes brilliant use of the PriorityClass and preemption from 06-05.
The idea
We deploy pods that do nothing — they sleep — but that reserve CPU and memory, with a negative priority. The effects:
- The CA counts those
requestsas occupancy, so it keeps nodes booted to host them. - When a real pod arrives (priority 0 or higher) and there is no room, the scheduler instantly evicts a filler pod and places the real one in its slot. Preemption takes seconds, not minutes.
- The evicted filler pods go to
Pending, which triggers the CA to boot new nodes... which will refill the cushion for next time.
It is a self-regenerating capacity cushion. You buy idle capacity in exchange for response time.
sequenceDiagram
participant HPA
participant Sched as Scheduler
participant Filler as Filler pods<br/>(priority -10)
participant CA as Cluster Autoscaler
Note over Filler: Normal state: 3 filler pods<br/>occupying 3 cores across the nodes
HPA->>Sched: Create 10 bookings-api pods (priority 0)
Sched->>Sched: There is no free CPU
Sched->>Filler: PREEMPTION: evict 3 filler pods
Note over Sched: Room appears in 2-5 seconds
Sched->>Sched: Schedule the bookings-api pods
Filler->>CA: 3 filler pods in Pending
CA->>CA: Request new nodes (3-4 min)
Note over Filler: The cushion regenerates on its own
Complete implementation
Step 1: the negative PriorityClass.
# k8s/base/priorityclass-capacity-filler.yaml
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
name: capacity-filler
# KEY POINT: a NEGATIVE value. Any normal pod (priority 0 by default) has
# more priority than these, so it will evict them without hesitation.
value: -10
# It is not the default class: only pods that ask for it explicitly use it.
globalDefault: false
# The default preemptionPolicy is PreemptLowerPriority, but these pods must
# not evict ANYONE: they are the last in the queue.
preemptionPolicy: Never
description: >
Filler pods that reserve capacity to absorb traffic spikes.
They are evicted instantly by any real pod. See 09-03.Two important details:
value: -10guarantees that any pod without apriorityClassName(priority 0) outranks them.preemptionPolicy: Nevermeans these pods, oncePending, will not try to evict others. It would be absurd for the cushion to throw out a real pod.
Step 2: the filler Deployment.
# k8s/environments/pro/deployment-capacity-filler.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: capacity-filler
namespace: rutas-norte-pro
labels:
app: capacity-filler
app.kubernetes.io/part-of: rutas-norte
spec:
# 6 filler pods. Each one reserves 1 core and 1 GiB.
# Total cushion: 6 cores and 6 GiB, roughly two nodes' worth.
# Enough to absorb 11 bookings-api replicas instantly.
replicas: 6
selector:
matchLabels:
app: capacity-filler
template:
metadata:
labels:
app: capacity-filler
app.kubernetes.io/part-of: rutas-norte
annotations:
# So the CA can freely remove nodes that only hold filler.
cluster-autoscaler.kubernetes.io/safe-to-evict: "true"
spec:
priorityClassName: capacity-filler
# Immediate termination: there is nothing to save, we do not wait a second.
# This is what makes preemption almost instantaneous.
terminationGracePeriodSeconds: 0
# Spread the cushion across nodes: 6 free cores are useless if they all
# sit on the same node when the new pods need to spread out.
topologySpreadConstraints:
- maxSkew: 1
topologyKey: kubernetes.io/hostname
whenUnsatisfiable: ScheduleAnyway
labelSelector:
matchLabels:
app: capacity-filler
containers:
- name: pause
# Minimal official image: a process that only sleeps. ~700 KB.
image: registry.k8s.io/pause:3.9
resources:
requests:
cpu: "1"
memory: 1Gi
limits:
cpu: "1"
memory: 1Gi
# It consumes NOTHING in practice: it only reserves the slot in the
# scheduler's eyes. Real CPU used is 0m.The pause image is the canonical choice: it is the same one Kubernetes uses internally for each pod's infrastructure container, it weighs less than a megabyte, it is already cached on every node, and its only job is to sleep.
Step 3: checking that it works.
# Initial state: 6 filler pods spread out
kubectl get pods -n rutas-norte-pro -l app=capacity-filler -o wideNAME READY STATUS NODE
capacity-filler-5d8f7c9b4-2xkjp 1/1 Running ip-10-0-1-42
capacity-filler-5d8f7c9b4-4mnqz 1/1 Running ip-10-0-1-88
capacity-filler-5d8f7c9b4-7wrtv 1/1 Running ip-10-0-2-19
capacity-filler-5d8f7c9b4-9hbcx 1/1 Running ip-10-0-2-73
capacity-filler-5d8f7c9b4-kp3ln 1/1 Running ip-10-0-3-47
capacity-filler-5d8f7c9b4-tv6ws 1/1 Running ip-10-0-3-91Now we force a spike:
kubectl scale deployment bookings-api -n rutas-norte-pro --replicas=20
kubectl get pods -n rutas-norte-pro -l app=capacity-filler --watchNAME READY STATUS AGE
capacity-filler-5d8f7c9b4-2xkjp 1/1 Terminating 4h
capacity-filler-5d8f7c9b4-4mnqz 1/1 Terminating 4h
capacity-filler-5d8f7c9b4-7wrtv 0/1 Pending 3s
capacity-filler-5d8f7c9b4-9hbcx 0/1 Pending 3sAnd in the events of the bookings-api pods:
Events:
Type Reason Age From Message
---- ------ ---- ---- -------
Normal Preempted 4s default-scheduler Preempted by pod
capacity-filler-5d8f7c9b4-2xkjp
on node ip-10-0-1-42
Normal Scheduled 3s default-scheduler Successfully assigned
rutas-norte-pro/bookings-api-... to ip-10-0-1-42Three seconds from the pod being created to being scheduled, instead of three and a half minutes. That is the difference between absorbing the opening of the sale and falling over.
The cost
Nothing is free:
Cushion: 6 cores and 6 GiB reserved permanently.
Equivalent to ~1.7 nodes of 4 cores.
Estimated cost of a node (4 cores, 16 GiB, on-demand): 95 EUR/month.
Cost of the cushion: 1.7 x 95 = 162 EUR/month = 1,940 EUR/year.
Benefit: instantly absorbing 11 bookings-api replicas.Is it worth it? At Rutas Norte the answer is clear: five minutes of outage at the opening of the May bank-holiday sale cost far more than €1,940 in unsold tickets. But it is a business decision, not a technical one, and it must be taken with the numbers on the table.
An obvious optimisation: the cushion does not have to be constant. A CronJob (06-03) that scales the filler Deployment from 2 to 6 replicas on Friday afternoons and returns it to 2 on Mondays reduces the annual cost to a fraction:
# k8s/environments/pro/cronjob-adjust-filler.yaml
apiVersion: batch/v1
kind: CronJob
metadata:
name: expand-filler-for-weekend
namespace: rutas-norte-pro
spec:
schedule: "0 16 * * 5" # Friday at 16:00
jobTemplate:
spec:
template:
spec:
serviceAccountName: scaling-operator
restartPolicy: OnFailure
containers:
- name: kubectl
image: bitnami/kubectl:1.30
command:
- kubectl
- scale
- deployment/capacity-filler
- --replicas=6
- -n
- rutas-norte-proAnd its counterpart on Monday at 6:00 with --replicas=2. The scaling-operator ServiceAccount needs a Role with permission over deployments/scale, exactly as we saw in 08-01.
- Karpenter: the modern alternative
The Cluster Autoscaler has a design limitation: it works with predefined node groups. Someone has to decide in advance which machine types exist, create a group for each of them, and maintain them.
Karpenter turns the approach around: it looks at the pending pods, works out which machine would be best for them, and boots it directly, with no groups.
How it works
Instead of node groups, you define constraints on which machine types are acceptable:
# NodePool: the decision space Karpenter is allowed to explore
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
name: rutas-norte-general
spec:
template:
metadata:
labels:
environment: pro
app.kubernetes.io/part-of: rutas-norte
spec:
requirements:
- key: kubernetes.io/arch
operator: In
values: ["amd64"]
- key: karpenter.sh/capacity-type
operator: In
values: ["spot", "on-demand"] # Prefers spot; falls back to on-demand
- key: karpenter.k8s.aws/instance-category
operator: In
values: ["c", "m", "r"] # Compute, general-purpose and memory families
- key: karpenter.k8s.aws/instance-size
operator: NotIn
values: ["nano", "micro", "small"]
- key: topology.kubernetes.io/zone
operator: In
values: ["eu-west-1a", "eu-west-1b", "eu-west-1c"]
nodeClassRef:
group: karpenter.k8s.aws
kind: EC2NodeClass
name: rutas-norte
# CONSOLIDATION: the flagship feature. Karpenter continuously re-evaluates
# whether the current pods would fit on fewer or cheaper nodes, and migrates.
disruption:
consolidationPolicy: WhenEmptyOrUnderutilized
consolidateAfter: 1m
budgets:
- nodes: "20%" # At most, touch 20% of the nodes at a time
limits:
cpu: "200" # Global safety ceiling
memory: 800GiWith this definition, when 15 bookings-api pods turn up pending, Karpenter works out that the best machine to host them is, say, a c6i.4xlarge spot instance in zone B, and boots it. Without anyone ever having created a c6i.4xlarge node group.
Consolidation
This is the most practical difference day to day. Karpenter continuously checks whether the current set of pods would fit on a cheaper node configuration, and if so, it migrates the pods and replaces the nodes.
The scenario it solves: after the May bank holiday, the CA leaves five nodes with one pod each, because none drops below the utilisation threshold in a way that allows draining. Karpenter sees that those five pods would fit on a single node, boots that node, migrates the pods and deletes the five.
It is exactly the problem we mentioned at the end of section 6, solved at the root.
Comparison
| Aspect | Cluster Autoscaler | Karpenter |
|---|---|---|
| Unit of work | Predefined node groups | Individual nodes |
| Choosing the machine type | Fixed per group | Computed from the pending pods |
| Prior configuration | One group per type/zone | One NodePool with constraints |
| Provisioning time | 3-6 min | 1-2 min (talks directly to the EC2 API) |
| Consolidation | No (only utilisation-based scale-down) | Yes, continuous |
| Cost optimisation | Limited (the price expander) |
Native: picks the cheapest type that works |
| Spot handling | Separate groups + the priority expander |
Native, with automatic fallback to on-demand |
| Provider support | Many (AWS, Azure, GCP, OpenStack, Cluster API...) | AWS mature; Azure available; others in development |
| Maturity | Very high, years in production | High on AWS, more recent |
| Operational complexity | Medium (maintaining groups) | Low after the initial setup |
| Predictability | High: you know which machines will appear | Lower: it can pick unexpected types |
When to choose each one:
- Cluster Autoscaler if you are outside AWS, if you have regulatory requirements about which machine types are admissible, if your provider already gives it to you configured and working, or if you value predictability over savings.
- Karpenter if you are on AWS, if cost is a genuine priority, if your workload is heterogeneous (small and large pods living together), and if the reduction in operational work pays off for you.
For Rutas Norte, with four nodes and a homogeneous workload, the CA is more than enough. If the platform grew to fifty nodes with varied workloads, Karpenter would be the obvious decision. We will come back to this kind of decision in 10-06 and 11-06.
An important note the comparison must not hide: Karpenter does not fix the forty-second spike either. One or two minutes is still too long. Over-provisioning is still necessary.
- How the three autoscalers fit together
We now have all three pieces. Let's see the full flow in time order.
flowchart TD
T0["t=0 s<br/>Traffic rises"] --> HPA["t=15 s<br/>HPA: average CPU exceeds the target<br/>Computes desired replicas"]
HPA --> DEP["t=16 s<br/>Deployment creates new pods"]
DEP --> SCHED{"t=17 s<br/>Does the scheduler<br/>find room?"}
SCHED -->|Yes| RUN["t=20-50 s<br/>Pod Running and Ready<br/>END: absorbed"]
SCHED -->|No| PRE{"Are there filler<br/>pods to evict?"}
PRE -->|Yes| PREEMPT["t=20 s<br/>PREEMPTION<br/>Real pod placed in 3 s"]
PREEMPT --> RUN2["t=50 s<br/>Pod Running<br/>END: absorbed with the cushion"]
PREEMPT --> PEND2["The filler pods<br/>go to Pending"]
PRE -->|No| PEND["Pod in Pending<br/>FailedScheduling event"]
PEND2 --> CA
PEND --> CA["t=30 s<br/>Cluster Autoscaler detects<br/>the unschedulable pods"]
CA --> SIM["Simulates: which node group<br/>would let them be scheduled?<br/>Applies the expander"]
SIM --> REQ["t=40 s<br/>Request to the cloud provider"]
REQ --> BOOT["t=40 s to 3 min<br/>Boot, joining the cluster,<br/>CNI and DaemonSets"]
BOOT --> READY["t=3 min<br/>Node Ready"]
READY --> SCHED2["t=3 min 5 s<br/>The pending pods are scheduled"]
SCHED2 --> PULL["t=3 to 4 min<br/>Image pull and start-up"]
PULL --> RUN3["t=4 min<br/>Pod Running<br/>END: absorbed with a new node"]
style RUN fill:#cfc,stroke:#393
style RUN2 fill:#cfc,stroke:#393
style RUN3 fill:#ffd,stroke:#c90
style PEND fill:#fdd,stroke:#c00
The three routes and their timings:
| Route | When it happens | Time until it serves traffic |
|---|---|---|
| Fits on the current nodes | There is free capacity | 20-50 s |
| Preemption of the cushion | It does not fit, but there are filler pods | 30-60 s |
| New node | It does not fit and there is no cushion | 3-6 min |
The strategic conclusion is direct: the over-provisioning cushion is not a luxury, it is what turns a five-minute incident into a thirty-second one.
And there is still a better route, which does not appear in the diagram because it is not reactive: scaling before the traffic arrives. If at 9:30 you already have 25 replicas and eight nodes because a cron trigger ordered it, at 10:00 there is no incident at all. That is 09-04.
A summary of the three levels
| Level | What changes | Trigger | Latency | Component |
|---|---|---|---|---|
| 1 | Number of pods | A CPU/business metric | 15-60 s | HPA (09-01) / KEDA (09-04) |
| 2 | Size of each pod | Consumption history | Hours-days | VPA (09-02) |
| 3 | Number of nodes | Unschedulable pods | 3-6 min | Cluster Autoscaler / Karpenter |
All three are necessary and none replaces another. The HPA without the CA has a hard ceiling. The CA without the HPA never fires. The VPA makes both more accurate, because correct requests are the foundation of both the HPA's calculation and the CA's simulation.
- What can be practised on minikube
Time for honesty: the Cluster Autoscaler cannot really be tested on minikube. There is no cloud provider to boot machines. The rutas-norte profile has the nodes you asked for when you created it, and that is all of them.
What you CAN practise
1. Producing the symptom and learning to read it.
# A cluster with 2 small nodes
minikube start -p rutas-norte --nodes=2 --cpus=2 --memory=4096
# Deploy something that does not fit
kubectl -n rutas-norte-dev create deployment occupier \
--image=registry.k8s.io/pause:3.9 --replicas=10
kubectl -n rutas-norte-dev set resources deployment/occupier \
--requests=cpu=800m,memory=512Mi
kubectl get pods -n rutas-norte-devNAME READY STATUS RESTARTS AGE
occupier-6f8d7c9b4-2mkjp 1/1 Running 0 15s
occupier-6f8d7c9b4-4nqzx 1/1 Running 0 15s
occupier-6f8d7c9b4-7wtrv 0/1 Pending 0 15s
occupier-6f8d7c9b4-9bhcx 0/1 Pending 0 15s
...You will see exactly the FailedScheduling with Insufficient cpu. It is the same message as in production.
2. The over-provisioning pattern with preemption. This works perfectly, because preemption is a function of the scheduler, not of the cloud provider.
# 1. Create the negative PriorityClass
kubectl apply -f k8s/base/priorityclass-capacity-filler.yaml
# 2. Deploy the filler (sized for minikube)
kubectl -n rutas-norte-dev create deployment filler \
--image=registry.k8s.io/pause:3.9 --replicas=2
kubectl -n rutas-norte-dev set resources deployment/filler \
--requests=cpu=500m,memory=256Mi
kubectl -n rutas-norte-dev patch deployment filler \
-p '{"spec":{"template":{"spec":{"priorityClassName":"capacity-filler","terminationGracePeriodSeconds":0}}}}'
# 3. Check that the filler is taking up room
kubectl get pods -n rutas-norte-dev -o wide
# 4. Create a real pod that would not fit without evicting the filler
kubectl -n rutas-norte-dev run important-pod \
--image=registry.k8s.io/pause:3.9 \
--overrides='{"spec":{"containers":[{"name":"important-pod","image":"registry.k8s.io/pause:3.9","resources":{"requests":{"cpu":"900m"}}}]}}'
# 5. Watch the preemption live
kubectl get events -n rutas-norte-dev --sort-by=.lastTimestamp | tail -20LAST SEEN TYPE REASON OBJECT MESSAGE
3s Normal Preempted pod/filler-7d8c9f4b6-2xkjp Preempted by pod
rutas-norte-dev/important-pod on node rutas-norte-m02
2s Normal Scheduled pod/important-pod Successfully assigned to rutas-norte-m02
1s Normal Scheduled pod/filler-7d8c9f4b6-9wqzn FailedSchedulingSeeing the Preempted with your own eyes and measuring that it happens in seconds is the most valuable part of the exercise, and you can practise it without a cloud.
3. Adding and removing nodes by hand, simulating what the CA would do.
# Add a node to the profile (equivalent to what the CA does when scaling up)
minikube -p rutas-norte node add
# See how the Pending pods get scheduled on their own
kubectl get pods -n rutas-norte-dev --watch
# Remove a node (equivalent to scale-down, draining included)
kubectl cordon rutas-norte-m03
kubectl drain rutas-norte-m03 --ignore-daemonsets --delete-emptydir-data
minikube -p rutas-norte node delete rutas-norte-m03That cordon → drain → delete sequence is exactly what the CA does internally. Practising it by hand teaches you more about node removal than reading any documentation, and it is the basis of the maintenance procedure we will see in 09-05.
4. Practising the safe-to-evict annotation and seeing how it blocks a drain.
kubectl annotate pod important-pod -n rutas-norte-dev \
cluster-autoscaler.kubernetes.io/safe-to-evict=false
# Note: this annotation is honoured by the CA, not by kubectl drain. To see a
# drain really blocked you need a PDB, which is what we will see in 09-05.What you CANNOT practise
| Not practicable | Alternative |
|---|---|
| Automatic node scale-up | minikube node add by hand |
| Expanders and group selection | Study the documentation and the logs of a real cluster |
| Automatic node scale-down | cordon + drain + node delete by hand |
The cluster-autoscaler-status ConfigMap |
It does not exist without the CA installed |
| Karpenter | Requires AWS |
To practise the CA for real you need: a managed cluster (10-06) with autoscaling enabled, or kind with the CA's test provider, or a personal cloud environment. A practical recommendation: a trial account with a provider for one afternoon, with a small cluster and a low maxSize, costs a few euros and teaches an enormous amount.
- The Rutas Norte capacity plan and its cost
We close with the concrete numbers. This is the work that has to be done before a May bank holiday, and that was not done last year.
Step 1: peak demand
From the maxReplicas of 09-01 and the requests recalibrated with the VPA in 09-02:
| Component | maxReplicas |
CPU/pod | Mem/pod | Max CPU | Max mem |
|---|---|---|---|---|---|
bookings-api (api) |
30 | 412m | 800Mi | 12.36 | 23.4 GiB |
bookings-api (exporter) |
30 | 6m | 21Mi | 0.18 | 0.6 GiB |
web-store |
15 | 70m | 64Mi | 1.05 | 0.9 GiB |
notifications-worker |
20 | 120m | 410Mi | 2.40 | 8.0 GiB |
bookings-postgres |
1 | 2000m | 4Gi | 2.00 | 4.0 GiB |
redis-cache |
1 | 300m | 2Gi | 0.30 | 2.0 GiB |
| Ingress controller | 4 | 200m | 256Mi | 0.80 | 1.0 GiB |
| DaemonSets (per node) | — | 450m | 512Mi | (var.) | (var.) |
| Application subtotal | 19.09 | 39.9 GiB |
Step 2: translating into nodes
General node: 4 cores, 16 GiB.
Allocatable after kubelet and system reservations: ~3.5 cores, ~14 GiB.
DaemonSets per node: 450m and 512Mi.
Usable capacity per node: 3.05 cores and 13.5 GiB.
Nodes by CPU: 19.09 / 3.05 = 6.26 -> 7 nodes
Nodes by memory: 39.9 / 13.5 = 2.96 -> 3 nodes
CPU is the limit: 7 general nodes are needed at the peak.
Also: bookings-postgres and redis-cache live in the "data" group with a taint.
postgres: 2 cores, 4 GiB. redis: 0.3 cores, 2 GiB.
Total: 2.3 cores and 6 GiB. They fit on 1 node of 8 cores, but we keep
2 for availability (one per zone).
Over-provisioning cushion: 6 cores = 2 additional nodes.
TOTAL AT THE PEAK: 9 general nodes + 2 data nodes = 11 nodes.Step 3: configuring the groups
rutas-norte-general-a min=1 max=4 (was max=8, we adjust: 3 zones x 4 = 12 > 9)
rutas-norte-general-b min=1 max=4
rutas-norte-general-c min=1 max=4
rutas-norte-data min=2 max=3
Total maximum capacity: 12 general nodes + 3 data nodes.
Margin over the calculation: 33%.The 33% margin is not a whim: it covers estimation errors, a deployment in progress during the peak (which temporarily doubles the pods of one component), and the loss of an entire zone.
Step 4: the costs
Fictitious prices, but of a realistic order of magnitude:
| Item | Amount | Unit cost/month | Total/month |
|---|---|---|---|
| Base general nodes (3, one per zone) | 3 | €95 | €285 |
| Data nodes (2) | 2 | €190 | €380 |
| Permanent over-provisioning cushion | 2 | €95 | €190 |
| Permanent subtotal | €855 | ||
| Extra nodes during the bank holiday (6 nodes × 3 days) | 6×3 days | €3.17/day | €57 |
| Total for the bank-holiday month | €912 |
And the comparison that justifies the whole module:
| Strategy | Annual cost | Capacity at the peak | Risk |
|---|---|---|---|
| Fixed, sized for a Tuesday (the starting point) | €8,000 | Insufficient | The platform falls over. The 2023 outage |
| Fixed, sized for the peak | €20,500 | Sufficient | None, but you pay €12,500 of idle capacity 362 days a year |
| Autoscaling + a permanent cushion | €10,260 + €57 | Sufficient | Low |
| Autoscaling + a cushion only in high season | €8,550 + €57 | Sufficient | Low |
The last row is the goal: practically the same cost as the insufficient fixed sizing, with the capacity of the peak sizing. That is the entire value of module 9 summed up in one line.
And a reminder of intellectual honesty: the reduced off-season cushion only works if the CronJob that expands it actually runs. Put an alert on it: if at 16:05 on Friday the filler Deployment does not have 6 replicas, somebody has to find out.
Common Mistakes and Tips
Mistake 1: believing the HPA can take you further than the capacity allows, and not noticing. The symptom is silent: kubectl get hpa shows 30 replicas and everything looks fine. Only kubectl get pods reveals that 15 are Pending. A mandatory alert on pending pods:
Mistake 2: HPA maxReplicas that is inconsistent with the node groups' maxSize. A maxReplicas: 100 with groups that top out at 8 nodes is a lie: you will never reach 100 replicas. Always compute Σ(maxReplicas × requests) and compare it with Σ(maxSize × allocatable capacity). Document that calculation next to the manifests.
Mistake 3: pods without requests in a cluster with a CA. A pod without requests is invisible to the CA's simulation: the autoscaler believes a node has room to spare when in fact it is saturated with pods consuming without declaring anything. The LimitRange of module 3 is the defence: it forces a defaultRequest onto everything that gets created.
Mistake 4: expecting the CA to save you from a sudden spike. Three to six minutes. If your traffic rises in forty seconds, the CA arrives late. Over-provisioning (section 10) and anticipatory scaling (09-04). There is no third way.
Mistake 5: impossible PDBs that pin nodes forever. A minAvailable equal to the replica count means no pod can ever be evicted, and the node hosting them will never be removed. You will be paying for underutilised nodes without knowing why. We develop this in 09-05, but the diagnosis is in the CA logs: pdb-blocked.
Mistake 6: a single node group for everything. Without separating the data from the application, an avalanche of bookings-api pods can land on the same node as bookings-postgres and compete for its CPU. Taints and separate groups (06-05).
Mistake 7: heterogeneous nodes inside the same group. The CA uses one template per group. If the group mixes 2-core and 8-core machines, the simulation lies and the decisions are erratic. One machine type per group, always.
Tip 1: --balance-similar-node-groups if you have one group per zone. Without it, the CA can pile all the new nodes into one zone and leave you exposed to a zone outage. It is one line of configuration with a direct impact on availability.
Tip 2: monitor cluster_autoscaler_unschedulable_pods_count. The CA exposes Prometheus metrics. The ones that matter:
# Pods the CA cannot manage to place
cluster_autoscaler_unschedulable_pods_count > 0
# Nodes per group, to see how full the groups are
cluster_autoscaler_nodes_count
# Scale-up errors (no capacity at the provider, quota exhausted)
rate(cluster_autoscaler_failed_scale_ups_total[15m]) > 0That third indicator is especially valuable: a failed_scale_up means the provider has said no. Quota exhausted, an instance type unavailable in the zone, an account limit. These are problems you only discover when it is too late, unless you watch for them.
Tip 3: rehearse the scaling before the peak. A week before the bank holiday, run a load test against rutas-norte-pre (09-06) and time how long the cluster takes to reach maximum capacity. That number, measured rather than estimated, is what tells you whether the cushion is enough.
Tip 4: raise the groups' maxSize before high season and lower it afterwards. It is a one-minute change that avoids max node group size reached at the worst possible moment. Put it on the pre-event checklist.
Tip 5: keep the CA logs. They are the only source that explains why it did not scale up or why a node was not removed. Without them, the postmortem of a capacity incident is guesswork. EFK (07-05).
Tip 6: review the scale-down threshold if you have many DaemonSets. With Fluentd, Falco, the CNI, the node exporter and perhaps a service-mesh proxy, the baseline utilisation of an empty node can be around 25%. With the default 50% threshold, scale-down almost never fires. Lowering it to 35% can save entire nodes.
Exercises
Exercise 1: computing the capacity and spotting the error
The Rutas Norte team has configured this for the May bank holiday:
HPAs:
| Component | minReplicas |
maxReplicas |
requests.cpu |
requests.memory |
|---|---|---|---|---|
bookings-api |
4 | 40 | 520m | 820Mi |
web-store |
3 | 20 | 70m | 64Mi |
notifications-worker |
2 | 25 | 120m | 410Mi |
Fixed components: bookings-postgres (1 replica, 2 cores, 4Gi), redis-cache (1 replica, 300m, 2Gi), Ingress (4 replicas, 200m, 256Mi).
Node groups:
rutas-norte-general-a min=1 max=3 4 cores / 16 GiB
rutas-norte-general-b min=1 max=3 4 cores / 16 GiB
rutas-norte-data min=1 max=2 8 cores / 32 GiB taint role=dataDaemonSets: 450m and 512Mi per node. System reservation: 500m and 2Gi per node.
Namespace ResourceQuota: requests.cpu: 40, requests.memory: 80Gi.
Compute: (a) the maximum CPU and memory the application can demand; (b) the real maximum allocatable capacity of the groups; (c) whether the configuration survives the peak; (d) the percentage of bookings-api's maxReplicas that is achievable in practice; (e) what you would fix and in what order.
Exercise 2: diagnosing why a node is not removed
After the May bank holiday, the cluster still has 9 nodes even though traffic returned to normal 5 hours ago. The bill is a concern. This is what you see:
Node ip-10-0-1-42: cpu 2890m (82%) memory 7Gi (50%)
Node ip-10-0-1-88: cpu 1210m (34%) memory 3Gi (21%)
Node ip-10-0-2-19: cpu 2650m (75%) memory 6Gi (43%)
Node ip-10-0-2-73: cpu 980m (28%) memory 2Gi (14%)
Node ip-10-0-3-47: cpu 1450m (41%) memory 4Gi (28%)
Node ip-10-0-3-91: cpu 2100m (60%) memory 5Gi (36%)
Node ip-10-0-1-15: cpu 760m (21%) memory 2Gi (14%)
Node ip-10-0-2-88: cpu 1890m (54%) memory 5Gi (36%)
Node ip-10-0-3-22: cpu 890m (25%) memory 2Gi (14%)scale_down.go: Node ip-10-0-1-88 is not suitable for removal:
pod rutas-norte-pro/occupancy-reports-28912440-x7kqp has
cluster-autoscaler.kubernetes.io/safe-to-evict annotation set to false
scale_down.go: Node ip-10-0-2-73 is not suitable for removal:
pod rutas-norte-pro/redis-cache-0 is not replicated and has local storage
scale_down.go: Node ip-10-0-1-15 is not suitable for removal:
pdb-blocked: not enough pod disruption budget to move bookings-api-7c9d4f8b6d-mn2vp
scale_down.go: Node ip-10-0-3-22 is not suitable for removal:
pod rutas-norte-pro/load-generator-5f8d9c-2wxkp is not replicatedAnalyse each blocked node, say whether the block is legitimate or a mistake, and propose the fix for each case. Compute how many nodes could be removed after the fixes and the monthly saving (€95/node/month).
Exercise 3: designing the strategy for a flash campaign
Rutas Norte is launching a Black Friday campaign unlike anything before:
- Duration: exactly 2 hours, from 20:00 to 22:00 on a Friday.
- Expected traffic: x25 compared with a normal day (far more than the May bank holiday).
- The peak arrives in 20 seconds: it is announced on the radio and everyone comes in at once.
- Tickets are limited: if the platform does not respond in the first 10 minutes, the campaign has failed.
- Budget: up to €400 for that night.
- It is known three weeks in advance.
Design the complete capacity strategy. What combination of HPA, CA, over-provisioning and scheduled scaling would you use? How many nodes, how much cushion, how far in advance? What would it cost? Write the key manifests and a checklist with concrete timings for that night.
Solutions
Solution 1
(a) Maximum demand from the application:
| Component | Max replicas | CPU/pod | Total CPU | Mem/pod | Total mem |
|---|---|---|---|---|---|
bookings-api |
40 | 520m | 20.80 | 820Mi | 32.03 GiB |
web-store |
20 | 70m | 1.40 | 64Mi | 1.25 GiB |
notifications-worker |
25 | 120m | 3.00 | 410Mi | 10.01 GiB |
| Ingress | 4 | 200m | 0.80 | 256Mi | 1.00 GiB |
| General subtotal | 26.00 | 44.29 GiB | |||
bookings-postgres |
1 | 2000m | 2.00 | 4Gi | 4.00 GiB |
redis-cache |
1 | 300m | 0.30 | 2Gi | 2.00 GiB |
| Data subtotal | 2.30 | 6.00 GiB |
(b) Real maximum allocatable capacity:
General node (4 cores, 16 GiB):
System reservation: 500m, 2 GiB
DaemonSets: 450m, 512Mi
Available for pods: 4 - 0.5 - 0.45 = 3.05 cores
16 - 2 - 0.5 = 13.5 GiB
General groups: a(max=3) + b(max=3) = 6 nodes
CPU: 6 x 3.05 = 18.30 cores
Memory: 6 x 13.5 = 81.00 GiB
Data node (8 cores, 32 GiB):
Available: 8 - 0.5 - 0.45 = 7.05 cores
32 - 2 - 0.5 = 29.5 GiB
Data group (max=2): 2 nodes
CPU: 14.10 cores
Memory: 59.00 GiB(c) Does it survive the peak? NO.
GENERAL:
Needed: 26.00 cores
Available: 18.30 cores
SHORTFALL: 7.70 cores (30% of what is needed)
Memory: needed 44.29 GiB, available 81 GiB. Plenty.
-> CPU IS THE LIMIT.
DATA:
Needed: 2.30 cores and 6 GiB
Available: 14.10 cores and 59 GiB
-> Enormously oversized. The data group is VERY over-provisioned.And there is a second problem many people would overlook: the ResourceQuota.
Total requests.cpu demand: 26.00 + 2.30 = 28.30
Quota: 40. It fits (71% used).
Total requests.memory demand: 44.29 + 6.00 = 50.29 GiB
Quota: 80 GiB. It fits (63% used).
The quota is NOT the problem here, but note that during a rolling update of
bookings-api the old and new pods coexist, which can add up to 25% more:
28.30 x 1.25 = 35.4 cores. That is 88% of the quota. Tight.(d) Percentage of bookings-api's maxReplicas that is achievable:
CPU available in the general group: 18.30 cores.
Consumption of the other general components (at their maximum):
web-store 1.40 + worker 3.00 + ingress 0.80 = 5.20 cores
CPU available for bookings-api: 18.30 - 5.20 = 13.10 cores
Achievable replicas: 13.10 / 0.52 = 25.19 -> 25 replicas
25 out of 40 = 62.5%The maxReplicas: 40 is fiction: the reality is 25 replicas. And worse: when the HPA asks for replicas 26 to 40, they will sit in Pending and nobody will find out unless there is an alert.
(e) Fixes, in order of priority:
Priority 1 (essential): raise the maxSize of the general groups.
Needed: 26.00 cores / 3.05 per node = 8.52 -> 9 nodes
With a third group per zone (advisable for 09-05): 3 groups x max=4 = 12 nodes
Capacity: 12 x 3.05 = 36.6 cores. A 41% margin over what is needed.rutas-norte-general-a min=1 max=4
rutas-norte-general-b min=1 max=4
rutas-norte-general-c min=1 max=4 <- new group, third zonePriority 2: reduce the maxSize of the data group.
The data group needs 2.30 cores and 6 GiB. A single 8-core node is more than
enough. With max=2 we keep availability (one per zone) but there is no reason
to allow more. It stays at max=2, which is already correct.
BUT: the machine type is oversized. A node of 4 cores and 16 GiB would be
enough for postgres (2 cores, 4 GiB) and redis (0.3, 2 GiB). Changing the type
from 8/32 to 4/16 would save ~95 EUR/month per node, 190 EUR/month in total.
A caveat: postgres benefits from memory for shared_buffers and the page cache.
Before reducing, measure with the VPA (09-02) whether 16 GiB is enough. With
requests of 4 GiB and a 16 GiB node, there is 10 GiB of system cache: probably yes.Priority 3: review bookings-api's maxReplicas: 40.
With 12 general nodes there is room for 25 replicas... wait, let's recompute with the new capacity:
New general capacity: 36.6 cores
Minus the other components: 36.6 - 5.20 = 31.4 cores
Achievable bookings-api replicas: 31.4 / 0.52 = 60
Now the 40 do fit. The maxReplicas: 40 is achievable.Priority 4: add the over-provisioning cushion. Without it, replicas 5 to 40 will take between 3 and 6 minutes to appear. With 6 filler pods of 1 core, the first 11 replicas are instantaneous.
Priority 5: the pending-pods alert. Without it, this whole calculation could be wrong and nobody would know until it was too late.
Solution 2
Node-by-node analysis:
| Node | Utilisation | State | Reason for the block | Legitimate? |
|---|---|---|---|---|
| ip-10-0-1-42 | 82% | Busy | Above the threshold | Yes, not a candidate |
| ip-10-0-1-88 | 34% | Blocked | safe-to-evict: false on occupancy-reports |
It depends |
| ip-10-0-2-19 | 75% | Busy | Above the threshold | Yes |
| ip-10-0-2-73 | 28% | Blocked | redis-cache-0, unreplicated and with local storage |
Yes, legitimate |
| ip-10-0-3-47 | 41% | Candidate | — | It should be removed |
| ip-10-0-3-91 | 60% | Busy | Above the threshold | Yes |
| ip-10-0-1-15 | 21% | Blocked | bookings-api PDB |
NO, it is a mistake |
| ip-10-0-2-88 | 54% | Busy | Above the threshold | Yes |
| ip-10-0-3-22 | 25% | Blocked | load-generator with no controller |
NO, it is junk |
Case 1: ip-10-0-1-88 — occupancy-reports with safe-to-evict: false.
Legitimacy: conditional. If the Job is running right now, the block is correct: killing it would lose the work. But the bank holiday ended 5 hours ago and occupancy-reports is a nightly CronJob.
# Is the Job still alive?
kubectl get jobs -n rutas-norte-pro
kubectl get pods -n rutas-norte-pro -l app=occupancy-reportsIf it shows as Running and has been for hours, it is a stuck Job: it probably failed and got wedged. The fix:
# 1. Investigate why it is still alive
kubectl logs -n rutas-norte-pro occupancy-reports-28912440-x7kqp --tail=50
# 2. If it is stuck, delete it
kubectl delete job occupancy-reports-28912440 -n rutas-norte-proAnd a structural prevention in the CronJob:
spec:
jobTemplate:
spec:
# If it has not finished in 2 hours, it is cut off. Avoids zombie jobs
# that pin nodes indefinitely.
activeDeadlineSeconds: 7200
ttlSecondsAfterFinished: 3600Node recoverable: YES, after cleaning up the Job.
Case 2: ip-10-0-2-73 — redis-cache-0.
Legitimacy: yes, completely. redis-cache is a single-replica StatefulSet with a volume. The CA cannot move it safely, and it certainly must not do so on its own initiative: restarting the cache dumps the whole load onto bookings-postgres.
Fix: none in the short term. In the medium term, the options are:
- Pin
redis-cacheto the data node group with affinity and a taint (06-05), so that it does not occupy a general node. - Accept it: a node dedicated to the cache is not waste if it is sized properly.
The best fix is the first: redis-cache should not be on a general node. It is an architectural decision we should already have taken.
Node recoverable: NOT directly, but yes once redis-cache is moved to the data group.
Case 3: ip-10-0-1-15 — the PDB blocking bookings-api.
Legitimacy: no, it is a configuration error. bookings-api is a Deployment with many replicas: moving one pod should be trivial. That the PDB prevents it means the PDB is wrong.
The likely suspect:
After the bank holiday, the HPA has come down to 4-6 replicas. With minAvailable: 30 and 6 live replicas, no pod can ever be evicted: we are already below the minimum.
The fix:
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: bookings-api
namespace: rutas-norte-pro
spec:
# A PERCENTAGE, not an absolute number. It adapts to the HPA's scaling on its own.
# With 6 replicas it allows evicting 1; with 30, it allows evicting 6.
maxUnavailable: 20%
selector:
matchLabels:
app: bookings-apiThis is the central lesson of 09-05, previewed: with an HPA, absolute-number PDBs are a trap. The replica count changes on its own, and a minimum that was reasonable at the peak becomes impossible in the trough.
Node recoverable: YES, immediately after fixing the PDB.
Case 4: ip-10-0-3-22 — load-generator with no controller.
Legitimacy: no, it is junk. It is the load-test pod from section 12 of 09-01, created with kubectl run without a controller and then forgotten. It has spent days there pinning an entire node.
Prevention:
- Always use
--rmwithkubectl runfor tests. - A Kyverno policy (08-03) that rejects pods without
ownerReferencesinrutas-norte-pro. - A periodic review of orphan pods:
kubectl get pods -n rutas-norte-pro -o json | \
python3 -c "import sys,json; [print(p['metadata']['name']) for p in json.load(sys.stdin)['items'] if not p['metadata'].get('ownerReferences')]"Node recoverable: YES, immediately.
Summary and saving:
| Node | Action | Removed? |
|---|---|---|
| ip-10-0-1-88 | Delete the stuck Job | Yes |
| ip-10-0-2-73 | Move redis-cache to the data group (medium term) |
Yes, in the medium term |
| ip-10-0-3-47 | None: it is a candidate, it just needs time | Yes |
| ip-10-0-1-15 | Fix the PDB to maxUnavailable: 20% |
Yes |
| ip-10-0-3-22 | Delete the orphan pod | Yes |
Nodes removable in the short term: 4 (ip-10-0-1-88, -3-47, -1-15, -3-22)
Nodes removable in the medium term: 1 more (ip-10-0-2-73)
Immediate saving: 4 x 95 EUR = 380 EUR/month = 4,560 EUR/year
Medium-term saving: 5 x 95 EUR = 475 EUR/month = 5,700 EUR/yearVerification after the fixes:
# Wait out the scale-down-unneeded-time (10 min by default) and verify
watch -n 30 'kubectl get configmap cluster-autoscaler-status -n kube-system \
-o jsonpath="{.data.status}" | grep -A2 ScaleDown'It should go from NoCandidates to CandidatesPresent and then to InProgress.
The underlying lesson of the exercise: zombie nodes are never the Cluster Autoscaler's fault. They are always the consequence of something misconfigured in the workloads: an impossible PDB, a stuck Job, an orphan pod, a stateful component in the wrong place. The CA is the messenger; the logs are the message.
Solution 3
Analysing the scenario. This case is qualitatively different from the May bank holiday:
| Factor | May bank holiday | Black Friday |
|---|---|---|
| Duration | 3 days | 2 hours |
| Multiplier | x10 | x25 |
| Rise speed | Minutes | 20 seconds |
| Notice | Known (the calendar) | 3 weeks |
| Tolerance for failure | Low | None: 10 minutes and it is over |
The key conclusion: reactive scaling is NO USE here. Neither the HPA (15 s of detection + 30 s of start-up) nor the CA (3-6 min) arrives in time for a rise of 20 seconds. All the capacity has to be booted and warm before 20:00.
Strategy: total pre-scaling.
Step 1: compute the capacity needed.
Normal Friday-night traffic: ~250 rps.
Expected traffic: 250 x 25 = 6,250 rps.
From the load test (09-06) we know: each bookings-api replica sustains
~85 rps with acceptable p95 latency.
Replicas needed: 6,250 / 85 = 73.5 -> with a 20% margin: 88 replicas.
CPU: 88 x 520m = 45.76 cores.
Memory: 88 x 820Mi = 70.4 GiB.
web-store: serves static assets, scales better. 6,250 rps / 400 rps per replica = 16
-> with margin: 20 replicas. 20 x 70m = 1.4 cores.
notifications-worker: the queue fills up but can be drained AFTER 22:00.
It does NOT need scaling during the campaign. It stays at 4 replicas and is
scaled to 25 from 22:00 onwards, when there is spare capacity.
A deliberate decision: absolute priority to sales.
Ingress: 8 replicas (double the usual). 1.6 cores.
TOTAL AT THE PEAK: 45.76 + 1.4 + 0.48 + 1.6 = 49.24 cores
70.4 + 1.25 + 1.6 + 2 = 75.25 GiBStep 2: translate into nodes.
Usable capacity per general node (4 cores): 3.05 cores, 13.5 GiB.
By CPU: 49.24 / 3.05 = 16.1 -> 17 nodes
By memory: 75.25 / 13.5 = 5.6 -> 6 nodes
CPU IS THE LIMIT: 17 general nodes.
Alternative: bigger nodes. With nodes of 8 cores / 32 GiB:
Usable per node: 8 - 0.5 - 0.45 = 7.05 cores
Nodes needed: 49.24 / 7.05 = 6.98 -> 7 nodes
7 large nodes instead of 17 small ones:
- Fewer nodes to boot = less total provisioning time
- Fewer replicated DaemonSets (17 x 450m = 7.65 cores versus 7 x 450m = 3.15)
- Worse granularity and worse fault tolerance
DECISION: 8-core nodes for that night. The DaemonSet saving (4.5 cores) is
decisive, and since all the capacity is booted BEFORE, the worse granularity
does not matter.Step 3: the manifests.
# k8s/environments/pro/campaign/hpa-bookings-api-blackfriday.yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: bookings-api
namespace: rutas-norte-pro
annotations:
rutasnorte.example/campaign: "black-friday-2026"
rutasnorte.example/revert-before: "2026-11-28T23:00:00Z"
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: bookings-api
# minReplicas RAISED TO 88. This is the key to the whole strategy:
# the capacity is BOOTED AND WARM before 20:00, we do not wait for the
# HPA to react. With a 20-second rise, any reactive mechanism
# arrives late.
minReplicas: 88
maxReplicas: 110 # A 25% margin in case the estimate falls short
metrics:
- type: ContainerResource
containerResource:
name: cpu
container: api
target:
type: Utilization
averageUtilization: 60
behavior:
scaleUp:
stabilizationWindowSeconds: 0
selectPolicy: Max
policies:
- type: Percent
value: 100
periodSeconds: 15
- type: Pods
value: 15
periodSeconds: 15
scaleDown:
# DISABLED during the campaign. Not a single replica fewer between
# 19:00 and 22:30, whatever the metrics do. A 30-second trough must
# not destroy capacity that took 20 minutes to raise.
selectPolicy: Disabled# k8s/environments/pro/campaign/deployment-capacity-filler-blackfriday.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: capacity-filler
namespace: rutas-norte-pro
spec:
# Cushion expanded to 10 pods of 1 core. It instantly absorbs
# replicas 89 to 108 if the estimate falls short.
replicas: 10
selector:
matchLabels:
app: capacity-filler
template:
metadata:
labels:
app: capacity-filler
annotations:
cluster-autoscaler.kubernetes.io/safe-to-evict: "true"
spec:
priorityClassName: capacity-filler
terminationGracePeriodSeconds: 0
topologySpreadConstraints:
- maxSkew: 1
topologyKey: kubernetes.io/hostname
whenUnsatisfiable: ScheduleAnyway
labelSelector:
matchLabels:
app: capacity-filler
containers:
- name: pause
image: registry.k8s.io/pause:3.9
resources:
requests: {cpu: "1", memory: 1Gi}
limits: {cpu: "1", memory: 1Gi}Node groups for that night:
rutas-norte-bf-a min=3 max=4 8 cores / 32 GiB zone A
rutas-norte-bf-b min=3 max=4 8 cores / 32 GiB zone B
rutas-norte-bf-c min=2 max=4 8 cores / 32 GiB zone C
Guaranteed minimum total: 8 nodes (already booted, not dependent on the CA)
Maximum total: 12 nodesNote the important detail: the groups' min is raised to 3, we do not rely on the CA. By setting the minimum high, the provider boots the machines and the CA cannot remove them. The capacity is guaranteed by construction, not by reaction.
Step 4: the cost.
Node of 8 cores / 32 GiB: ~0.26 EUR/hour on demand.
Expanded-capacity window: 18:00 to 23:00 = 5 hours.
Extra nodes over the baseline (baseline: 3 nodes of 4 cores):
8 nodes of 8 cores x 5 hours x 0.26 = 10.40 EUR
Estimated additional egress traffic: ~15 EUR
Margin for incidents (nodes up to the maximum of 12): +4 nodes x 5 h x 0.26 = 5.20 EUR
TOTAL ESTIMATE: ~31 EUR
BUDGET: 400 EUR.
Margin: 369 EUR (92%).The result is surprising and it is the lesson of the exercise: sizing generously for two hours is dirt cheap. The cost of cloud infrastructure is proportional to time, and two hours of over-provisioning cost less than a coffee per node. With a €400 budget, there is no excuse whatsoever for skimping on capacity that night.
With the spare margin you can afford improvements:
- Going to 12 nodes from the start instead of 8: +€5.20.
- Doubling the
redis-cachereplicas as a read replica: +€2. - Keeping the capacity until 2:00 to drain the notifications queue without rushing: +€15.
All of that fits comfortably.
Step 5: the checklist for the night.
| Time | Action | Owner | Verification |
|---|---|---|---|
| 3 weeks before | Load test on rutas-norte-pre at 6,250 rps (09-06) |
Platform | Measure real rps per replica |
| 2 weeks before | Check the provider quotas: does it allow 12 nodes of 8 cores? | Platform | Raise a quota ticket if not |
| 1 week before | Full dress rehearsal: apply the manifests in pre and measure the complete start-up time |
Platform | Stopwatch |
| 1 week before | Review the PDBs: maxUnavailable as a percentage, not absolute |
Platform | kubectl get pdb |
| 3 days before | Freeze deployments. No code changes until Monday | The whole team | Block in CI |
| 17:00 | Raise the node groups' min to 3/3/2 |
Platform | kubectl get nodes = 8 |
| 17:30 | Verify that the 8 nodes are Ready and have the images preloaded |
Platform | kubectl get nodes |
| 18:00 | Apply the campaign HPA (minReplicas: 88) |
Platform | kubectl get hpa |
| 18:00 | Expand the filler to 10 replicas | Platform | kubectl get pods -l app=capacity-filler |
| 18:30 | Verify 88 pods Running and Ready, not merely created |
Platform | kubectl get pods --field-selector status.phase=Running | wc -l |
| 18:45 | Smoke test: 100 real end-to-end test bookings | QA | All OK |
| 19:00 | Check Grafana: p95 latency stable, zero errors, zero Pending |
Platform | Dashboards |
| 19:30 | Pre-warm the caches: query the 50 best-selling routes | Platform | redis-cli DBSIZE |
| 19:45 | War room open. The whole team connected | Everyone | — |
| 19:55 | Final check. kubectl get pods --all-namespaces | grep -v Running |
Platform | Empty |
| 20:00 | OPENING. Watch: p95 latency, error rate, Pending pods, PostgreSQL connections |
Everyone | Dashboards |
| 20:00-22:00 | Continuous watch. The rule: touch nothing unless there is a confirmed incident | Everyone | — |
| 22:00 | Campaign closes | — | — |
| 22:15 | Scale notifications-worker to 25 replicas to drain the queue |
Platform | Queue length |
| 23:00 | Restore the normal HPA (minReplicas: 4) |
Platform | kubectl apply of the base manifest |
| 23:30 | Lower the node groups' min. The CA will remove the rest |
Platform | kubectl get nodes |
| The next day | Analysis: real rps, latencies, comparison against the estimate | Everyone | Report |
Critical notes on the plan:
-
The 18:30 verification is the most important one. 88 pods created is not the same as 88 pods
Ready. If at 18:30 there are 20 inPending, there are two hours to fix it. If you find out at 20:01, there is nothing to be done. -
The
scaleDown: Disabledis non-negotiable. Without it, a one-minute trough between 20:15 and 20:16 could destroy 30 replicas that would take minutes to come back. -
notifications-workeris not scaled during the campaign. It is a deliberate decision: the confirmation emails can wait two hours, the sale cannot. All the capacity goes to sales. This is the kind of explicit trade-off that distinguishes a capacity plan from a wish list. -
The "touch nothing" rule during the campaign. Almost every serious incident at events like this is caused by somebody trying to fix something under pressure. If the dashboards are green, you look and you do not touch.
-
Pre-warming the caches at 19:30 saves the first 5,000 users from paying the cost of cold queries against PostgreSQL. It is half an hour of work with an enormous impact on the first few minutes.
Conclusion
Cluster autoscaling is the level that holds up the other two. Without it, the HorizontalPodAutoscaler has a hard, invisible ceiling: it asks for replicas, the Deployment creates them, and they sit in Pending while the dashboards show numbers that do not correspond to reality.
The essentials:
- The symptom is always the same: pods in
PendingwithFailedSchedulingandInsufficient cpu. Recognising it and alerting on it comes first. - The Cluster Autoscaler watches unschedulable pods, simulates which node group would host them, and chooses with an
expander.least-wasteis the most sensible in general. - Scaling up is easy; scaling down is where the problems are. A utilisation threshold, a waiting time, and a list of blockers you have to know: pods without a controller, local storage, restrictive PDBs, system pods, and the
safe-to-evict: "false"annotation. - Diagnosis means reading three places: the
cluster-autoscaler-statusConfigMap, theNotTriggerScaleUpevents on the pending pods, and the CA logs with theirnot suitable for removal. - A node takes between 3 and 6 minutes to go from being requested to serving traffic. That does not save you from a spike that rises in forty seconds, however well configured everything is.
- Over-provisioning with negative-priority pods turns three and a half minutes into three seconds, in exchange for paying for idle capacity. It is the most elegant use of the
PriorityClassand preemption of 06-05, and it regenerates itself. - Karpenter picks the machine according to the pending pods instead of working with fixed groups, and consolidates continuously. Better on AWS and with heterogeneous workloads; the CA remains the solid, predictable option everywhere else.
- The three autoscalers fit together as a cascade: the HPA creates pods → if they do not fit,
Pending→ the CA adds nodes. And the cushion cuts out the long path. - The Rutas Norte capacity plan costs practically the same as last year's insufficient fixed sizing, but with the capacity of the peak sizing. That is the entire value of the module.
But we still have an underlying problem that none of the three lessons has solved. Everything we have built is reactive: it waits for CPU to rise before acting. And there are two cases where that simply does not work.
The first is notifications-worker. During the May bank holiday it can have forty thousand emails waiting in the queue and CPU at 20%, because its job is to wait for responses from the mail server, not to compute. A CPU-based HPA would see that 20% against a 70% target and conclude that there are replicas to spare: it would scale down precisely when the most consumers are needed.
The second is the very moment the sale opens. We know it happens at 10:00 sharp. It is information we have months in advance. And yet our entire system waits until 10:00:15 to discover, through CPU, something we already knew in March.
In the next lesson, Event-Driven and Custom-Metric Scaling with KEDA, we solve both. We will see how the HPA can consume business metrics through the aggregation API, what KEDA is and why it does not replace the HPA but feeds it, how to scale notifications-worker by the real length of its queue — scale-to-zero included —, how to scale bookings-api by requests per second using the metrics we instrumented in 07-03, and how to schedule a trigger that pre-warms the whole platform half an hour before the May bank-holiday sale opens.
Kubernetes Course
Module 1: Introduction to Kubernetes
- What Is Kubernetes?
- Kubernetes Architecture
- Key Concepts and Terminology
- Setting Up a Kubernetes Cluster
- The Kubernetes CLI: kubectl
- Objects, YAML Manifests and the Declarative Model
- The Course Project: the Rutas Norte Platform
Module 2: Core Kubernetes Components
- Pods
- ReplicaSets
- Deployments
- Updates, Rollbacks and Deployment Strategies
- Services
- Namespaces
- Labels, Selectors and Annotations
Module 3: Configuration and Secret Management
- ConfigMaps
- Secrets
- Environment Variables
- Resource Quotas and Limits
- LimitRanges and Quality of Service (QoS) Classes
- ServiceAccounts and API Access from Pods
Module 4: Networking in Kubernetes
- Cluster Networking
- Service Types
- Internal DNS and Service Discovery
- Ingress Controllers
- TLS and Certificate Management with cert-manager
- Network Policies
Module 5: Storage in Kubernetes
- Volumes
- Persistent Volumes
- Persistent Volume Claims
- Storage Classes
- Dynamic Provisioning, Expansion and Snapshots
- Backup and Restore of Persistent Data
Module 6: Advanced Kubernetes Concepts
- StatefulSets
- DaemonSets
- Jobs and CronJobs
- Init Containers, Sidecars and Multi-Container Patterns
- Scheduling: Affinity, Taints and Tolerations
- Custom Resource Definitions (CRDs)
- Operators and the Controller Pattern
Module 7: Monitoring and Logging
- Health Checks and Probes
- Metrics Server and kubectl top
- Monitoring with Prometheus
- Visualization and Alerting with Grafana and Alertmanager
- Centralized Logging with Elasticsearch, Fluentd and Kibana (EFK)
- Application Debugging and Cluster Events
Module 8: Kubernetes Security
- Role-Based Access Control (RBAC)
- Security Contexts and Container Hardening
- Pod Security Policies and Pod Security Standards
- Network Security
- Image Security
- Auditing, Scanning and Vulnerability Management
Module 9: Scaling and Performance
- Horizontal Pod Autoscaling
- Vertical Pod Autoscaling
- Cluster Autoscaling
- Event-Driven and Custom-Metric Scaling with KEDA
- High Availability: PodDisruptionBudgets and Topology
- Performance Tuning
Module 10: Kubernetes Ecosystem and Tooling
- Minikube and Local Environments with kind
- Kubeadm
- Helm
- Kustomize
- GitOps with Argo CD and Flux
- Managed Kubernetes: EKS, AKS and GKE
Module 11: Case Studies and Real-World Applications
- Deploying a Web Application
- Running Stateful Applications
- CI/CD with Kubernetes
- Deployment Strategies: Blue-Green and Canary
- Multi-Cluster Management
- Production Operations: Incidents, Runbooks and Costs
