The previous lesson ended by pointing out the hidden assumption behind everything we have built so far: we have taken it for granted that, when the HorizontalPodAutoscaler asks for thirty replicas of bookings-api, there is somewhere to put them. There is not. The Rutas Norte cluster has four nodes, and thirty bookings-api pods plus fifteen web-store plus twenty notifications-worker do not fit in them, not by a long way.

That situation has a name and a very concrete appearance: pods in Pending, with a FailedScheduling event saying Insufficient cpu. We already know how to read it from 06-05 and 07-06. What we do not yet know is what to do about it automatically.

This is the third and final level of scaling. The HPA changes the number of pods. The VPA changes the size of each pod. Cluster autoscaling changes the number of nodes. Without it, the other two have a hard ceiling they cannot break through.

We are going to look at the Cluster Autoscaler in detail: how it decides to scale up, the delicate business of removing nodes (which is where all the real problems live), how long a node genuinely takes to become available, the over-provisioning pattern that compensates for those times, and Karpenter as the modern alternative. We will finish with Rutas Norte's concrete capacity plan for the May bank-holiday weekend, with its numbers and its cost.

Contents

  1. The hard ceiling: when the HPA asks and there is no room
  2. The exact symptom: Pending and FailedScheduling
  3. What the Cluster Autoscaler is and where it lives
  4. How it decides to scale up: simulation and expanders
  5. Node groups and their relationship with the provider
  6. Removing nodes: the delicate part
  7. The reasons a node cannot be drained
  8. Diagnosis: the status ConfigMap and the events
  9. The real timings: how long a node genuinely takes
  10. Over-provisioning with filler pods
  11. Karpenter: the modern alternative
  12. How the three autoscalers fit together
  13. What can be practised on minikube
  14. The Rutas Norte capacity plan and its cost
  15. Common Mistakes and Tips
  16. Exercises
  17. Conclusion

  1. The hard ceiling: when the HPA asks and there is no room

Let's set the scene with real numbers. The Rutas Norte production cluster:

Amount
Nodes 4
CPU per node 4 cores
Memory per node 16 GiB
Total raw CPU 16 cores
Reserved for kubelet and the system ~0.5 cores and 1.5 GiB per node
Total allocatable CPU ~14 cores
Total allocatable memory ~58 GiB

And the idle consumption, with every component at minReplicas:

Component Replicas requests.cpu Total CPU
bookings-api (api + sidecar) 4 520m 2.08
web-store 3 200m 0.60
notifications-worker 2 320m 0.64
bookings-postgres 1 2000m 2.00
redis-cache 1 300m 0.30
DaemonSets (Fluentd, Falco, CNI) 4×3 ~150m 1.80
Ingress controller 2 200m 0.40
Total at idle 7.82 cores

That leaves about 6.2 free cores. Now the May bank holiday arrives:

The bookings-api HPA wants 30 replicas.
Current: 4. New: 26. Each one: 520m.
CPU needed: 26 x 520m = 13.52 cores.

CPU available: 6.2 cores.
CPU missing: 7.32 cores.

Replicas that fit: 6.2 / 0.52 = 11.9 -> 11 new replicas.
Replicas that do NOT fit: 15.

Fifteen bookings-api pods will be left in Pending. The HPA will have done its job: it asked for 30 replicas and the Deployment created them. The ReplicaSet has them registered. But the scheduler cannot find anywhere to put them, and there they stay.

And here is the worst of it: the HPA does not know. In kubectl get hpa you will see REPLICAS 30. In Grafana, the replicas panel will say 30. The platform will have 15 pods ready and 15 ghosts. And since the HPA computes the CPU average only across the pods that do exist and are ready, it will see them at 90% and will ask for... nothing more, because it is already at maxReplicas.

The result is a platform that falls over under the illusion of being scaled.

  1. The exact symptom: Pending and FailedScheduling

Let's learn to recognise it precisely, because it is the trigger for everything else.

kubectl get pods -n rutas-norte-pro -l app=bookings-api
NAME                            READY   STATUS    RESTARTS   AGE
bookings-api-7c9d4f8b6d-2mk8p   2/2     Running   0          4h
bookings-api-7c9d4f8b6d-5xqzn   2/2     Running   0          4h
...
bookings-api-7c9d4f8b6d-qw8vz   0/2     Pending   0          92s
bookings-api-7c9d4f8b6d-rt4nk   0/2     Pending   0          92s
bookings-api-7c9d4f8b6d-sv7mx   0/2     Pending   0          91s

Pending with 0/2 and no restarts: the pod exists in the API but no node has accepted it. There are no containers running because there is nowhere to run them.

kubectl describe pod bookings-api-7c9d4f8b6d-qw8vz -n rutas-norte-pro
Name:             bookings-api-7c9d4f8b6d-qw8vz
Namespace:        rutas-norte-pro
Priority:         0
Node:             <none>
Status:           Pending
...
Events:
  Type     Reason            Age    From               Message
  ----     ------            ----   ----               -------
  Warning  FailedScheduling  95s    default-scheduler  0/4 nodes are available:
           4 Insufficient cpu. preemption: 0/4 nodes are available:
           4 No preemption victims found for incoming pod.
  Normal   NotTriggerScaleUp 90s    cluster-autoscaler  pod didn't trigger scale-up:
           1 max node group size reached

Two golden lines:

Node: <none> — confirmation that it is not assigned to any node.

0/4 nodes are available: 4 Insufficient cpu — the scheduler evaluated the four nodes and all four failed on insufficient CPU. The message always has the shape X/Y nodes are available: <reasons>, and the reasons are grouped by cause. The ones you will meet in practice:

Message Meaning Does the CA fix it?
Insufficient cpu There is not enough allocatable CPU Yes
Insufficient memory There is not enough allocatable memory Yes
node(s) had untolerated taint {...} Taints without a toleration (06-05) Depends on the node group
node(s) didn't match Pod's node affinity/selector Node affinity not satisfied (06-05) Only if there is a group that satisfies it
node(s) didn't match pod topology spread constraints Topology spread (09-05) Sometimes
node(s) had volume node affinity conflict The PV is in another zone (module 5) No
pod has unbound immediate PersistentVolumeClaims The PVC has not been provisioned No
Insufficient nvidia.com/gpu Extended resource exhausted Yes, if there is a GPU group

The NotTriggerScaleUp line is the Cluster Autoscaler's, and in this example it says max node group size reached: the CA is installed, it has seen the pending pod, and it can do nothing because the node group is already at its maximum size. We will come back to this message in section 8, because it is the main diagnostic tool.

A very useful command to see every pending pod at a glance:

kubectl get pods --all-namespaces --field-selector status.phase=Pending

And to see the real occupancy of the nodes:

kubectl describe nodes | grep -A 8 "Allocated resources"
Allocated resources:
  (Total limits may be over 100 percent, i.e., overcommitted.)
  Resource           Requests      Limits
  --------           --------      ------
  cpu                3410m (97%)   6200m (177%)
  memory             9856Mi (67%)  14Gi (98%)

Note that the column that matters is Requests, not Limits. The scheduler places pods according to requests. A node at 97% of requests is full for scheduling purposes, even if real consumption is 30%. It is the lesson of module 3, now with capacity consequences.

  1. What the Cluster Autoscaler is and where it lives

The Cluster Autoscaler (CA) is a component that adjusts the number of nodes in the cluster. Like the VPA, it lives in the kubernetes/autoscaler repository and is not part of the Kubernetes core.

How it works, in one sentence: it watches the pods that cannot be scheduled and adds nodes; it watches the underutilised nodes and removes them.

Where it runs

It runs as a Deployment inside the cluster itself, normally in kube-system, with a ServiceAccount that holds broad permissions (read pods and nodes, create events, evict pods) and cloud-provider credentials to manipulate the node groups.

On managed Kubernetes (10-06) the provider installs and configures it for you: on EKS, AKS and GKE you switch autoscaling on with a toggle and the provider takes care of the rest. That is the most common scenario, and it is why you should understand its behaviour even if you never install it by hand.

What it needs in order to work

Requirement Why
An API to create and destroy nodes The CA does not boot machines: it asks the provider to do it
Node groups with a minimum and maximum size That is the unit it operates on
Homogeneous nodes inside each group It simulates using one template node per group
That every pod declares requests Without requests it cannot simulate whether a pod would fit

That third requirement matters more than it looks. The CA assumes that every node in a group is identical. When it simulates whether a pod would fit on a new node from group A, it uses group A's template. If the group has heterogeneous machines, the simulation lies and the CA makes the wrong decisions. Hence the universal recommendation: one machine type per node group.

And the fourth requirement connects with the whole of module 3: a pod without requests is invisible to the scheduler for capacity purposes, and to the CA as well. Here, once again, requests is the foundation of everything.

  1. How it decides to scale up: simulation and expanders

The CA loop runs every 10 seconds by default (--scan-interval).

The decision process

  1. It lists the unschedulable pods. The ones that have been Pending with FailedScheduling for at least --max-pod-provisioning-time.
  2. It groups equivalent pods. Pods from the same controller with the same requirements are treated as one group, so as not to simulate the same thing twenty times.
  3. For each node group, it simulates. "If I add a node from group A, how many of these pending pods would fit?" The simulation uses the real scheduler, with all its predicates: resources, taints, affinities, topology spread, host ports.
  4. It discards the groups that do not help. If adding a node from group B would not allow a single pod to be scheduled (because it has a taint the pod does not tolerate, for example), that group is discarded.
  5. It chooses among the viable groups using the expander.
  6. It computes how many nodes are needed and asks the provider to grow the group.

An important detail of step 6: the CA can add several nodes at once. If there are 15 pending pods of 520m each and 6 fit on a 4-core node, it will ask for 3 nodes in one go, not one at a time. This is bounded by --max-nodes-total and by the group's maximum.

The expanders

When several node groups would do the job, the expander decides which one.

Expander Criterion When to use it
random Picks at random among the viable ones The default. Only valid if every group is equivalent
most-pods The group that would let it schedule the most pending pods When the priority is to unblock fast
least-waste The group that would leave the least idle CPU and memory after placing the pods The most sensible in general. It optimises the fit
price The cheapest group (requires provider support) When cost rules and there are varied machine types
priority According to a list of priorities in a ConfigMap Explicit control: try spot first, then on-demand

A concrete example of least-waste. There are 3 pending pods needing 500m and 512Mi each (total: 1.5 cores and 1.5 GiB), and two groups available:

Group Node type Allocatable After placing the 3 pods Waste
A 2 cores, 8 GiB 1.8 cores / 7 GiB 0.3 cores and 5.5 GiB free CPU: 17%, RAM: 79%
B 8 cores, 32 GiB 7.7 cores / 30 GiB 6.2 cores and 28.5 GiB free CPU: 81%, RAM: 95%

least-waste picks group A: it fits better and you do not pay for 6 idle cores. most-pods would pick B (many more future pods would fit). price would pick A if it is cheaper.

Configuring the expander:

# Fragment of the cluster-autoscaler Deployment
spec:
  containers:
    - name: cluster-autoscaler
      image: registry.k8s.io/autoscaling/cluster-autoscaler:v1.30.0
      command:
        - ./cluster-autoscaler
        - --cloud-provider=aws
        - --namespace=kube-system
        - --expander=least-waste
        - --scan-interval=10s
        - --balance-similar-node-groups         # Spreads across equivalent groups (zones)
        - --skip-nodes-with-local-storage=false
        - --skip-nodes-with-system-pods=false
        - --node-group-auto-discovery=asg:tag=k8s.io/cluster-autoscaler/enabled,k8s.io/cluster-autoscaler/rutas-norte

The --balance-similar-node-groups option deserves a mention because it solves a real availability problem: if you have one node group per availability zone, this flag makes the CA spread new nodes evenly across zones instead of piling them into one. Without it, you can end up with twelve nodes in zone A and none in zone B, and a zone outage leaves you with no platform. It connects directly with what we will see in 09-05.

The priority expander: try the cheap machines first

A very profitable pattern in the cloud. Spot (interruptible) instances cost a fraction of the price, but the provider can reclaim them with two minutes' notice.

apiVersion: v1
kind: ConfigMap
metadata:
  name: cluster-autoscaler-priority-expander
  namespace: kube-system
data:
  priorities: |-
    100:
      - rutas-norte-spot-.*        # First choice: spot, much cheaper
    50:
      - rutas-norte-ondemand-.*    # Fallback: on-demand, always available

The CA tries the priority-100 group first. If the provider has no spot capacity available, it falls back to the priority-50 one. For notifications-worker, which tolerates interruptions without trouble, it is the obvious choice. For bookings-postgres, never.

  1. Node groups and their relationship with the provider

The CA does not create machines: it manipulates the provider's node groups. Every provider has its own name for the same thing:

Provider Name of the concept
AWS Auto Scaling Group (ASG)
Azure Virtual Machine Scale Set (VMSS) / Agent Pool
Google Cloud Managed Instance Group (MIG) / Node Pool
On-premise / kubeadm Requires a custom provider or Cluster API

A node group has a minimum size, a maximum and a current one. The CA only changes the current one, within those bounds. Everything else (the machine type, the image, the initial taints and labels) is defined by the group.

The Rutas Norte groups:

rutas-norte-general-a   min=1  max=8   4 cores / 16 GiB   zone A
rutas-norte-general-b   min=1  max=8   4 cores / 16 GiB   zone B
rutas-norte-general-c   min=1  max=8   4 cores / 16 GiB   zone C
rutas-norte-data        min=2  max=3   8 cores / 32 GiB   taint: role=data:NoSchedule

Three design decisions worth understanding:

One group per zone. Necessary so that --balance-similar-node-groups can spread across zones and so that zone affinity works. With a single multi-zone group, the CA cannot guarantee which zone the new node will appear in, and that breaks the topology spread of 09-05.

A dedicated data group with a taint. bookings-postgres and redis-cache have tolerations for role=data (06-05) and node affinity towards that group. That way the databases live on big machines with fast disks and do not share a node with the avalanche of bookings-api pods. The taint stops the application pods from landing there.

The data group has max=3, which is very low. That is deliberate: we do not want the CA booting big, expensive machines by mistake. bookings-postgres does not scale horizontally (09-01), so that group almost never needs to grow.

Labels and taints on the new nodes

When the CA simulates whether a pod would fit on a node from group A, it uses a template. If the real node, on boot, receives labels or taints the template did not know about, the simulation lies.

A typical and very frustrating case: a pod with nodeSelector: disk=ssd is Pending. The CA simulates with the group template, which does not carry that label, concludes that the pod would not fit even on a new node, and does not scale up. But the real nodes in the group do receive that label on boot, through a start-up script.

The fix is to declare the labels and taints in the provider's own node-group definition (on AWS, with k8s.io/cluster-autoscaler/node-template/label/... tags on the ASG), so that the CA knows about them before booting the node.

  1. Removing nodes: the delicate part

Scaling up is easy: there are pending pods, you add nodes. Scaling down is where all the problems are, because removing a node means evicting the pods it holds and trusting that they will reappear somewhere else without breaking anything.

The algorithm

Every 10 seconds, for each node:

  1. Is it below the utilisation threshold? By default, --scale-down-utilization-threshold=0.5: the sum of its pods' requests is less than 50% of its allocatable capacity.
  2. Has it been that way long enough? --scale-down-unneeded-time=10m: ten consecutive minutes below the threshold.
  3. Can it be drained? It simulates whether all its pods would fit on other existing nodes. If they do not fit, it is left alone.
  4. Is there any pod that prevents eviction? The list in section 7.
  5. If everything passes: it cordons the node, evicts its pods respecting the PDBs, waits, and asks the provider to delete it.

Main parameters:

Parameter Default What it controls
--scale-down-enabled true Master switch for scale-down
--scale-down-utilization-threshold 0.5 Underutilisation threshold
--scale-down-unneeded-time 10m Consecutive time before acting
--scale-down-delay-after-add 10m Wait after having added a node
--scale-down-delay-after-delete 0s Wait after having deleted a node
--scale-down-delay-after-failure 3m Wait after a scale-down failure
--max-graceful-termination-sec 600 Maximum wait for a pod to terminate
--max-empty-bulk-delete 10 Empty nodes deletable at once

--scale-down-delay-after-add is the key parameter for avoiding node flapping. Without it, the CA could add a node, see that the cluster is now underutilised and remove it immediately. Ten minutes of grace break that cycle.

About the utilisation threshold, an important calculation:

Default threshold: 50%.

Node with 4 allocatable cores holding these pods:
  - 2 bookings-api pods:  2 x 520m = 1040m
  - 1 web-store pod:                 200m
  - DaemonSets (fluentd, falco, cni): 450m
  Total requests: 1690m out of 3500m allocatable = 48.3%

48.3% < 50%  ->  scale-down candidate.

The CA simulates: do those 4 pods fit on the other nodes?
If yes -> it evicts them and deletes the node.
If no  -> it leaves it alone.

Note that DaemonSets count towards utilisation but do not prevent deletion (they are ignored in step 4, because they disappear with the node). With many DaemonSets, the baseline utilisation of an empty node is already high, and that makes the CA scale down less than you would expect. In clusters with many agents (logging, security, service mesh, monitoring), it is common to have to lower the threshold to 30-40% for scale-down to work at all.

The special case of empty nodes

A node that only holds DaemonSets is considered empty and is deleted via a fast path, without the full simulation. That is why, after a traffic peak, the nodes that end up completely free disappear fairly quickly, while the ones with a stray pod take far longer.

This explains a phenomenon that puzzles a lot of people: after the May bank holiday, eight nodes go away in twenty minutes and two stay for hours with a single pod each. The real fix is not to fiddle with the CA: it is topology spread and consolidation (Karpenter, section 11).

  1. The reasons a node cannot be drained

This is the list to know by heart, because it explains 90% of the cases of "I have empty nodes the autoscaler will not remove and I am paying for them".

# Reason Detail How to fix it
1 Pods without a controller A pod created by hand (with no Deployment, ReplicaSet, Job or StatefulSet) cannot be recreated elsewhere: if you evict it, it is gone for good Always use a controller. Never kubectl run without --restart in production
2 Pods with local storage emptyDir, hostPath or local volumes: the data lives on that node and would be lost --skip-nodes-with-local-storage=false if you accept the loss; or migrate to a network PV
3 A restrictive PodDisruptionBudget The PDB does not allow evicting that pod without dropping below the minimum (09-05) Review the PDB; make sure there are enough replicas
4 System pods in kube-system Critical components with no PDB --skip-nodes-with-system-pods=false (carefully), or give them a PDB
5 The safe-to-evict: "false" annotation An explicit "do not move me" marker Remove the annotation if it no longer applies
6 Pods that do not fit on any other node The simulation fails: the pod is too big or has constraints Add capacity or relax the constraints
7 A node with the scale-down-disabled: "true" annotation Explicit exclusion of the node Remove the annotation

The safe-to-evict annotation

It is the most direct control you have:

apiVersion: v1
kind: Pod
metadata:
  annotations:
    # This pod must NOT be evicted by the Cluster Autoscaler.
    # The node hosting it will never be removed while it is here.
    cluster-autoscaler.kubernetes.io/safe-to-evict: "false"

When "false" makes sense:

  • A long batch job that would lose hours of compute if restarted. occupancy-reports is a candidate: if it takes 40 minutes and the CA kills it at minute 35, the night's work is lost.
  • A database migration in flight.
  • A pod with irreproducible local state.

And the opposite value, "true", serves to unblock situations:

metadata:
  annotations:
    # This pod CAN be evicted even though it uses emptyDir.
    # We know the contents of its emptyDir are a regenerable cache.
    cluster-autoscaler.kubernetes.io/safe-to-evict: "true"

It is the clean way to resolve reason 2 without changing the CA's global flag. In Rutas Norte, web-store uses an emptyDir for the nginx cache: marking it as safe-to-evict: "true" lets the CA remove its nodes with no trouble, because that cache regenerates itself.

The PDB as a blocker (a preview of 09-05)

The most frequent case and the most dangerous. A PDB like this:

apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: bookings-postgres
  namespace: rutas-norte-pro
spec:
  minAvailable: 1
  selector:
    matchLabels:
      app: bookings-postgres

With a single replica of bookings-postgres and minAvailable: 1, that pod can never be evicted: doing so would leave 0 available. The node hosting it is pinned permanently. The CA will try, will fail, and will record it in its events.

In this particular case the block is desirable (we do not want the CA moving the database on its own), but the same pattern applied by mistake to an ordinary service produces zombie nodes that nobody can explain. We develop this fully in 09-05.

  1. Diagnosis: the status ConfigMap and the events

The CA publishes its internal state in a ConfigMap. It is the first place to look.

kubectl get configmap cluster-autoscaler-status -n kube-system -o yaml
apiVersion: v1
kind: ConfigMap
metadata:
  name: cluster-autoscaler-status
  namespace: kube-system
data:
  status: |
    Cluster-autoscaler status at 2026-05-01 10:14:32:
    Cluster-wide:
      Health:      Healthy (ready=7 unready=0 notStarted=1 registered=8 longNotStarted=0)
                   LastProbeTime: 2026-05-01 10:14:31
      ScaleUp:     InProgress (ready=7 registered=8)
                   LastProbeTime: 2026-05-01 10:14:31
      ScaleDown:   NoCandidates (candidates=0)
                   LastProbeTime: 2026-05-01 10:14:31

    NodeGroups:
      Name:        rutas-norte-general-a
      Health:      Healthy (ready=3 unready=0 notStarted=1 registered=4
                   cloudProviderTarget=4 (minSize=1, maxSize=8))
      ScaleUp:     InProgress (ready=3 registered=4)
      ScaleDown:   NoCandidates (candidates=0)

      Name:        rutas-norte-general-b
      Health:      Healthy (ready=2 unready=0 notStarted=0 registered=2
                   cloudProviderTarget=2 (minSize=1, maxSize=8))
      ScaleUp:     NoActivity
      ScaleDown:   NoCandidates (candidates=0)

      Name:        rutas-norte-data
      Health:      Healthy (ready=2 unready=0 notStarted=0 registered=2
                   cloudProviderTarget=2 (minSize=2, maxSize=3))
      ScaleUp:     NoActivity
      ScaleDown:   NoCandidates (candidates=0)

How to read it:

Field Meaning
ready Nodes operational and accepting pods
unready Nodes registered but not ready (a problem)
notStarted Nodes requested that are still booting. This is the one to watch during a peak
registered Total known to the cluster
cloudProviderTarget How many the CA has asked the provider for
ScaleUp: InProgress A scale-up is under way
ScaleDown: NoCandidates No node meets the scale-down criteria

In the output above: group A has 3 ready and 1 booting. The CA has already requested the node; all you have to do is wait. This information is exactly what you need during an incident to know whether the CA is working or stuck.

ScaleDown states you will see:

State Meaning
NoCandidates No node is below the threshold
CandidatesPresent There are candidates, waiting for the unneeded-time
InProgress Draining a node right now

The events: why it is NOT scaling up

When the CA decides not to do anything, it says so in an event on the pending pod:

kubectl get events -n rutas-norte-pro --field-selector reason=NotTriggerScaleUp \
  --sort-by=.lastTimestamp
LAST SEEN   TYPE     REASON              OBJECT                              MESSAGE
23s         Normal   NotTriggerScaleUp   pod/bookings-api-7c9d4f8b6d-qw8vz   pod didn't trigger scale-up:
                                                                             1 max node group size reached,
                                                                             1 node(s) had untolerated taint {role: data}

This message is a gift: it tells you exactly why each node group was discarded. Here: one group is at its maximum and another has a taint the pod does not tolerate.

Common NotTriggerScaleUp messages and what they mean:

Message What it really means Action
max node group size reached The group is at its maxSize Raise the group's maximum
node(s) had untolerated taint {...} The group's nodes carry a taint the pod does not tolerate Add a toleration or create another group
node(s) didn't match Pod's node affinity The nodeSelector/affinity does not match the template Review the group's labels
Insufficient cpu (in the CA's message) Not even a new node from the group would have enough CPU for that pod The pod is too big. Reduce its requests or use a larger group
in backoff after failed scale-up A previous attempt failed (no capacity at the provider, quota exhausted) Look at the CA logs and the provider quota
pod has unbound immediate PersistentVolumeClaims It is a storage problem, not a capacity one Review the StorageClass (module 5)

And for blocked scale-down:

kubectl logs -n kube-system deployment/cluster-autoscaler | grep -i "scale.down\|unremovable"
I0501 10:22:14.882  scale_down.go: Node ip-10-0-3-47 is not suitable for removal:
    pod redis-cache-0 is not replicated and has local storage
I0501 10:22:14.883  scale_down.go: Node ip-10-0-2-19 is not suitable for removal:
    pdb-blocked: not enough pod disruption budget to move bookings-postgres-0
I0501 10:22:14.884  scale_down.go: Node ip-10-0-1-88 is not suitable for removal:
    cluster-autoscaler.kubernetes.io/safe-to-evict annotation set to false

The CA logs are the definitive source when the ConfigMap and the events are not enough. Keep them in EFK (07-05): they are indispensable for the postmortem of a capacity incident.

  1. The real timings: how long a node genuinely takes

Here is the figure that changes the strategy, and the one almost nobody mentions.

Breaking the time down

gantt
    title Time from Pending pod to Running pod on a new node
    dateFormat  s
    axisFormat  %S s

    section Detection
    Pod goes to Pending          :a1, 0, 5s
    The CA detects it (scan)     :a2, after a1, 10s

    section Provisioning
    Request to the provider      :b1, after a2, 5s
    Machine boot                 :b2, after b1, 45s
    Operating system boot        :b3, after b2, 20s

    section Joining the cluster
    kubelet starts and registers :c1, after b3, 15s
    CNI and DaemonSets ready     :c2, after c1, 30s
    Node becomes Ready           :c3, after c2, 5s

    section Pod
    Scheduling                   :d1, after c3, 2s
    Image pull                   :d2, after d1, 40s
    Container start-up           :d3, after d2, 25s
    Readiness probe OK           :d4, after d3, 10s

Adding it up: between 3 and 4 minutes in the good case. In the bad case (a large uncached image, a saturated region, heavy DaemonSets) it can exceed 6 minutes.

Typical breakdown in a table:

Phase Typical time What makes it worse
Detecting the pending pod 10-15 s A high --scan-interval
Request to the provider 5-10 s Retries because of quota
Machine boot 30-60 s Machine type, saturated region
Operating-system boot 15-30 s An unoptimised image
kubelet registration 10-20 s Complex configuration, slow bootstrap
CNI + DaemonSets ready 20-60 s Many heavy DaemonSets
Pulling the pod's image 20-120 s A large image, a distant registry
Container start-up + readiness 20-40 s An application that is slow to start
Total 2-6 min

Why this does not save you from a sudden spike

Let's go back to the May bank holiday. The sale opens at 10:00. The real traffic profile:

10:00:00   The sale opens. Traffic goes from 200 rps to 2400 rps in 40 seconds.
10:00:15   The HPA detects CPU at 280%. It asks for 12 replicas.
10:00:20   The Deployment creates 8 pods. 4 fit on the existing nodes.
           4 are left Pending.
10:00:30   The CA detects the pending pods and asks for 2 nodes.
10:00:45   The HPA evaluates again: still saturated. It asks for 24 replicas.
           More Pending pods.
10:01:00   The CA asks for 3 more nodes.
10:03:30   The first node arrives. 6 pods are scheduled.
10:04:10   The first pods on the new node become Ready.
10:05:00   The other nodes arrive. The platform reaches capacity.

TOTAL DEGRADATION TIME: almost FIVE MINUTES.

Five minutes of degraded latency and errors at the most commercially valuable moment of the year. Thousands of users abandoning their purchase. The Cluster Autoscaler, however well configured, cannot solve a spike that rises in forty seconds, because the physics of booting a virtual machine does not allow it.

Two strategies come out of this, and you need both:

  1. Over-provisioning: having capacity already booted and waiting (section 10).
  2. Anticipatory scaling: scaling before the peak using a time-based trigger rather than a reactive one. That is KEDA's cron trigger, which we will see in 09-04.

  1. Over-provisioning with filler pods

The pattern that solves the problem of the previous section, and which makes brilliant use of the PriorityClass and preemption from 06-05.

The idea

We deploy pods that do nothing — they sleep — but that reserve CPU and memory, with a negative priority. The effects:

  1. The CA counts those requests as occupancy, so it keeps nodes booted to host them.
  2. When a real pod arrives (priority 0 or higher) and there is no room, the scheduler instantly evicts a filler pod and places the real one in its slot. Preemption takes seconds, not minutes.
  3. The evicted filler pods go to Pending, which triggers the CA to boot new nodes... which will refill the cushion for next time.

It is a self-regenerating capacity cushion. You buy idle capacity in exchange for response time.

sequenceDiagram
    participant HPA
    participant Sched as Scheduler
    participant Filler as Filler pods<br/>(priority -10)
    participant CA as Cluster Autoscaler

    Note over Filler: Normal state: 3 filler pods<br/>occupying 3 cores across the nodes
    HPA->>Sched: Create 10 bookings-api pods (priority 0)
    Sched->>Sched: There is no free CPU
    Sched->>Filler: PREEMPTION: evict 3 filler pods
    Note over Sched: Room appears in 2-5 seconds
    Sched->>Sched: Schedule the bookings-api pods
    Filler->>CA: 3 filler pods in Pending
    CA->>CA: Request new nodes (3-4 min)
    Note over Filler: The cushion regenerates on its own

Complete implementation

Step 1: the negative PriorityClass.

# k8s/base/priorityclass-capacity-filler.yaml
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
  name: capacity-filler
# KEY POINT: a NEGATIVE value. Any normal pod (priority 0 by default) has
# more priority than these, so it will evict them without hesitation.
value: -10
# It is not the default class: only pods that ask for it explicitly use it.
globalDefault: false
# The default preemptionPolicy is PreemptLowerPriority, but these pods must
# not evict ANYONE: they are the last in the queue.
preemptionPolicy: Never
description: >
  Filler pods that reserve capacity to absorb traffic spikes.
  They are evicted instantly by any real pod. See 09-03.

Two important details:

  • value: -10 guarantees that any pod without a priorityClassName (priority 0) outranks them.
  • preemptionPolicy: Never means these pods, once Pending, will not try to evict others. It would be absurd for the cushion to throw out a real pod.

Step 2: the filler Deployment.

# k8s/environments/pro/deployment-capacity-filler.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: capacity-filler
  namespace: rutas-norte-pro
  labels:
    app: capacity-filler
    app.kubernetes.io/part-of: rutas-norte
spec:
  # 6 filler pods. Each one reserves 1 core and 1 GiB.
  # Total cushion: 6 cores and 6 GiB, roughly two nodes' worth.
  # Enough to absorb 11 bookings-api replicas instantly.
  replicas: 6
  selector:
    matchLabels:
      app: capacity-filler
  template:
    metadata:
      labels:
        app: capacity-filler
        app.kubernetes.io/part-of: rutas-norte
      annotations:
        # So the CA can freely remove nodes that only hold filler.
        cluster-autoscaler.kubernetes.io/safe-to-evict: "true"
    spec:
      priorityClassName: capacity-filler

      # Immediate termination: there is nothing to save, we do not wait a second.
      # This is what makes preemption almost instantaneous.
      terminationGracePeriodSeconds: 0

      # Spread the cushion across nodes: 6 free cores are useless if they all
      # sit on the same node when the new pods need to spread out.
      topologySpreadConstraints:
        - maxSkew: 1
          topologyKey: kubernetes.io/hostname
          whenUnsatisfiable: ScheduleAnyway
          labelSelector:
            matchLabels:
              app: capacity-filler

      containers:
        - name: pause
          # Minimal official image: a process that only sleeps. ~700 KB.
          image: registry.k8s.io/pause:3.9
          resources:
            requests:
              cpu: "1"
              memory: 1Gi
            limits:
              cpu: "1"
              memory: 1Gi
          # It consumes NOTHING in practice: it only reserves the slot in the
          # scheduler's eyes. Real CPU used is 0m.

The pause image is the canonical choice: it is the same one Kubernetes uses internally for each pod's infrastructure container, it weighs less than a megabyte, it is already cached on every node, and its only job is to sleep.

Step 3: checking that it works.

# Initial state: 6 filler pods spread out
kubectl get pods -n rutas-norte-pro -l app=capacity-filler -o wide
NAME                               READY   STATUS    NODE
capacity-filler-5d8f7c9b4-2xkjp    1/1     Running   ip-10-0-1-42
capacity-filler-5d8f7c9b4-4mnqz    1/1     Running   ip-10-0-1-88
capacity-filler-5d8f7c9b4-7wrtv    1/1     Running   ip-10-0-2-19
capacity-filler-5d8f7c9b4-9hbcx    1/1     Running   ip-10-0-2-73
capacity-filler-5d8f7c9b4-kp3ln    1/1     Running   ip-10-0-3-47
capacity-filler-5d8f7c9b4-tv6ws    1/1     Running   ip-10-0-3-91

Now we force a spike:

kubectl scale deployment bookings-api -n rutas-norte-pro --replicas=20
kubectl get pods -n rutas-norte-pro -l app=capacity-filler --watch
NAME                               READY   STATUS        AGE
capacity-filler-5d8f7c9b4-2xkjp    1/1     Terminating   4h
capacity-filler-5d8f7c9b4-4mnqz    1/1     Terminating   4h
capacity-filler-5d8f7c9b4-7wrtv    0/1     Pending       3s
capacity-filler-5d8f7c9b4-9hbcx    0/1     Pending       3s

And in the events of the bookings-api pods:

Events:
  Type     Reason      Age   From               Message
  ----     ------      ----  ----               -------
  Normal   Preempted   4s    default-scheduler  Preempted by pod
                                                capacity-filler-5d8f7c9b4-2xkjp
                                                on node ip-10-0-1-42
  Normal   Scheduled   3s    default-scheduler  Successfully assigned
                                                rutas-norte-pro/bookings-api-... to ip-10-0-1-42

Three seconds from the pod being created to being scheduled, instead of three and a half minutes. That is the difference between absorbing the opening of the sale and falling over.

The cost

Nothing is free:

Cushion: 6 cores and 6 GiB reserved permanently.
Equivalent to ~1.7 nodes of 4 cores.

Estimated cost of a node (4 cores, 16 GiB, on-demand): 95 EUR/month.
Cost of the cushion: 1.7 x 95 = 162 EUR/month = 1,940 EUR/year.

Benefit: instantly absorbing 11 bookings-api replicas.

Is it worth it? At Rutas Norte the answer is clear: five minutes of outage at the opening of the May bank-holiday sale cost far more than €1,940 in unsold tickets. But it is a business decision, not a technical one, and it must be taken with the numbers on the table.

An obvious optimisation: the cushion does not have to be constant. A CronJob (06-03) that scales the filler Deployment from 2 to 6 replicas on Friday afternoons and returns it to 2 on Mondays reduces the annual cost to a fraction:

# k8s/environments/pro/cronjob-adjust-filler.yaml
apiVersion: batch/v1
kind: CronJob
metadata:
  name: expand-filler-for-weekend
  namespace: rutas-norte-pro
spec:
  schedule: "0 16 * * 5"       # Friday at 16:00
  jobTemplate:
    spec:
      template:
        spec:
          serviceAccountName: scaling-operator
          restartPolicy: OnFailure
          containers:
            - name: kubectl
              image: bitnami/kubectl:1.30
              command:
                - kubectl
                - scale
                - deployment/capacity-filler
                - --replicas=6
                - -n
                - rutas-norte-pro

And its counterpart on Monday at 6:00 with --replicas=2. The scaling-operator ServiceAccount needs a Role with permission over deployments/scale, exactly as we saw in 08-01.

  1. Karpenter: the modern alternative

The Cluster Autoscaler has a design limitation: it works with predefined node groups. Someone has to decide in advance which machine types exist, create a group for each of them, and maintain them.

Karpenter turns the approach around: it looks at the pending pods, works out which machine would be best for them, and boots it directly, with no groups.

How it works

Instead of node groups, you define constraints on which machine types are acceptable:

# NodePool: the decision space Karpenter is allowed to explore
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
  name: rutas-norte-general
spec:
  template:
    metadata:
      labels:
        environment: pro
        app.kubernetes.io/part-of: rutas-norte
    spec:
      requirements:
        - key: kubernetes.io/arch
          operator: In
          values: ["amd64"]
        - key: karpenter.sh/capacity-type
          operator: In
          values: ["spot", "on-demand"]     # Prefers spot; falls back to on-demand
        - key: karpenter.k8s.aws/instance-category
          operator: In
          values: ["c", "m", "r"]           # Compute, general-purpose and memory families
        - key: karpenter.k8s.aws/instance-size
          operator: NotIn
          values: ["nano", "micro", "small"]
        - key: topology.kubernetes.io/zone
          operator: In
          values: ["eu-west-1a", "eu-west-1b", "eu-west-1c"]
      nodeClassRef:
        group: karpenter.k8s.aws
        kind: EC2NodeClass
        name: rutas-norte

  # CONSOLIDATION: the flagship feature. Karpenter continuously re-evaluates
  # whether the current pods would fit on fewer or cheaper nodes, and migrates.
  disruption:
    consolidationPolicy: WhenEmptyOrUnderutilized
    consolidateAfter: 1m
    budgets:
      - nodes: "20%"           # At most, touch 20% of the nodes at a time

  limits:
    cpu: "200"                 # Global safety ceiling
    memory: 800Gi

With this definition, when 15 bookings-api pods turn up pending, Karpenter works out that the best machine to host them is, say, a c6i.4xlarge spot instance in zone B, and boots it. Without anyone ever having created a c6i.4xlarge node group.

Consolidation

This is the most practical difference day to day. Karpenter continuously checks whether the current set of pods would fit on a cheaper node configuration, and if so, it migrates the pods and replaces the nodes.

The scenario it solves: after the May bank holiday, the CA leaves five nodes with one pod each, because none drops below the utilisation threshold in a way that allows draining. Karpenter sees that those five pods would fit on a single node, boots that node, migrates the pods and deletes the five.

It is exactly the problem we mentioned at the end of section 6, solved at the root.

Comparison

Aspect Cluster Autoscaler Karpenter
Unit of work Predefined node groups Individual nodes
Choosing the machine type Fixed per group Computed from the pending pods
Prior configuration One group per type/zone One NodePool with constraints
Provisioning time 3-6 min 1-2 min (talks directly to the EC2 API)
Consolidation No (only utilisation-based scale-down) Yes, continuous
Cost optimisation Limited (the price expander) Native: picks the cheapest type that works
Spot handling Separate groups + the priority expander Native, with automatic fallback to on-demand
Provider support Many (AWS, Azure, GCP, OpenStack, Cluster API...) AWS mature; Azure available; others in development
Maturity Very high, years in production High on AWS, more recent
Operational complexity Medium (maintaining groups) Low after the initial setup
Predictability High: you know which machines will appear Lower: it can pick unexpected types

When to choose each one:

  • Cluster Autoscaler if you are outside AWS, if you have regulatory requirements about which machine types are admissible, if your provider already gives it to you configured and working, or if you value predictability over savings.
  • Karpenter if you are on AWS, if cost is a genuine priority, if your workload is heterogeneous (small and large pods living together), and if the reduction in operational work pays off for you.

For Rutas Norte, with four nodes and a homogeneous workload, the CA is more than enough. If the platform grew to fifty nodes with varied workloads, Karpenter would be the obvious decision. We will come back to this kind of decision in 10-06 and 11-06.

An important note the comparison must not hide: Karpenter does not fix the forty-second spike either. One or two minutes is still too long. Over-provisioning is still necessary.

  1. How the three autoscalers fit together

We now have all three pieces. Let's see the full flow in time order.

flowchart TD
    T0["t=0 s<br/>Traffic rises"] --> HPA["t=15 s<br/>HPA: average CPU exceeds the target<br/>Computes desired replicas"]
    HPA --> DEP["t=16 s<br/>Deployment creates new pods"]
    DEP --> SCHED{"t=17 s<br/>Does the scheduler<br/>find room?"}

    SCHED -->|Yes| RUN["t=20-50 s<br/>Pod Running and Ready<br/>END: absorbed"]

    SCHED -->|No| PRE{"Are there filler<br/>pods to evict?"}
    PRE -->|Yes| PREEMPT["t=20 s<br/>PREEMPTION<br/>Real pod placed in 3 s"]
    PREEMPT --> RUN2["t=50 s<br/>Pod Running<br/>END: absorbed with the cushion"]
    PREEMPT --> PEND2["The filler pods<br/>go to Pending"]

    PRE -->|No| PEND["Pod in Pending<br/>FailedScheduling event"]
    PEND2 --> CA
    PEND --> CA["t=30 s<br/>Cluster Autoscaler detects<br/>the unschedulable pods"]

    CA --> SIM["Simulates: which node group<br/>would let them be scheduled?<br/>Applies the expander"]
    SIM --> REQ["t=40 s<br/>Request to the cloud provider"]
    REQ --> BOOT["t=40 s to 3 min<br/>Boot, joining the cluster,<br/>CNI and DaemonSets"]
    BOOT --> READY["t=3 min<br/>Node Ready"]
    READY --> SCHED2["t=3 min 5 s<br/>The pending pods are scheduled"]
    SCHED2 --> PULL["t=3 to 4 min<br/>Image pull and start-up"]
    PULL --> RUN3["t=4 min<br/>Pod Running<br/>END: absorbed with a new node"]

    style RUN fill:#cfc,stroke:#393
    style RUN2 fill:#cfc,stroke:#393
    style RUN3 fill:#ffd,stroke:#c90
    style PEND fill:#fdd,stroke:#c00

The three routes and their timings:

Route When it happens Time until it serves traffic
Fits on the current nodes There is free capacity 20-50 s
Preemption of the cushion It does not fit, but there are filler pods 30-60 s
New node It does not fit and there is no cushion 3-6 min

The strategic conclusion is direct: the over-provisioning cushion is not a luxury, it is what turns a five-minute incident into a thirty-second one.

And there is still a better route, which does not appear in the diagram because it is not reactive: scaling before the traffic arrives. If at 9:30 you already have 25 replicas and eight nodes because a cron trigger ordered it, at 10:00 there is no incident at all. That is 09-04.

A summary of the three levels

Level What changes Trigger Latency Component
1 Number of pods A CPU/business metric 15-60 s HPA (09-01) / KEDA (09-04)
2 Size of each pod Consumption history Hours-days VPA (09-02)
3 Number of nodes Unschedulable pods 3-6 min Cluster Autoscaler / Karpenter

All three are necessary and none replaces another. The HPA without the CA has a hard ceiling. The CA without the HPA never fires. The VPA makes both more accurate, because correct requests are the foundation of both the HPA's calculation and the CA's simulation.

  1. What can be practised on minikube

Time for honesty: the Cluster Autoscaler cannot really be tested on minikube. There is no cloud provider to boot machines. The rutas-norte profile has the nodes you asked for when you created it, and that is all of them.

What you CAN practise

1. Producing the symptom and learning to read it.

# A cluster with 2 small nodes
minikube start -p rutas-norte --nodes=2 --cpus=2 --memory=4096

# Deploy something that does not fit
kubectl -n rutas-norte-dev create deployment occupier \
  --image=registry.k8s.io/pause:3.9 --replicas=10

kubectl -n rutas-norte-dev set resources deployment/occupier \
  --requests=cpu=800m,memory=512Mi

kubectl get pods -n rutas-norte-dev
NAME                        READY   STATUS    RESTARTS   AGE
occupier-6f8d7c9b4-2mkjp    1/1     Running   0          15s
occupier-6f8d7c9b4-4nqzx    1/1     Running   0          15s
occupier-6f8d7c9b4-7wtrv    0/1     Pending   0          15s
occupier-6f8d7c9b4-9bhcx    0/1     Pending   0          15s
...
kubectl describe pod -n rutas-norte-dev -l app=occupier | grep -A5 Events

You will see exactly the FailedScheduling with Insufficient cpu. It is the same message as in production.

2. The over-provisioning pattern with preemption. This works perfectly, because preemption is a function of the scheduler, not of the cloud provider.

# 1. Create the negative PriorityClass
kubectl apply -f k8s/base/priorityclass-capacity-filler.yaml

# 2. Deploy the filler (sized for minikube)
kubectl -n rutas-norte-dev create deployment filler \
  --image=registry.k8s.io/pause:3.9 --replicas=2
kubectl -n rutas-norte-dev set resources deployment/filler \
  --requests=cpu=500m,memory=256Mi
kubectl -n rutas-norte-dev patch deployment filler \
  -p '{"spec":{"template":{"spec":{"priorityClassName":"capacity-filler","terminationGracePeriodSeconds":0}}}}'

# 3. Check that the filler is taking up room
kubectl get pods -n rutas-norte-dev -o wide

# 4. Create a real pod that would not fit without evicting the filler
kubectl -n rutas-norte-dev run important-pod \
  --image=registry.k8s.io/pause:3.9 \
  --overrides='{"spec":{"containers":[{"name":"important-pod","image":"registry.k8s.io/pause:3.9","resources":{"requests":{"cpu":"900m"}}}]}}'

# 5. Watch the preemption live
kubectl get events -n rutas-norte-dev --sort-by=.lastTimestamp | tail -20
LAST SEEN   TYPE     REASON      OBJECT                      MESSAGE
3s          Normal   Preempted   pod/filler-7d8c9f4b6-2xkjp  Preempted by pod
                                                             rutas-norte-dev/important-pod on node rutas-norte-m02
2s          Normal   Scheduled   pod/important-pod           Successfully assigned to rutas-norte-m02
1s          Normal   Scheduled   pod/filler-7d8c9f4b6-9wqzn  FailedScheduling

Seeing the Preempted with your own eyes and measuring that it happens in seconds is the most valuable part of the exercise, and you can practise it without a cloud.

3. Adding and removing nodes by hand, simulating what the CA would do.

# Add a node to the profile (equivalent to what the CA does when scaling up)
minikube -p rutas-norte node add

# See how the Pending pods get scheduled on their own
kubectl get pods -n rutas-norte-dev --watch

# Remove a node (equivalent to scale-down, draining included)
kubectl cordon rutas-norte-m03
kubectl drain rutas-norte-m03 --ignore-daemonsets --delete-emptydir-data
minikube -p rutas-norte node delete rutas-norte-m03

That cordon → drain → delete sequence is exactly what the CA does internally. Practising it by hand teaches you more about node removal than reading any documentation, and it is the basis of the maintenance procedure we will see in 09-05.

4. Practising the safe-to-evict annotation and seeing how it blocks a drain.

kubectl annotate pod important-pod -n rutas-norte-dev \
  cluster-autoscaler.kubernetes.io/safe-to-evict=false

# Note: this annotation is honoured by the CA, not by kubectl drain. To see a
# drain really blocked you need a PDB, which is what we will see in 09-05.

What you CANNOT practise

Not practicable Alternative
Automatic node scale-up minikube node add by hand
Expanders and group selection Study the documentation and the logs of a real cluster
Automatic node scale-down cordon + drain + node delete by hand
The cluster-autoscaler-status ConfigMap It does not exist without the CA installed
Karpenter Requires AWS

To practise the CA for real you need: a managed cluster (10-06) with autoscaling enabled, or kind with the CA's test provider, or a personal cloud environment. A practical recommendation: a trial account with a provider for one afternoon, with a small cluster and a low maxSize, costs a few euros and teaches an enormous amount.

  1. The Rutas Norte capacity plan and its cost

We close with the concrete numbers. This is the work that has to be done before a May bank holiday, and that was not done last year.

Step 1: peak demand

From the maxReplicas of 09-01 and the requests recalibrated with the VPA in 09-02:

Component maxReplicas CPU/pod Mem/pod Max CPU Max mem
bookings-api (api) 30 412m 800Mi 12.36 23.4 GiB
bookings-api (exporter) 30 6m 21Mi 0.18 0.6 GiB
web-store 15 70m 64Mi 1.05 0.9 GiB
notifications-worker 20 120m 410Mi 2.40 8.0 GiB
bookings-postgres 1 2000m 4Gi 2.00 4.0 GiB
redis-cache 1 300m 2Gi 0.30 2.0 GiB
Ingress controller 4 200m 256Mi 0.80 1.0 GiB
DaemonSets (per node) — 450m 512Mi (var.) (var.)
Application subtotal 19.09 39.9 GiB

Step 2: translating into nodes

General node: 4 cores, 16 GiB.
Allocatable after kubelet and system reservations: ~3.5 cores, ~14 GiB.
DaemonSets per node: 450m and 512Mi.
Usable capacity per node: 3.05 cores and 13.5 GiB.

Nodes by CPU:    19.09 / 3.05 = 6.26  ->  7 nodes
Nodes by memory: 39.9  / 13.5 = 2.96  ->  3 nodes

CPU is the limit: 7 general nodes are needed at the peak.

Also: bookings-postgres and redis-cache live in the "data" group with a taint.
  postgres: 2 cores, 4 GiB.  redis: 0.3 cores, 2 GiB.
  Total: 2.3 cores and 6 GiB.  They fit on 1 node of 8 cores, but we keep
  2 for availability (one per zone).

Over-provisioning cushion: 6 cores = 2 additional nodes.

TOTAL AT THE PEAK: 9 general nodes + 2 data nodes = 11 nodes.

Step 3: configuring the groups

rutas-norte-general-a   min=1  max=4    (was max=8, we adjust: 3 zones x 4 = 12 > 9)
rutas-norte-general-b   min=1  max=4
rutas-norte-general-c   min=1  max=4
rutas-norte-data        min=2  max=3

Total maximum capacity: 12 general nodes + 3 data nodes.
Margin over the calculation: 33%.

The 33% margin is not a whim: it covers estimation errors, a deployment in progress during the peak (which temporarily doubles the pods of one component), and the loss of an entire zone.

Step 4: the costs

Fictitious prices, but of a realistic order of magnitude:

Item Amount Unit cost/month Total/month
Base general nodes (3, one per zone) 3 €95 €285
Data nodes (2) 2 €190 €380
Permanent over-provisioning cushion 2 €95 €190
Permanent subtotal €855
Extra nodes during the bank holiday (6 nodes × 3 days) 6×3 days €3.17/day €57
Total for the bank-holiday month €912

And the comparison that justifies the whole module:

Strategy Annual cost Capacity at the peak Risk
Fixed, sized for a Tuesday (the starting point) €8,000 Insufficient The platform falls over. The 2023 outage
Fixed, sized for the peak €20,500 Sufficient None, but you pay €12,500 of idle capacity 362 days a year
Autoscaling + a permanent cushion €10,260 + €57 Sufficient Low
Autoscaling + a cushion only in high season €8,550 + €57 Sufficient Low

The last row is the goal: practically the same cost as the insufficient fixed sizing, with the capacity of the peak sizing. That is the entire value of module 9 summed up in one line.

And a reminder of intellectual honesty: the reduced off-season cushion only works if the CronJob that expands it actually runs. Put an alert on it: if at 16:05 on Friday the filler Deployment does not have 6 replicas, somebody has to find out.

Common Mistakes and Tips

Mistake 1: believing the HPA can take you further than the capacity allows, and not noticing. The symptom is silent: kubectl get hpa shows 30 replicas and everything looks fine. Only kubectl get pods reveals that 15 are Pending. A mandatory alert on pending pods:

sum(kube_pod_status_phase{phase="Pending", namespace="rutas-norte-pro"}) > 0

Mistake 2: HPA maxReplicas that is inconsistent with the node groups' maxSize. A maxReplicas: 100 with groups that top out at 8 nodes is a lie: you will never reach 100 replicas. Always compute Σ(maxReplicas × requests) and compare it with Σ(maxSize × allocatable capacity). Document that calculation next to the manifests.

Mistake 3: pods without requests in a cluster with a CA. A pod without requests is invisible to the CA's simulation: the autoscaler believes a node has room to spare when in fact it is saturated with pods consuming without declaring anything. The LimitRange of module 3 is the defence: it forces a defaultRequest onto everything that gets created.

Mistake 4: expecting the CA to save you from a sudden spike. Three to six minutes. If your traffic rises in forty seconds, the CA arrives late. Over-provisioning (section 10) and anticipatory scaling (09-04). There is no third way.

Mistake 5: impossible PDBs that pin nodes forever. A minAvailable equal to the replica count means no pod can ever be evicted, and the node hosting them will never be removed. You will be paying for underutilised nodes without knowing why. We develop this in 09-05, but the diagnosis is in the CA logs: pdb-blocked.

Mistake 6: a single node group for everything. Without separating the data from the application, an avalanche of bookings-api pods can land on the same node as bookings-postgres and compete for its CPU. Taints and separate groups (06-05).

Mistake 7: heterogeneous nodes inside the same group. The CA uses one template per group. If the group mixes 2-core and 8-core machines, the simulation lies and the decisions are erratic. One machine type per group, always.

Tip 1: --balance-similar-node-groups if you have one group per zone. Without it, the CA can pile all the new nodes into one zone and leave you exposed to a zone outage. It is one line of configuration with a direct impact on availability.

Tip 2: monitor cluster_autoscaler_unschedulable_pods_count. The CA exposes Prometheus metrics. The ones that matter:

# Pods the CA cannot manage to place
cluster_autoscaler_unschedulable_pods_count > 0

# Nodes per group, to see how full the groups are
cluster_autoscaler_nodes_count

# Scale-up errors (no capacity at the provider, quota exhausted)
rate(cluster_autoscaler_failed_scale_ups_total[15m]) > 0

That third indicator is especially valuable: a failed_scale_up means the provider has said no. Quota exhausted, an instance type unavailable in the zone, an account limit. These are problems you only discover when it is too late, unless you watch for them.

Tip 3: rehearse the scaling before the peak. A week before the bank holiday, run a load test against rutas-norte-pre (09-06) and time how long the cluster takes to reach maximum capacity. That number, measured rather than estimated, is what tells you whether the cushion is enough.

Tip 4: raise the groups' maxSize before high season and lower it afterwards. It is a one-minute change that avoids max node group size reached at the worst possible moment. Put it on the pre-event checklist.

Tip 5: keep the CA logs. They are the only source that explains why it did not scale up or why a node was not removed. Without them, the postmortem of a capacity incident is guesswork. EFK (07-05).

Tip 6: review the scale-down threshold if you have many DaemonSets. With Fluentd, Falco, the CNI, the node exporter and perhaps a service-mesh proxy, the baseline utilisation of an empty node can be around 25%. With the default 50% threshold, scale-down almost never fires. Lowering it to 35% can save entire nodes.

Exercises

Exercise 1: computing the capacity and spotting the error

The Rutas Norte team has configured this for the May bank holiday:

HPAs:

Component minReplicas maxReplicas requests.cpu requests.memory
bookings-api 4 40 520m 820Mi
web-store 3 20 70m 64Mi
notifications-worker 2 25 120m 410Mi

Fixed components: bookings-postgres (1 replica, 2 cores, 4Gi), redis-cache (1 replica, 300m, 2Gi), Ingress (4 replicas, 200m, 256Mi).

Node groups:

rutas-norte-general-a   min=1  max=3   4 cores / 16 GiB
rutas-norte-general-b   min=1  max=3   4 cores / 16 GiB
rutas-norte-data        min=1  max=2   8 cores / 32 GiB   taint role=data

DaemonSets: 450m and 512Mi per node. System reservation: 500m and 2Gi per node.

Namespace ResourceQuota: requests.cpu: 40, requests.memory: 80Gi.

Compute: (a) the maximum CPU and memory the application can demand; (b) the real maximum allocatable capacity of the groups; (c) whether the configuration survives the peak; (d) the percentage of bookings-api's maxReplicas that is achievable in practice; (e) what you would fix and in what order.

Exercise 2: diagnosing why a node is not removed

After the May bank holiday, the cluster still has 9 nodes even though traffic returned to normal 5 hours ago. The bill is a concern. This is what you see:

kubectl get configmap cluster-autoscaler-status -n kube-system -o yaml | grep -A4 "ScaleDown"
      ScaleDown:   NoCandidates (candidates=0)
kubectl describe nodes | grep -A6 "Allocated resources"
Node ip-10-0-1-42:  cpu 2890m (82%)  memory 7Gi (50%)
Node ip-10-0-1-88:  cpu 1210m (34%)  memory 3Gi (21%)
Node ip-10-0-2-19:  cpu 2650m (75%)  memory 6Gi (43%)
Node ip-10-0-2-73:  cpu  980m (28%)  memory 2Gi (14%)
Node ip-10-0-3-47:  cpu 1450m (41%)  memory 4Gi (28%)
Node ip-10-0-3-91:  cpu 2100m (60%)  memory 5Gi (36%)
Node ip-10-0-1-15:  cpu  760m (21%)  memory 2Gi (14%)
Node ip-10-0-2-88:  cpu 1890m (54%)  memory 5Gi (36%)
Node ip-10-0-3-22:  cpu  890m (25%)  memory 2Gi (14%)
kubectl logs -n kube-system deployment/cluster-autoscaler --tail=50 | grep -i "not suitable"
scale_down.go: Node ip-10-0-1-88 is not suitable for removal:
    pod rutas-norte-pro/occupancy-reports-28912440-x7kqp has
    cluster-autoscaler.kubernetes.io/safe-to-evict annotation set to false
scale_down.go: Node ip-10-0-2-73 is not suitable for removal:
    pod rutas-norte-pro/redis-cache-0 is not replicated and has local storage
scale_down.go: Node ip-10-0-1-15 is not suitable for removal:
    pdb-blocked: not enough pod disruption budget to move bookings-api-7c9d4f8b6d-mn2vp
scale_down.go: Node ip-10-0-3-22 is not suitable for removal:
    pod rutas-norte-pro/load-generator-5f8d9c-2wxkp is not replicated

Analyse each blocked node, say whether the block is legitimate or a mistake, and propose the fix for each case. Compute how many nodes could be removed after the fixes and the monthly saving (€95/node/month).

Exercise 3: designing the strategy for a flash campaign

Rutas Norte is launching a Black Friday campaign unlike anything before:

  • Duration: exactly 2 hours, from 20:00 to 22:00 on a Friday.
  • Expected traffic: x25 compared with a normal day (far more than the May bank holiday).
  • The peak arrives in 20 seconds: it is announced on the radio and everyone comes in at once.
  • Tickets are limited: if the platform does not respond in the first 10 minutes, the campaign has failed.
  • Budget: up to €400 for that night.
  • It is known three weeks in advance.

Design the complete capacity strategy. What combination of HPA, CA, over-provisioning and scheduled scaling would you use? How many nodes, how much cushion, how far in advance? What would it cost? Write the key manifests and a checklist with concrete timings for that night.


Solutions

Solution 1

(a) Maximum demand from the application:

Component Max replicas CPU/pod Total CPU Mem/pod Total mem
bookings-api 40 520m 20.80 820Mi 32.03 GiB
web-store 20 70m 1.40 64Mi 1.25 GiB
notifications-worker 25 120m 3.00 410Mi 10.01 GiB
Ingress 4 200m 0.80 256Mi 1.00 GiB
General subtotal 26.00 44.29 GiB
bookings-postgres 1 2000m 2.00 4Gi 4.00 GiB
redis-cache 1 300m 0.30 2Gi 2.00 GiB
Data subtotal 2.30 6.00 GiB

(b) Real maximum allocatable capacity:

General node (4 cores, 16 GiB):
  System reservation:   500m, 2 GiB
  DaemonSets:           450m, 512Mi
  Available for pods:   4 - 0.5 - 0.45 = 3.05 cores
                        16 - 2 - 0.5 = 13.5 GiB

General groups: a(max=3) + b(max=3) = 6 nodes
  CPU:    6 x 3.05 = 18.30 cores
  Memory: 6 x 13.5 = 81.00 GiB

Data node (8 cores, 32 GiB):
  Available: 8 - 0.5 - 0.45 = 7.05 cores
             32 - 2 - 0.5 = 29.5 GiB

Data group (max=2): 2 nodes
  CPU:    14.10 cores
  Memory: 59.00 GiB

(c) Does it survive the peak? NO.

GENERAL:
  Needed:    26.00 cores
  Available: 18.30 cores
  SHORTFALL: 7.70 cores (30% of what is needed)

  Memory: needed 44.29 GiB, available 81 GiB. Plenty.
  -> CPU IS THE LIMIT.

DATA:
  Needed:    2.30 cores and 6 GiB
  Available: 14.10 cores and 59 GiB
  -> Enormously oversized. The data group is VERY over-provisioned.

And there is a second problem many people would overlook: the ResourceQuota.

Total requests.cpu demand: 26.00 + 2.30 = 28.30
Quota: 40. It fits (71% used).

Total requests.memory demand: 44.29 + 6.00 = 50.29 GiB
Quota: 80 GiB. It fits (63% used).

The quota is NOT the problem here, but note that during a rolling update of
bookings-api the old and new pods coexist, which can add up to 25% more:
28.30 x 1.25 = 35.4 cores. That is 88% of the quota. Tight.

(d) Percentage of bookings-api's maxReplicas that is achievable:

CPU available in the general group: 18.30 cores.
Consumption of the other general components (at their maximum):
  web-store 1.40 + worker 3.00 + ingress 0.80 = 5.20 cores

CPU available for bookings-api: 18.30 - 5.20 = 13.10 cores
Achievable replicas: 13.10 / 0.52 = 25.19  ->  25 replicas

25 out of 40 = 62.5%

The maxReplicas: 40 is fiction: the reality is 25 replicas. And worse: when the HPA asks for replicas 26 to 40, they will sit in Pending and nobody will find out unless there is an alert.

(e) Fixes, in order of priority:

Priority 1 (essential): raise the maxSize of the general groups.

Needed: 26.00 cores / 3.05 per node = 8.52  ->  9 nodes
With a third group per zone (advisable for 09-05): 3 groups x max=4 = 12 nodes
  Capacity: 12 x 3.05 = 36.6 cores. A 41% margin over what is needed.
rutas-norte-general-a   min=1  max=4
rutas-norte-general-b   min=1  max=4
rutas-norte-general-c   min=1  max=4     <- new group, third zone

Priority 2: reduce the maxSize of the data group.

The data group needs 2.30 cores and 6 GiB. A single 8-core node is more than
enough. With max=2 we keep availability (one per zone) but there is no reason
to allow more. It stays at max=2, which is already correct.

BUT: the machine type is oversized. A node of 4 cores and 16 GiB would be
enough for postgres (2 cores, 4 GiB) and redis (0.3, 2 GiB). Changing the type
from 8/32 to 4/16 would save ~95 EUR/month per node, 190 EUR/month in total.

A caveat: postgres benefits from memory for shared_buffers and the page cache.
Before reducing, measure with the VPA (09-02) whether 16 GiB is enough. With
requests of 4 GiB and a 16 GiB node, there is 10 GiB of system cache: probably yes.

Priority 3: review bookings-api's maxReplicas: 40.

With 12 general nodes there is room for 25 replicas... wait, let's recompute with the new capacity:

New general capacity: 36.6 cores
Minus the other components: 36.6 - 5.20 = 31.4 cores
Achievable bookings-api replicas: 31.4 / 0.52 = 60

Now the 40 do fit. The maxReplicas: 40 is achievable.

Priority 4: add the over-provisioning cushion. Without it, replicas 5 to 40 will take between 3 and 6 minutes to appear. With 6 filler pods of 1 core, the first 11 replicas are instantaneous.

Priority 5: the pending-pods alert. Without it, this whole calculation could be wrong and nobody would know until it was too late.

Solution 2

Node-by-node analysis:

Node Utilisation State Reason for the block Legitimate?
ip-10-0-1-42 82% Busy Above the threshold Yes, not a candidate
ip-10-0-1-88 34% Blocked safe-to-evict: false on occupancy-reports It depends
ip-10-0-2-19 75% Busy Above the threshold Yes
ip-10-0-2-73 28% Blocked redis-cache-0, unreplicated and with local storage Yes, legitimate
ip-10-0-3-47 41% Candidate — It should be removed
ip-10-0-3-91 60% Busy Above the threshold Yes
ip-10-0-1-15 21% Blocked bookings-api PDB NO, it is a mistake
ip-10-0-2-88 54% Busy Above the threshold Yes
ip-10-0-3-22 25% Blocked load-generator with no controller NO, it is junk

Case 1: ip-10-0-1-88 — occupancy-reports with safe-to-evict: false.

Legitimacy: conditional. If the Job is running right now, the block is correct: killing it would lose the work. But the bank holiday ended 5 hours ago and occupancy-reports is a nightly CronJob.

# Is the Job still alive?
kubectl get jobs -n rutas-norte-pro
kubectl get pods -n rutas-norte-pro -l app=occupancy-reports

If it shows as Running and has been for hours, it is a stuck Job: it probably failed and got wedged. The fix:

# 1. Investigate why it is still alive
kubectl logs -n rutas-norte-pro occupancy-reports-28912440-x7kqp --tail=50

# 2. If it is stuck, delete it
kubectl delete job occupancy-reports-28912440 -n rutas-norte-pro

And a structural prevention in the CronJob:

spec:
  jobTemplate:
    spec:
      # If it has not finished in 2 hours, it is cut off. Avoids zombie jobs
      # that pin nodes indefinitely.
      activeDeadlineSeconds: 7200
      ttlSecondsAfterFinished: 3600

Node recoverable: YES, after cleaning up the Job.

Case 2: ip-10-0-2-73 — redis-cache-0.

Legitimacy: yes, completely. redis-cache is a single-replica StatefulSet with a volume. The CA cannot move it safely, and it certainly must not do so on its own initiative: restarting the cache dumps the whole load onto bookings-postgres.

Fix: none in the short term. In the medium term, the options are:

  • Pin redis-cache to the data node group with affinity and a taint (06-05), so that it does not occupy a general node.
  • Accept it: a node dedicated to the cache is not waste if it is sized properly.

The best fix is the first: redis-cache should not be on a general node. It is an architectural decision we should already have taken.

Node recoverable: NOT directly, but yes once redis-cache is moved to the data group.

Case 3: ip-10-0-1-15 — the PDB blocking bookings-api.

Legitimacy: no, it is a configuration error. bookings-api is a Deployment with many replicas: moving one pod should be trivial. That the PDB prevents it means the PDB is wrong.

kubectl get pdb -n rutas-norte-pro bookings-api -o yaml

The likely suspect:

spec:
  minAvailable: 30        # <-- Set during the bank holiday, when there were 30 replicas

After the bank holiday, the HPA has come down to 4-6 replicas. With minAvailable: 30 and 6 live replicas, no pod can ever be evicted: we are already below the minimum.

The fix:

apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: bookings-api
  namespace: rutas-norte-pro
spec:
  # A PERCENTAGE, not an absolute number. It adapts to the HPA's scaling on its own.
  # With 6 replicas it allows evicting 1; with 30, it allows evicting 6.
  maxUnavailable: 20%
  selector:
    matchLabels:
      app: bookings-api

This is the central lesson of 09-05, previewed: with an HPA, absolute-number PDBs are a trap. The replica count changes on its own, and a minimum that was reasonable at the peak becomes impossible in the trough.

Node recoverable: YES, immediately after fixing the PDB.

Case 4: ip-10-0-3-22 — load-generator with no controller.

Legitimacy: no, it is junk. It is the load-test pod from section 12 of 09-01, created with kubectl run without a controller and then forgotten. It has spent days there pinning an entire node.

kubectl delete pod load-generator-5f8d9c-2wxkp -n rutas-norte-pro

Prevention:

  • Always use --rm with kubectl run for tests.
  • A Kyverno policy (08-03) that rejects pods without ownerReferences in rutas-norte-pro.
  • A periodic review of orphan pods:
kubectl get pods -n rutas-norte-pro -o json | \
  python3 -c "import sys,json; [print(p['metadata']['name']) for p in json.load(sys.stdin)['items'] if not p['metadata'].get('ownerReferences')]"

Node recoverable: YES, immediately.

Summary and saving:

Node Action Removed?
ip-10-0-1-88 Delete the stuck Job Yes
ip-10-0-2-73 Move redis-cache to the data group (medium term) Yes, in the medium term
ip-10-0-3-47 None: it is a candidate, it just needs time Yes
ip-10-0-1-15 Fix the PDB to maxUnavailable: 20% Yes
ip-10-0-3-22 Delete the orphan pod Yes
Nodes removable in the short term: 4 (ip-10-0-1-88, -3-47, -1-15, -3-22)
Nodes removable in the medium term: 1 more (ip-10-0-2-73)

Immediate saving:    4 x 95 EUR = 380 EUR/month = 4,560 EUR/year
Medium-term saving:  5 x 95 EUR = 475 EUR/month = 5,700 EUR/year

Verification after the fixes:

# Wait out the scale-down-unneeded-time (10 min by default) and verify
watch -n 30 'kubectl get configmap cluster-autoscaler-status -n kube-system \
  -o jsonpath="{.data.status}" | grep -A2 ScaleDown'

It should go from NoCandidates to CandidatesPresent and then to InProgress.

The underlying lesson of the exercise: zombie nodes are never the Cluster Autoscaler's fault. They are always the consequence of something misconfigured in the workloads: an impossible PDB, a stuck Job, an orphan pod, a stateful component in the wrong place. The CA is the messenger; the logs are the message.

Solution 3

Analysing the scenario. This case is qualitatively different from the May bank holiday:

Factor May bank holiday Black Friday
Duration 3 days 2 hours
Multiplier x10 x25
Rise speed Minutes 20 seconds
Notice Known (the calendar) 3 weeks
Tolerance for failure Low None: 10 minutes and it is over

The key conclusion: reactive scaling is NO USE here. Neither the HPA (15 s of detection + 30 s of start-up) nor the CA (3-6 min) arrives in time for a rise of 20 seconds. All the capacity has to be booted and warm before 20:00.

Strategy: total pre-scaling.

Step 1: compute the capacity needed.

Normal Friday-night traffic: ~250 rps.
Expected traffic: 250 x 25 = 6,250 rps.

From the load test (09-06) we know: each bookings-api replica sustains
~85 rps with acceptable p95 latency.

Replicas needed: 6,250 / 85 = 73.5  ->  with a 20% margin: 88 replicas.

CPU: 88 x 520m = 45.76 cores.
Memory: 88 x 820Mi = 70.4 GiB.

web-store: serves static assets, scales better. 6,250 rps / 400 rps per replica = 16
  -> with margin: 20 replicas. 20 x 70m = 1.4 cores.

notifications-worker: the queue fills up but can be drained AFTER 22:00.
  It does NOT need scaling during the campaign. It stays at 4 replicas and is
  scaled to 25 from 22:00 onwards, when there is spare capacity.
  A deliberate decision: absolute priority to sales.

Ingress: 8 replicas (double the usual). 1.6 cores.

TOTAL AT THE PEAK: 45.76 + 1.4 + 0.48 + 1.6 = 49.24 cores
                   70.4 + 1.25 + 1.6 + 2 = 75.25 GiB

Step 2: translate into nodes.

Usable capacity per general node (4 cores): 3.05 cores, 13.5 GiB.

By CPU:    49.24 / 3.05 = 16.1  ->  17 nodes
By memory: 75.25 / 13.5 = 5.6   ->  6 nodes

CPU IS THE LIMIT: 17 general nodes.

Alternative: bigger nodes. With nodes of 8 cores / 32 GiB:
  Usable per node: 8 - 0.5 - 0.45 = 7.05 cores
  Nodes needed: 49.24 / 7.05 = 6.98  ->  7 nodes

7 large nodes instead of 17 small ones:
  - Fewer nodes to boot = less total provisioning time
  - Fewer replicated DaemonSets (17 x 450m = 7.65 cores versus 7 x 450m = 3.15)
  - Worse granularity and worse fault tolerance

DECISION: 8-core nodes for that night. The DaemonSet saving (4.5 cores) is
decisive, and since all the capacity is booted BEFORE, the worse granularity
does not matter.

Step 3: the manifests.

# k8s/environments/pro/campaign/hpa-bookings-api-blackfriday.yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: bookings-api
  namespace: rutas-norte-pro
  annotations:
    rutasnorte.example/campaign: "black-friday-2026"
    rutasnorte.example/revert-before: "2026-11-28T23:00:00Z"
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: bookings-api
  # minReplicas RAISED TO 88. This is the key to the whole strategy:
  # the capacity is BOOTED AND WARM before 20:00, we do not wait for the
  # HPA to react. With a 20-second rise, any reactive mechanism
  # arrives late.
  minReplicas: 88
  maxReplicas: 110          # A 25% margin in case the estimate falls short
  metrics:
    - type: ContainerResource
      containerResource:
        name: cpu
        container: api
        target:
          type: Utilization
          averageUtilization: 60
  behavior:
    scaleUp:
      stabilizationWindowSeconds: 0
      selectPolicy: Max
      policies:
        - type: Percent
          value: 100
          periodSeconds: 15
        - type: Pods
          value: 15
          periodSeconds: 15
    scaleDown:
      # DISABLED during the campaign. Not a single replica fewer between
      # 19:00 and 22:30, whatever the metrics do. A 30-second trough must
      # not destroy capacity that took 20 minutes to raise.
      selectPolicy: Disabled
# k8s/environments/pro/campaign/deployment-capacity-filler-blackfriday.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: capacity-filler
  namespace: rutas-norte-pro
spec:
  # Cushion expanded to 10 pods of 1 core. It instantly absorbs
  # replicas 89 to 108 if the estimate falls short.
  replicas: 10
  selector:
    matchLabels:
      app: capacity-filler
  template:
    metadata:
      labels:
        app: capacity-filler
      annotations:
        cluster-autoscaler.kubernetes.io/safe-to-evict: "true"
    spec:
      priorityClassName: capacity-filler
      terminationGracePeriodSeconds: 0
      topologySpreadConstraints:
        - maxSkew: 1
          topologyKey: kubernetes.io/hostname
          whenUnsatisfiable: ScheduleAnyway
          labelSelector:
            matchLabels:
              app: capacity-filler
      containers:
        - name: pause
          image: registry.k8s.io/pause:3.9
          resources:
            requests: {cpu: "1", memory: 1Gi}
            limits: {cpu: "1", memory: 1Gi}

Node groups for that night:

rutas-norte-bf-a   min=3  max=4   8 cores / 32 GiB   zone A
rutas-norte-bf-b   min=3  max=4   8 cores / 32 GiB   zone B
rutas-norte-bf-c   min=2  max=4   8 cores / 32 GiB   zone C

Guaranteed minimum total: 8 nodes (already booted, not dependent on the CA)
Maximum total: 12 nodes

Note the important detail: the groups' min is raised to 3, we do not rely on the CA. By setting the minimum high, the provider boots the machines and the CA cannot remove them. The capacity is guaranteed by construction, not by reaction.

Step 4: the cost.

Node of 8 cores / 32 GiB: ~0.26 EUR/hour on demand.

Expanded-capacity window: 18:00 to 23:00 = 5 hours.
Extra nodes over the baseline (baseline: 3 nodes of 4 cores):
  8 nodes of 8 cores x 5 hours x 0.26 = 10.40 EUR

Estimated additional egress traffic: ~15 EUR
Margin for incidents (nodes up to the maximum of 12): +4 nodes x 5 h x 0.26 = 5.20 EUR

TOTAL ESTIMATE: ~31 EUR

BUDGET: 400 EUR.
Margin: 369 EUR (92%).

The result is surprising and it is the lesson of the exercise: sizing generously for two hours is dirt cheap. The cost of cloud infrastructure is proportional to time, and two hours of over-provisioning cost less than a coffee per node. With a €400 budget, there is no excuse whatsoever for skimping on capacity that night.

With the spare margin you can afford improvements:

  • Going to 12 nodes from the start instead of 8: +€5.20.
  • Doubling the redis-cache replicas as a read replica: +€2.
  • Keeping the capacity until 2:00 to drain the notifications queue without rushing: +€15.

All of that fits comfortably.

Step 5: the checklist for the night.

Time Action Owner Verification
3 weeks before Load test on rutas-norte-pre at 6,250 rps (09-06) Platform Measure real rps per replica
2 weeks before Check the provider quotas: does it allow 12 nodes of 8 cores? Platform Raise a quota ticket if not
1 week before Full dress rehearsal: apply the manifests in pre and measure the complete start-up time Platform Stopwatch
1 week before Review the PDBs: maxUnavailable as a percentage, not absolute Platform kubectl get pdb
3 days before Freeze deployments. No code changes until Monday The whole team Block in CI
17:00 Raise the node groups' min to 3/3/2 Platform kubectl get nodes = 8
17:30 Verify that the 8 nodes are Ready and have the images preloaded Platform kubectl get nodes
18:00 Apply the campaign HPA (minReplicas: 88) Platform kubectl get hpa
18:00 Expand the filler to 10 replicas Platform kubectl get pods -l app=capacity-filler
18:30 Verify 88 pods Running and Ready, not merely created Platform kubectl get pods --field-selector status.phase=Running | wc -l
18:45 Smoke test: 100 real end-to-end test bookings QA All OK
19:00 Check Grafana: p95 latency stable, zero errors, zero Pending Platform Dashboards
19:30 Pre-warm the caches: query the 50 best-selling routes Platform redis-cli DBSIZE
19:45 War room open. The whole team connected Everyone —
19:55 Final check. kubectl get pods --all-namespaces | grep -v Running Platform Empty
20:00 OPENING. Watch: p95 latency, error rate, Pending pods, PostgreSQL connections Everyone Dashboards
20:00-22:00 Continuous watch. The rule: touch nothing unless there is a confirmed incident Everyone —
22:00 Campaign closes — —
22:15 Scale notifications-worker to 25 replicas to drain the queue Platform Queue length
23:00 Restore the normal HPA (minReplicas: 4) Platform kubectl apply of the base manifest
23:30 Lower the node groups' min. The CA will remove the rest Platform kubectl get nodes
The next day Analysis: real rps, latencies, comparison against the estimate Everyone Report

Critical notes on the plan:

  1. The 18:30 verification is the most important one. 88 pods created is not the same as 88 pods Ready. If at 18:30 there are 20 in Pending, there are two hours to fix it. If you find out at 20:01, there is nothing to be done.

  2. The scaleDown: Disabled is non-negotiable. Without it, a one-minute trough between 20:15 and 20:16 could destroy 30 replicas that would take minutes to come back.

  3. notifications-worker is not scaled during the campaign. It is a deliberate decision: the confirmation emails can wait two hours, the sale cannot. All the capacity goes to sales. This is the kind of explicit trade-off that distinguishes a capacity plan from a wish list.

  4. The "touch nothing" rule during the campaign. Almost every serious incident at events like this is caused by somebody trying to fix something under pressure. If the dashboards are green, you look and you do not touch.

  5. Pre-warming the caches at 19:30 saves the first 5,000 users from paying the cost of cold queries against PostgreSQL. It is half an hour of work with an enormous impact on the first few minutes.

Conclusion

Cluster autoscaling is the level that holds up the other two. Without it, the HorizontalPodAutoscaler has a hard, invisible ceiling: it asks for replicas, the Deployment creates them, and they sit in Pending while the dashboards show numbers that do not correspond to reality.

The essentials:

  • The symptom is always the same: pods in Pending with FailedScheduling and Insufficient cpu. Recognising it and alerting on it comes first.
  • The Cluster Autoscaler watches unschedulable pods, simulates which node group would host them, and chooses with an expander. least-waste is the most sensible in general.
  • Scaling up is easy; scaling down is where the problems are. A utilisation threshold, a waiting time, and a list of blockers you have to know: pods without a controller, local storage, restrictive PDBs, system pods, and the safe-to-evict: "false" annotation.
  • Diagnosis means reading three places: the cluster-autoscaler-status ConfigMap, the NotTriggerScaleUp events on the pending pods, and the CA logs with their not suitable for removal.
  • A node takes between 3 and 6 minutes to go from being requested to serving traffic. That does not save you from a spike that rises in forty seconds, however well configured everything is.
  • Over-provisioning with negative-priority pods turns three and a half minutes into three seconds, in exchange for paying for idle capacity. It is the most elegant use of the PriorityClass and preemption of 06-05, and it regenerates itself.
  • Karpenter picks the machine according to the pending pods instead of working with fixed groups, and consolidates continuously. Better on AWS and with heterogeneous workloads; the CA remains the solid, predictable option everywhere else.
  • The three autoscalers fit together as a cascade: the HPA creates pods → if they do not fit, Pending → the CA adds nodes. And the cushion cuts out the long path.
  • The Rutas Norte capacity plan costs practically the same as last year's insufficient fixed sizing, but with the capacity of the peak sizing. That is the entire value of the module.

But we still have an underlying problem that none of the three lessons has solved. Everything we have built is reactive: it waits for CPU to rise before acting. And there are two cases where that simply does not work.

The first is notifications-worker. During the May bank holiday it can have forty thousand emails waiting in the queue and CPU at 20%, because its job is to wait for responses from the mail server, not to compute. A CPU-based HPA would see that 20% against a 70% target and conclude that there are replicas to spare: it would scale down precisely when the most consumers are needed.

The second is the very moment the sale opens. We know it happens at 10:00 sharp. It is information we have months in advance. And yet our entire system waits until 10:00:15 to discover, through CPU, something we already knew in March.

In the next lesson, Event-Driven and Custom-Metric Scaling with KEDA, we solve both. We will see how the HPA can consume business metrics through the aggregation API, what KEDA is and why it does not replace the HPA but feeds it, how to scale notifications-worker by the real length of its queue — scale-to-zero included —, how to scale bookings-api by requests per second using the metrics we instrumented in 07-03, and how to schedule a trigger that pre-warms the whole platform half an hour before the May bank-holiday sale opens.

Kubernetes Course

Module 1: Introduction to Kubernetes

Module 2: Core Kubernetes Components

Module 3: Configuration and Secret Management

Module 4: Networking in Kubernetes

Module 5: Storage in Kubernetes

Module 6: Advanced Kubernetes Concepts

Module 7: Monitoring and Logging

Module 8: Kubernetes Security

Module 9: Scaling and Performance

Module 10: Kubernetes Ecosystem and Tooling

Module 11: Case Studies and Real-World Applications

Module 12: Preparing for Kubernetes Certification

© Copyright 2026. All rights reserved