If you already know how to build an image with Docker and run a container, you are halfway there: you know how to package software. What Kubernetes solves is the other half, the half that shows up when that container is no longer alone on your laptop and becomes one of forty that must be alive, reachable and up to date at three in the morning on the Friday of a bank-holiday weekend. This lesson explains exactly what problem container orchestration solves, where Kubernetes comes from, what it offers compared with the alternatives and —just as important— when it is not the right choice. It is the lesson that gives meaning to all the others: without understanding the problem, the solutions in the next eleven modules look like gratuitous complications.
Contents
- From a single container to a fleet
- The leap from Docker and Docker Compose
- A short history: from Borg to the CNCF
- What Kubernetes actually gives you
- Kubernetes compared with the alternatives
- When Kubernetes is NOT a good idea
- The course scenario: Rutas Norte S.L.
- From a single container to a fleet
A container solves one very specific problem: packaging an application together with its dependencies so that it runs identically on any machine with a container runtime. That is huge, but it is only the packaging.
As soon as you move to production, questions appear that the container alone does not answer:
- Placement: I have 5 servers and 40 containers. Which one goes where? And when I add a sixth server?
- Survival: the process died at 03:14. Who brings it back up? And if the whole machine has shut down?
- Scaling: today I need 3 replicas of the API, and on Easter Friday I will need 20. Who starts them and who shuts them down afterwards?
- Discovery: the API has restarted and now has a different IP. How does the frontend find out?
- Load distribution: with 20 replicas, who decides which one serves each request?
- Deployment without downtime: I want to release version 2.4.0 without a single customer seeing a 502 error.
- Configuration and secrets: the database password cannot live inside the image, nor in the
git log. - Resources: one container going haywire must not take down the other nine on the same machine.
Each of those questions, solved by hand, turns into a script. The sum of those scripts, maintained by one person who then goes on holiday, is the reason container orchestration exists.
An orchestrator is, in a single sentence: a system you tell what state you want, and which then works permanently to make reality match that description. You say "I want 5 replicas of the API on version 2.4.0, reachable at api.rutasnorte.example"; the orchestrator works out where they fit, starts them, watches them, replaces the ones that die and routes traffic to them.
The change of mindset: imperative versus declarative
This is the concept that is hardest at first and pays off most later.
| Approach | How it is expressed | Example | Who holds the state |
|---|---|---|---|
| Imperative | A sequence of orders | "start this container", "stop that one", "copy this file" | You, in your head and with scripts |
| Declarative | A description of the outcome | "there must be 5 replicas of this image with this configuration" | The system, in a continuous loop |
Docker is fundamentally imperative (docker run). Kubernetes is fundamentally declarative: you write a file describing what you want, you hand it to the cluster, and from then on the cluster works for you permanently. If somebody deletes a container, it comes back. If a server goes down, its workloads reappear elsewhere. This mechanism is called the reconciliation loop and we will study it in depth in Kubernetes Architecture.
- The leap from Docker and Docker Compose
One very common confusion is worth clearing up straight away: Kubernetes does not replace Docker, it replaces what you used to do around Docker.
- Docker (or any OCI-compatible tool) is still what you use to build images.
- The image registry still exists.
- What changes is who runs and governs those containers.
Docker Compose is the natural intermediate step. With a docker-compose.yml you describe several services, their networks and volumes, and one command brings the whole stack up. It is declarative... within a single machine. That is exactly its ceiling:
# Simplified docker-compose.yml of the current Rutas Norte environment
# It works perfectly... on ONE machine.
services:
web-store:
image: registry.rutasnorte.example/web-store:1.8.0
ports: ["80:80"]
bookings-api:
image: registry.rutasnorte.example/bookings-api:2.3.1
environment:
DB_HOST: bookings-postgres
bookings-postgres:
image: postgres:16
volumes: ["data:/var/lib/postgresql/data"]
volumes:
data:What this file cannot do:
- Spread those services across 5 machines.
- Bring the stack back up on another server if the current one shuts down.
- Scale
bookings-apito 20 replicas, distributed and load-balanced. - Replace image
2.3.1with2.4.0gradually, rolling back automatically if it fails. - Isolate environments (
dev,pre,pro) with different permissions and quotas.
Kubernetes does exactly those five things, and that is the leap. The price is a learning curve and a layer of new concepts, which is precisely what this course is going to give you.
- A short history: from Borg to the CNCF
Kubernetes was not born as an experiment. It is the third generation of an idea proven over more than a decade.
| Stage | What happened | Why it matters |
|---|---|---|
| ~2003–2014 | Google runs Borg internally, and later Omega, systems that run billions of containers a week | Kubernetes inherits their ideas: pods, labels, controllers, a central scheduler |
| 2013 | Docker makes containers popular for everyone | The standard packaging arrives; the governance is still missing |
| June 2014 | Google releases Kubernetes as an open source project | The name comes from the Greek κυβερνήτης, "helmsman". Hence the shorthand K8s (K + 8 letters + s) |
| July 2015 | Version 1.0 and donation to the CNCF (Cloud Native Computing Foundation), under the Linux Foundation | It stops being "Google's": neutral governance, which enabled adoption by AWS, Microsoft, Red Hat, VMware... |
| 2016–2018 | Docker Swarm, Mesos and others compete; the ecosystem converges on Kubernetes | It becomes the de facto standard |
| 2020–today | dockershim is removed and the runtime is now addressed through CRI (containerd, CRI-O). Roughly 3 minor releases a year |
This explains why people say "Kubernetes no longer uses Docker" (it uses containerd directly) |
The point that matters to you as a professional: because it sits under the CNCF with neutral governance, Kubernetes is the only orchestration layer that runs practically the same way on your laptop, in your own data centre and on AWS, Azure or Google Cloud. That portability is a business argument, not just a technical one.
- What Kubernetes actually gives you
Let us get specific about the capabilities, because "orchestrating" is far too vague.
4.1. Declarative state and reconciliation
You describe the outcome in a YAML file and the cluster maintains it. You do not execute steps: you publish intentions. Everything else on this list follows from that.
4.2. Self-healing
- If a container exits with an error, it is restarted.
- If a container is alive but unresponsive (detected by probes, module 7), it is restarted.
- If a whole node stops responding, its workloads are recreated on other nodes.
- If somebody deletes a replica by hand, it comes back.
4.3. Horizontal scaling, manual and automatic
You can go from 3 to 20 replicas with a single command, or let the cluster do it on its own based on CPU, memory or business metrics (module 9). For Rutas Norte, whose traffic peaks fall on bank-holiday weekends and holidays, this translates directly into money: you do not pay for idle capacity in February just to survive August.
4.4. Service discovery and load balancing
Replicas are born and die with different IPs. Kubernetes gives you a stable name (bookings-api) that resolves through internal DNS and spreads traffic across the healthy replicas. The frontend never knows an IP. You will see this in module 4.
4.5. Progressive, downtime-free deployments
You change the image version and the cluster replaces the replicas little by little, checking that the new ones are healthy before retiring the old ones. If something goes wrong, rollback to the previous version. Module 2.
4.6. Configuration and secret management
Configuration is kept separate from the image (ConfigMaps) and credentials are handled apart, with access control (Secrets). Since bookings-postgres stores customers' personal data, this is not optional for Rutas Norte: it is regulatory compliance. Module 3.
4.7. Portability
The same manifests work on minikube, on a self-managed cluster built with kubeadm, or on EKS/AKS/GKE. What changes are the infrastructure details (storage type, load balancer), not the description of the application.
4.8. Extensibility
You can add your own object types (CRDs) and your own reconciliation logic (operators, module 6). Kubernetes is not just an orchestrator: it is a platform for building platforms.
- Kubernetes compared with the alternatives
No tool is the best in the abstract. This table will help you justify a decision to a team or a client.
| Criterion | Kubernetes | Docker Compose | Docker Swarm | HashiCorp Nomad | PaaS (Heroku, App Engine, Cloud Run) |
|---|---|---|---|---|---|
| Scope | Multi-node cluster | One machine | Multi-node cluster | Multi-node cluster | Managed service |
| Learning curve | Steep | Very gentle | Gentle | Moderate | Very gentle |
| Self-healing | Yes, complete | No (only restart) |
Yes, basic | Yes | Yes (opaque) |
| Autoscaling | Yes (pods, nodes, events) | No | Not natively | With integrations | Yes, managed |
| Non-containerised workloads | No (containers only) | No | No | Yes (binaries, Java, QEMU) | No |
| Ecosystem and community | Enormous (CNCF) | Large but bounded | In decline | Modest | Closed, vendor-specific |
| Portability across clouds | Very high | High (but no cluster) | Medium | High | None: vendor lock-in |
| Operational cost | High | Almost nil | Low | Medium | Low (paid on the invoice) |
| Fine-grained control (network, storage, security) | Total | Minimal | Limited | Good | Minimal |
| Typical fit | Serious production, several teams and services | Local development and demos | Small stacks already using it | Mixed container + non-container environments | Small teams that prioritise speed |
Quick readings of the table:
- Docker Compose does not compete with Kubernetes: it coexists with it. You will keep using it locally.
- Swarm is simpler, but its community and ecosystem have shrunk drastically; starting a new project on Swarm today means taking on debt.
- Nomad is the serious alternative if you have workloads that are not containers.
- A PaaS may well be the right answer, and admitting it is not a failure. You pay more per unit of compute, but you save on people.
- When Kubernetes is NOT a good idea
A professional is recognised by knowing when not to apply a tool. Clear signs that Kubernetes is overkill:
- A small team with no platform skills. Kubernetes needs somebody to look after it: upgrades every few months, certificates, monitoring, permissions. If nobody has that time assigned, the cluster degrades.
- A single monolithic application with steady traffic. If two machines and a load balancer cover your load all year round, a cluster only adds pieces that can fail.
- Trivial or short-lived workloads. A daily cron job that takes 40 seconds does not justify a control plane.
- A need for immediate results. The migration is a project of weeks or months. If the business needs to deliver in two weeks, a PaaS delivers sooner.
- Applications tightly bound to one specific server (per-machine licences, special hardware, state on local disk with no replication). It can be done, but you are fighting the design.
- Real cost badly estimated. A managed cluster costs money for the control plane, nodes with reserved capacity, traffic and storage. For three containers, it works out expensive.
Rule of thumb: Kubernetes starts to pay off when you have several services, several environments and more than one person deploying. Before that, it is usually a solution looking for a problem.
- The course scenario: Rutas Norte S.L.
The whole course revolves around a fictional company: Rutas Norte S.L., which sells coach tickets online through the Rutas Norte platform. Here we only introduce it; the full detail is in the lesson The Course Project: the Rutas Norte Platform.
Its starting point is the one many real companies find themselves in:
- The platform runs on Docker Compose across two rented machines.
- Deployments are manual: somebody logs in over SSH and runs
docker compose pullandup -d. There is a service outage of one or two minutes, so releases happen at night. - On bank-holiday weekends and during the holidays traffic multiplies; the website slows down and some bookings are lost. The only answer so far has been to rent a bigger machine and leave it idle the rest of the year.
- When a container dies in the small hours, nobody brings it back until the next morning.
- The database password lives in a
.envfile that has been emailed around more times than anybody wants to admit; and that database stores customers' personal data.
The technical leadership has decided to migrate to Kubernetes for four reasons, which are exactly the capabilities in section 4:
| Current problem | Capability that solves it | Module where it is covered |
|---|---|---|
| Night-time outages with no response | Self-healing | 2 and 7 |
| Peaks on bank holidays and in the holiday season | Horizontal autoscaling | 9 |
| Deployments with service downtime | Progressive updates and rollback | 2 and 11 |
| Credentials and personal data exposed | Secrets, RBAC and network policies | 3, 4 and 8 |
From the next lesson onwards we will build that platform piece by piece, and every new concept will be justified by a concrete Rutas Norte need.
Common Mistakes and Tips
- "Kubernetes replaces Docker". It does not. You still build images the same way. What Kubernetes replaces is
docker run, the start-up scripts and the hand-rolled load balancer. Since version 1.24 the cluster talks to containerd or CRI-O through the CRI interface, not to the Docker daemon: that is whydockershimdisappeared. - Translating
docker-compose.ymlmechanically. Conversion tools exist, but the result is usually a poor deployment: no probes, no resource limits, no externalised configuration. It pays to redesign, not to translate. - Starting with the production cluster. The right order is: local environment (module 1), understand objects (modules 2 and 3), and only then think about production.
- Believing Kubernetes makes your application resilient. The cluster restarts containers; it does not fix an application that loses data on restart or cannot tolerate running several instances. Kubernetes rewards well-designed applications and punishes the ones that are not.
- Ignoring the cost of people. A cluster's budget is not just infrastructure: it is training, on-call rotas and maintenance. Put it in writing before proposing the migration.
- Vocabulary tip: get into the habit, starting today, of saying "desired state" and "reconciliation". When something does not work in the cluster, the right question is almost always: what state have I declared, and why can the cluster not reach it?
Exercises
Exercise 1: Diagnosing the scenario
Without writing a single manifest, draw up a table with the five operational problems Rutas Norte suffers today and, for each one, state: (a) which Kubernetes capability would solve it, and (b) what would happen if the company decided not to migrate and simply rented bigger servers.
Exercise 2: An architecture decision
For each of these three cases, decide whether you would recommend Kubernetes, a PaaS or Docker Compose, and justify the decision in three lines:
- A two-person startup with one web application and one database, that needs to reach the market within a month.
- An insurance company with 60 microservices, three environments, audit requirements and teams in two countries.
- A static corporate blog with 200 visits a day.
Exercise 3: A pitch for the board
Write a paragraph of at most 150 words addressed to the non-technical leadership of Rutas Norte explaining why the company should migrate to Kubernetes, without using the words "pod", "container" or "cluster". It must mention cost, availability and customer data.
Solutions
Solution 1
| Current problem | Kubernetes capability | If they do not migrate and only grow the server |
|---|---|---|
| Container down overnight with no recovery | Self-healing: automatic restart and recreation | A bigger server restarts nothing: the outage lasts just as long |
| Peaks on bank holidays and in the holiday season | Horizontal autoscaling | You pay for peak capacity all year round; there is still a fixed ceiling |
| Deployment with service downtime | Progressive update with health checks and rollback | The downtime stays; it just moves to another machine |
Credentials in a shared .env |
Secrets with access control and RBAC | No improvement whatsoever; it is a process problem, not a sizing one |
| Single point of failure (two machines) | Workloads rescheduled onto other nodes | One big server is an even bigger single point of failure |
Solution 2
- PaaS. Two people cannot maintain a cluster and reach the market within a month. The priority is to ship; vendor lock-in is an acceptable problem to solve later if the product works.
- Kubernetes. 60 microservices, three environments and auditing are exactly the point where the operational cost pays for itself: RBAC, a namespace per environment, network policies and consistent deployments across teams and countries.
- None of the three, or Docker Compose at most. 200 visits a day on a static site are served by static hosting or a CDN. Any orchestrator is pure cost with no benefit.
Solution 3 (example of a valid answer)
Today, ticket sales depend on two machines we manage by hand. When something fails at night, it is not recovered until the next morning, and on bank holidays and during the holiday season —exactly when we sell the most— the website slows down and we lose bookings. On top of that, we pay all year for capacity we only need for a few weeks. The new platform we are proposing watches the service continuously and restores it by itself if it fails, expands capacity automatically when demand rises and reduces it when demand falls, and lets us publish improvements without taking the website down. It also separates the keys that unlock our customers' data from the code and records who may consult them, which strengthens our position in a data protection audit. The initial cost is training and setup; the saving is in capacity not wasted and in sales we lose today.
Conclusion
Kubernetes is a container orchestrator: a system to which you declare the state you want and which works continuously to maintain it, providing self-healing, scaling, service discovery, downtime-free deployments and portability across clouds. It grew out of Google's experience with Borg, it is today a neutral CNCF project and it has become the de facto standard, but it carries a real operational cost that makes it a poor fit for small teams and trivial workloads. Rutas Norte falls squarely into the profile that does justify it: several services, several environments, pronounced demand peaks and personal data to protect.
You now know what Kubernetes does. The next question is how it does it: which pieces make up a cluster, what the reconciliation loop that underpins the whole declarative model looks like, and what happens exactly, step by step, from the moment you submit a manifest until a container is running on a node. That is what we will see in Kubernetes Architecture.
Kubernetes Course
Module 1: Introduction to Kubernetes
- What Is Kubernetes?
- Kubernetes Architecture
- Key Concepts and Terminology
- Setting Up a Kubernetes Cluster
- The Kubernetes CLI: kubectl
- Objects, YAML Manifests and the Declarative Model
- The Course Project: the Rutas Norte Platform
Module 2: Core Kubernetes Components
- Pods
- ReplicaSets
- Deployments
- Updates, Rollbacks and Deployment Strategies
- Services
- Namespaces
- Labels, Selectors and Annotations
Module 3: Configuration and Secret Management
- ConfigMaps
- Secrets
- Environment Variables
- Resource Quotas and Limits
- LimitRanges and Quality of Service (QoS) Classes
- ServiceAccounts and API Access from Pods
Module 4: Networking in Kubernetes
- Cluster Networking
- Service Types
- Internal DNS and Service Discovery
- Ingress Controllers
- TLS and Certificate Management with cert-manager
- Network Policies
Module 5: Storage in Kubernetes
- Volumes
- Persistent Volumes
- Persistent Volume Claims
- Storage Classes
- Dynamic Provisioning, Expansion and Snapshots
- Backup and Restore of Persistent Data
Module 6: Advanced Kubernetes Concepts
- StatefulSets
- DaemonSets
- Jobs and CronJobs
- Init Containers, Sidecars and Multi-Container Patterns
- Scheduling: Affinity, Taints and Tolerations
- Custom Resource Definitions (CRDs)
- Operators and the Controller Pattern
Module 7: Monitoring and Logging
- Health Checks and Probes
- Metrics Server and kubectl top
- Monitoring with Prometheus
- Visualization and Alerting with Grafana and Alertmanager
- Centralized Logging with Elasticsearch, Fluentd and Kibana (EFK)
- Application Debugging and Cluster Events
Module 8: Kubernetes Security
- Role-Based Access Control (RBAC)
- Security Contexts and Container Hardening
- Pod Security Policies and Pod Security Standards
- Network Security
- Image Security
- Auditing, Scanning and Vulnerability Management
Module 9: Scaling and Performance
- Horizontal Pod Autoscaling
- Vertical Pod Autoscaling
- Cluster Autoscaling
- Event-Driven and Custom-Metric Scaling with KEDA
- High Availability: PodDisruptionBudgets and Topology
- Performance Tuning
Module 10: Kubernetes Ecosystem and Tooling
- Minikube and Local Environments with kind
- Kubeadm
- Helm
- Kustomize
- GitOps with Argo CD and Flux
- Managed Kubernetes: EKS, AKS and GKE
Module 11: Case Studies and Real-World Applications
- Deploying a Web Application
- Running Stateful Applications
- CI/CD with Kubernetes
- Deployment Strategies: Blue-Green and Canary
- Multi-Cluster Management
- Production Operations: Incidents, Runbooks and Costs
