We closed module 3 by pointing out that the Rutas Norte network is completely flat: any pod in rutas-norte-pro can open a connection to bookings-postgres:5432, and no client on the internet can reach www.rutasnorte.example yet. Before fitting doors (the network policies of 04-06) and windows (the Ingress of 04-04), you need to understand how the house is built. This is the lesson that explains what really sits behind that 10.244.x.x you see in kubectl get pods -o wide, who creates that network interface, why a Service IP appears on no network card of any machine, and how a packet leaving bookings-api ends up arriving at bookings-postgres. It is the most "plumbing" lesson of the course, and also the one that stops everything else from looking like magic.
Contents
- The Kubernetes network model and its four rules
- Why that model simplifies everything else
- The four communication planes
- CNI: the interface that plugs the pod into the network
- The plugins: Flannel, Calico, Cilium and Weave Net
- Overlay (VXLAN), native routing (BGP) and eBPF
- The cluster's address ranges
- How a ClusterIP is really implemented: kube-proxy
- The full journey of a packet in Rutas Norte
- Hands-on verification in minikube
- The Kubernetes network model and its four rules
Kubernetes does not implement the network. What it does is impose a contract: it defines a model that any network implementation must satisfy, and lets others satisfy it however they like. The contract has four rules.
| # | Rule | What it means in practice |
|---|---|---|
| 1 | Every pod has its own IP address | The node's IP is not shared and there is no port mapping. bookings-api listens on 3000 inside its own IP, even with five replicas on the same node |
| 2 | Every pod can talk to every pod without NAT | It makes no difference whether they are on the same node or on different nodes: the connection is direct, with no address translation |
| 3 | Node agents can reach the pods on that node | The kubelet, and therefore the health probes of 07-01, can connect to pods without tricks |
| 4 | The IP a pod sees for itself is the one everybody else sees | If hostname -i inside the container says 10.244.1.37, the other pods see it as exactly 10.244.1.37 |
The fourth rule looks redundant, but it is the one that explicitly outlaws the classic Docker model. In a docker run -p 8080:80, the container believes it is 172.17.0.4:80 while the rest of the world sees it as 192.168.1.10:8080. That mismatch breaks any protocol that advertises its own address: clustered databases, discovery systems, service registries. Kubernetes eliminates it at the root.
Notice what the contract does not say: nothing about how packets are routed, whether there is encapsulation, which IP range is used or whether there is a firewall. All of that is left to the implementation.
- Why that model simplifies everything else
The model boils down to a single phrase: "one IP per pod, flat network". Its consequences are enormous.
- Ports stop being a scarce resource. In a port-mapping world, deploying three replicas of
bookings-apion the same node forces you to assign 3000, 3001 and 3002, and to have somebody keep track. With one IP per pod, all three listen on 3000. The manifest does not change depending on where the pod lands. - Applications do not need to know they are in Kubernetes. An application that works on a virtual machine works in a pod: it opens its port and connects to
host:port. This is what made it possible to migrate so much software without rewriting it. - Discovery becomes trivial. Since there is no NAT, the IP that the
EndpointSliceof 02-05 points to is directly connectable. Without that guarantee, every Service would have to translate addresses and ports per node. - Pod migration is transparent. A pod dies and another is born with a different IP, but the mechanics of connecting are identical.
The price is that by default there is no isolation whatsoever. Rule 2 says "every pod can talk to every pod", and that is exactly what it means: the most insignificant pod in rutas-norte-pro can try to connect to bookings-postgres. The NetworkPolicies of 04-06 exist precisely to trim back that universal permission.
- The four communication planes
When somebody says "the Kubernetes network" they are usually mixing up four different problems, with different solutions.
flowchart TB
subgraph N1["Node 1"]
subgraph P1["Pod bookings-api 10.244.1.5"]
C1["api container"]
C2["logs sidecar"]
end
P2["Pod redis-cache<br/>10.244.1.9"]
end
subgraph N2["Node 2"]
P3["Pod bookings-postgres<br/>10.244.2.4"]
end
EXT["Internet client"]
SVC["ClusterIP Service<br/>10.96.31.72"]
C1 -.->|"1. localhost"| C2
P1 -->|"2. pod to pod, no NAT"| P3
P1 -->|"3. pod to Service"| SVC
SVC --> P2
EXT -->|"4. outside to Service"| SVC
| Plane | How it is solved | Where it is covered |
|---|---|---|
| 1. Container ↔ container in the same pod | They share the network namespace: they see each other over localhost and compete for the same ports |
Already seen in 02-01; patterns in 06-04 |
| 2. Pod ↔ pod | Implemented by the CNI plugin. It is the central topic of this lesson | Here, sections 4 to 7 |
| 3. Pod ↔ Service | Implemented by kube-proxy (or by the CNI, if it replaces kube-proxy) rewriting destinations | Here, section 8 |
| 4. Outside ↔ Service | Service types NodePort/LoadBalancer and Ingress |
04-02 and 04-04 |
A very common diagnostic mistake is attacking the wrong plane. If bookings-api cannot reach bookings-postgres by pod IP, the problem belongs to the CNI. If it reaches it by pod IP but not by the Service name, the problem belongs to kube-proxy or to the DNS of 04-03. Separating the planes saves hours.
- CNI: the interface that plugs the pod into the network
CNI (Container Network Interface) is a CNCF specification, very short and deliberately boring: it defines how a container runtime asks an external program to connect or disconnect a container from a network. It is not a product: it is a contract between two parties.
The call chain when a pod is created is this:
sequenceDiagram
participant K as kubelet
participant CR as containerd
participant CNI as CNI plugin
participant IPAM as IPAM
K->>CR: create the pod sandbox
CR->>CR: create empty network namespace
CR->>CNI: ADD (namespace, pod ID)
CNI->>IPAM: give me an IP from the node podCIDR
IPAM-->>CNI: 10.244.1.37/24
CNI->>CNI: create veth, move one end into the ns
CNI->>CNI: assign IP, default route
CNI-->>CR: OK, IP 10.244.1.37
CR-->>K: sandbox ready
K->>CR: start the pod containers
The points worth fixing in your mind:
- The plugin is invoked by the kubelet through the runtime (containerd, in our cluster). The apiserver never touches the network.
- The plugin is installed as a binary in
/opt/cni/binand configured with a JSON file in/etc/cni/net.d. That is why plugins are almost always deployed as a DaemonSet (06-02): one pod per node that copies the binary and the configuration file at startup and then keeps its agent running. - The operations are
ADD,DEL,CHECKandVERSION. When the pod is deleted,DELis invoked and the IP returns to the pool. - IPAM (IP Address Management) is a subcomponent: it decides which specific IP is assigned within the range allotted to the node.
What the plugin does in a typical ADD, translated into Linux commands you already know:
- It creates a veth pair (two virtual interfaces joined like a cable).
- It leaves one end in the host's network namespace (with a name such as
veth3a7f2c1@if3) and moves the other inside the pod's namespace, where it is calledeth0. - It asks IPAM for an IP and assigns it to
eth0. - It adds the pod's default route, pointing at the bridge or at the node's gateway.
- On the host, it attaches its end to the bridge (
cni0,docker0,cbr0) or creates the corresponding route. - It returns the assigned IP to the runtime, which ends up in the pod's
status.podIP.
If the CNI fails or is not installed, you will see the classic symptom:
NAME READY STATUS RESTARTS AGE
bookings-api-7d9f5c8b4-x2kp9 0/1 ContainerCreating 0 3m
Events:
Warning FailedCreatePodSandBox kubelet Failed to create pod sandbox:
plugin type="bridge" failed (add): failed to set bridge addr: could not add IPPods stuck forever in ContainerCreating with sandbox errors are nearly always a node networking problem, not a problem with your manifest.
- The plugins: Flannel, Calico, Cilium and Weave Net
Choosing a plugin is one of the longest-lasting decisions in a cluster: changing it with production traffic running is a delicate operation. These are the four most widely used historically.
| Flannel | Calico | Cilium | Weave Net | |
|---|---|---|---|---|
| Data model | VXLAN overlay (or host-gw) |
Native L3 routing with BGP; VXLAN/IPIP option | eBPF in the kernel; optional VXLAN/geneve or native | Its own overlay (VXLAN "fast datapath") |
| NetworkPolicy | Does not implement it | Yes, plus richer policies of its own | Yes, including L7 policies (HTTP, Kafka, DNS) | Yes (basic support) |
| Performance | Good; encapsulation penalty | Very good in native mode (no encapsulation) | The best: it can replace kube-proxy | Acceptable; the slowest of the four |
| Complexity | Minimal | Medium (BGP requires understanding the physical network) | High; requires a modern kernel | Low |
| Encryption in transit | No | WireGuard | WireGuard or IPsec | Yes, built in |
| Observability | None | Prometheus metrics | Hubble: flows, service maps, L7 | Basic |
| When to choose it | Learning, lab, cluster with no isolation requirements | Production with NetworkPolicy and a network under your control | Demanding production, observability, L7 policies, mesh without sidecars | Simple installations where it is already in use |
The practical consequence for Rutas Norte, and it deserves underlining: if the cluster uses Flannel, the NetworkPolicies you write in 04-06 will be applied without error and will do absolutely nothing. The object is created, kubectl get networkpolicy lists it, and the traffic keeps flowing. It is the most dangerous silent failure in this part of the course, which is why we will verify it explicitly.
The managed clusters of 10-06 come with their own plugin (Amazon VPC CNI, Azure CNI, the GKE one), which usually assigns pods IPs from the real cloud network. That makes a pod directly routable from other machines in the VPC, at the cost of consuming addresses from the corporate space.
- Overlay (VXLAN), native routing (BGP) and eBPF
The problem they all solve is the same: pod 10.244.1.5 is on node A and wants to talk to 10.244.2.4, which is on node B. The physical network between nodes knows nothing about 10.244.0.0/16: if you hand it that packet as it is, it drops it.
Overlay network (VXLAN)
A virtual network is built on top of the physical one. The source node puts the entire original packet inside a UDP packet addressed to the real IP of the destination node (port 8472 in VXLAN), which decapsulates it and delivers it to the pod.
[ IP nodeA -> IP nodeB | UDP 8472 | VXLAN | IP 10.244.1.5 -> 10.244.2.4 | TCP 5432 | data ]
\_______________ outer envelope _______________/ \____________ original packet ____/- Decisive advantage: it works over any network, without touching routers or asking anything of the networking team. That is why it is the default mode in so many installations.
- Cost: the header adds about 50 bytes, which reduces the effective MTU (typically from 1500 to 1450) and causes fragmentation if something ignores it; there is extra work encapsulating and decapsulating every packet; and the traffic is opaque to the firewalls and analysers of the physical network, which only see UDP between nodes.
Native routing (BGP)
Nothing is encapsulated. Each node advertises over BGP to the routers of the physical network: "the range 10.244.2.0/24 is reachable through me". The physical network learns the routes and delivers pod packets directly.
- Advantage: no overhead, full MTU, pod traffic is visible and filterable with the existing network tools.
- Requirement: the network must allow it. In your own data centre with routers that speak BGP, it is ideal. In many clouds or in closed corporate networks it is not possible and you fall back to encapsulation.
eBPF
eBPF makes it possible to load verified programs inside the Linux kernel and hook them to specific points of the network stack. Cilium uses it to make routing, load-balancing and policy decisions before the packet travels through the whole iptables machinery.
- It can replace kube-proxy entirely, removing the
iptableschains we will see in section 8. - It scales far better: where
iptablesdegrades with thousands of Services, eBPF uses constant-cost hash tables. - It enables HTTP-aware policies (by path and method) and workload identity.
- In exchange, it demands a reasonably recent kernel and raises the bar of knowledge needed to debug it.
| Criterion | VXLAN | Native BGP | eBPF |
|---|---|---|---|
| Per-packet overhead | High (~50 B + encapsulation) | None | None or negative |
| Physical network requirements | None | Must route/speak BGP | None |
| Visibility from the physical network | Low | High | Medium |
| Scaling with many Services | Depends on kube-proxy | Depends on kube-proxy | Excellent |
- The cluster's address ranges
Three address spaces coexist in a cluster, and they must never be confused.
| Range | What it addresses | Example in minikube | Routable outside? |
|---|---|---|---|
| Node network | The physical or virtual machines | 192.168.49.0/24 |
Yes, it is the real network |
podCIDR / cluster CIDR |
The pod IPs | 10.244.0.0/16, sliced into a /24 per node |
Only inside the cluster |
serviceCIDR |
The virtual IPs of the Services | 10.96.0.0/12 |
It exists on no interface |
The global cluster CIDR is shared out: the kube-controller-manager, with --allocate-node-cidrs, assigns each node a slice (a /24 by default, that is 254 usable pods per node) and writes it into spec.podCIDR of the Node object. The CNI plugin's IPAM hands out IPs within that slice. This guarantees that two nodes never assign the same IP.
Where they are configured:
# When creating the cluster with kubeadm (lesson 10-02)
kubeadm init --pod-network-cidr=10.244.0.0/16 --service-cidr=10.96.0.0/12
# In minikube, when creating the profile
minikube start -p rutas-norte --extra-config=kubeadm.pod-network-cidr=10.244.0.0/16Three warnings that save incidents:
- They cannot be changed on the fly. Changing the
serviceCIDRof a live cluster is not a supported operation; you plan it before creating the cluster. - They must not overlap with the corporate network. If your company uses
10.96.0.0/16for servers, the Rutas Norte pods will not be able to talk to them: the node will believe those IPs are cluster Services. - The
/24per node is a real limit. A large node with 300 pods runs out of addresses even if it has CPU and memory to spare.
- How a ClusterIP is really implemented: kube-proxy
Here is the revelation of the lesson. When in 02-05 you created the bookings-postgres Service with clusterIP: 10.96.140.22, that address was not assigned to any interface on any machine. It is a fiction. It does not answer ping. There is no process listening on it.
What does exist are destination-rewriting rules installed on every node by kube-proxy, a DaemonSet in the kube-system namespace that watches the API and, every time a Service or an EndpointSlice changes, updates those rules.
flowchart LR
API["kube-apiserver<br/>Services + EndpointSlices"] -->|watch| KP["kube-proxy<br/>(one pod per node)"]
KP -->|programmes rules| DP["iptables / IPVS<br/>in the node kernel"]
POD["Source pod"] -->|"connects to 10.96.140.22:5432"| DP
DP -->|"DNAT to 10.244.2.4:5432"| DEST["Pod bookings-postgres"]
iptables mode (the most common)
For each Service, kube-proxy creates a KUBE-SERVICES chain that matches on destination IP and port, jumps to a KUBE-SVC-XXXX chain belonging to the Service, and from there spreads across the KUBE-SEP-YYYY chains, one per endpoint, each with its own DNAT rule.
# Entry chain: captures traffic towards the Service IP
-A KUBE-SERVICES -d 10.96.140.22/32 -p tcp --dport 5432
-j KUBE-SVC-QW3RTY5432ABCD
# Spread across 2 endpoints with statistical probability
-A KUBE-SVC-QW3RTY5432ABCD -m statistic --mode random --probability 0.50000
-j KUBE-SEP-AAA111
-A KUBE-SVC-QW3RTY5432ABCD -j KUBE-SEP-BBB222
# Each endpoint: destination translation to the real pod IP
-A KUBE-SEP-AAA111 -p tcp -j DNAT --to-destination 10.244.2.4:5432
-A KUBE-SEP-BBB222 -p tcp -j DNAT --to-destination 10.244.3.7:5432Three important facts come out of this:
- Balancing is random per connection, not per request. A long TCP connection (a PostgreSQL session, a WebSocket) stays stuck to one pod until it closes. That is why an API that multiplexes requests over a few connections can spread load badly.
- The probability is cascaded: with 3 endpoints, the rules carry 1/3, then 1/2, then the remainder. That comes out uniform.
- The cost is linear:
iptablesevaluates rules in order. With tens of thousands of Services, every update forces the reprogramming of huge tables and synchronisation latencies appear. That is the reason IPVS and eBPF exist.
IPVS mode
IPVS is the Linux kernel's L4 load balancer, with hash tables instead of lists of rules.
iptables |
IPVS |
|
|---|---|---|
| Structure | List of rules evaluated in order | Hash table, constant cost |
| Scaling | Degrades with thousands of Services | Stable with tens of thousands |
| Algorithms | Random only | rr, lc, dh, sh, sed, nq |
| Debugging | iptables-save |
ipvsadm -Ln |
| Requirements | None | ip_vs* modules loaded |
In IPVS mode a kube-ipvs0 interface does appear on the node with the Service IPs assigned to it, but it is a dummy interface: it exists so the kernel accepts the packets, not to answer them.
The mental rule to take away: a Service is a rule, not a machine. If a ClusterIP does not answer, do not look for a crashed process: look for empty endpoints (a selector that does not match, as you saw in 02-05) or a kube-proxy that is not synchronising.
- The full journey of a packet in Rutas Norte
Let us follow a real connection: pod bookings-api-7d9f5c8b4-x2kp9 (10.244.1.5, node 1) opens a connection to bookings-postgres:5432 (ClusterIP 10.96.140.22), whose only endpoint is 10.244.2.4 on node 2. The cluster uses VXLAN.
sequenceDiagram
participant APP as bookings-api (10.244.1.5)
participant DNS as CoreDNS
participant NS1 as Node 1 kernel
participant NET as Physical network
participant NS2 as Node 2 kernel
participant PG as bookings-postgres (10.244.2.4)
APP->>DNS: bookings-postgres.rutas-norte-pro.svc.cluster.local?
DNS-->>APP: 10.96.140.22
APP->>NS1: SYN to 10.96.140.22:5432 (leaves via eth0/veth)
NS1->>NS1: KUBE-SERVICES -> KUBE-SVC -> KUBE-SEP<br/>DNAT destination = 10.244.2.4:5432
NS1->>NS1: route: 10.244.2.0/24 is via VXLAN on node2
NS1->>NET: encapsulates UDP 8472 node1 -> node2
NET->>NS2: delivers the outer packet
NS2->>NS2: decapsulates: IP 10.244.1.5 -> 10.244.2.4
NS2->>PG: SYN through the pod veth
PG-->>NS2: SYN-ACK to 10.244.1.5
NS2->>NET: encapsulated return
NET->>NS1: delivery
NS1->>NS1: conntrack undoes the DNAT:<br/>source becomes 10.96.140.22:5432
NS1-->>APP: SYN-ACK apparently from the ClusterIP
The details that matter:
- Name resolution happens first and is independent of everything else (04-03).
- The DNAT happens on the source node, before routing. Node 2 never sees the Service IP.
bookings-postgressees10.244.1.5as the source, the real IP of the client pod, not the node's. This is rule 2 (no source NAT) and it is what allows the NetworkPolicies of 04-06 to identify the sender by its labels. Careful: this changes when traffic comes in from outside withexternalTrafficPolicy: Cluster(04-02).conntrackis indispensable: the kernel remembers the translation so it can undo it in the reply. A fullconntracktable causes connections that are lost at random under load, a classic symptom during the Rutas Norte bank-holiday peaks.
- Hands-on verification in minikube
Everything above can be touched with your own hands on the rutas-norte profile.
The cluster ranges
# podCIDR assigned to each node
kubectl get nodes -o jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.spec.podCIDR}{"\n"}{end}'# The serviceCIDR is not exposed as a field; it is deduced by provoking an error
kubectl create svc clusterip cidr-test --tcp=80:80 --dry-run=server \
-o yaml --clusterip=1.2.3.4The Service "cidr-test" is invalid: spec.clusterIPs[0]:
Invalid value: []string{"1.2.3.4"}: failed to allocate IP 1.2.3.4:
the provided IP (1.2.3.4) is not in the valid range. The range of valid IPs is 10.96.0.0/12It is a very handy trick: the error message reveals the exact range to you.
The real IPs of the platform
NAME READY STATUS IP NODE
bookings-api-7d9f5c8b4-x2kp9 1/1 Running 10.244.0.31 rutas-norte
bookings-postgres-5c9d7f-4nm2q 1/1 Running 10.244.0.14 rutas-norte
NAME TYPE CLUSTER-IP PORT(S)
bookings-postgres ClusterIP 10.96.140.22 5432/TCPThe pod IPs are in 10.244.0.x (the podCIDR of the only node) and the Service one is in 10.96.x.x. Different ranges, different natures.
The rules that do not exist and the ones that do
# The Service IP is on no interface of the node
minikube -p rutas-norte ssh -- ip -4 addr show | grep -E "inet " inet 127.0.0.1/8 scope host lo
inet 192.168.49.2/24 brd 192.168.49.255 scope global eth0
inet 10.244.0.1/24 brd 10.244.0.255 scope global bridgeNot a trace of 10.96.140.22. Now the rules:
-A KUBE-SERVICES -d 10.96.140.22/32 -p tcp -m comment
--comment "rutas-norte-pro/bookings-postgres cluster IP" -m tcp --dport 5432
-j KUBE-SVC-6GHJ2KLM4NOPQRSTThe comment includes <namespace>/<service>: it is the quickest way to locate the rules of a specific Service.
An ephemeral pod with networking tools
The nicolaka/netshoot image ships dig, curl, nc, tcpdump, traceroute and ipvsadm. It is the Swiss army knife for debugging networks in Kubernetes.
kubectl run netshoot --rm -it --restart=Never \
-n rutas-norte-pro --image=nicolaka/netshoot -- bashInside the pod:
# 1. My own IP: it must be in the podCIDR of the node
ip -4 addr show eth0 | grep inet
# inet 10.244.0.42/24 scope global eth0
# 2. DIRECT pod-to-pod connectivity (plane 2, skipping the Service)
nc -zv 10.244.0.14 5432
# Connection to 10.244.0.14 5432 port [tcp/postgresql] succeeded!
# 3. Pod-to-Service connectivity (plane 3)
nc -zv bookings-postgres 5432
# Connection to bookings-postgres 5432 port [tcp/postgresql] succeeded!
# 4. What proves the warning from module 3: any old pod,
# with no relationship whatsoever to the platform, reaches the database
# holding the customers' personal data. That is what 04-06 fixes.If step 2 works and step 3 does not, the CNI is fine and the problem is the Service or DNS. If step 2 does not work either, it is the CNI or a network policy. That fork covers 80% of network diagnosis in Kubernetes.
Common Mistakes and Tips
- Confusing the three ranges. Seeing
10.96.x.xand hunting for the node that pod is on is a classic. Rule of thumb:10.244.x.xis a pod (it exists),10.96.x.xis a Service (it is a rule),192.168.x.xis a node. - Pinging a ClusterIP. It does not answer and that is normal: kube-proxy's rules only match TCP/UDP on the declared port, not ICMP. A failed ping to a Service proves nothing. Use
nc -zvorcurl. - Overlapping the
podCIDRor theserviceCIDRwith the corporate network. It causes baffling failures when talking to external services, such as the payment gateway. Check it before creating the cluster, not afterwards. - Ignoring the MTU with VXLAN. Connections that open fine but hang when transferring large blocks (a PostgreSQL restore, a 2 MB JSON response) usually mean a badly tuned MTU. Diagnosis:
ping -M do -s 1400 <ip>. - Choosing Flannel and writing NetworkPolicies. They apply without error and filter nothing. Before trusting a policy, verify that the CNI implements it with a real test.
- Blaming the CNI for everything. Before that, check the planes in order:
localhostinside the pod, pod IP, ClusterIP, DNS name. The first one that fails points at the guilty layer. - Tip: keep an alias for the debugging pod.
alias kshoot='kubectl run netshoot-$RANDOM --rm -it --restart=Never --image=nicolaka/netshoot -- bash'. With--rmyou leave no rubbish behind in the cluster. - Tip: in production, watch the
conntrackmetrics (nf_conntrack_countagainstnf_conntrack_max). During the Rutas Norte bank-holiday peaks, a full table shows up as intermittent connection errors with no pod down at all.
Exercises
Exercise 1: Map out the cluster network
On the rutas-norte profile, document: the node's podCIDR, the serviceCIDR, the node's IP, and the IPs of the pods and Services in rutas-norte-pro. Classify each address into its range and explain which ones exist physically and which do not.
Exercise 2: Prove that the ClusterIP is a fiction
Pick the redis-cache Service. Prove with three checks that its IP exists on no interface, that iptables rules do exist for it, and that you can nevertheless connect to it. Explain who does the translation and on which node it happens.
Exercise 3: Separate the planes during a breakdown
A colleague says that bookings-api "cannot see the database". Design a sequence of checks, from the lowest layer to the highest, that lets you decide whether the problem is the CNI, the Service, DNS or the application itself. Run it with netshoot and note what output you would expect in each failure case.
Solutions
Exercise 1
kubectl get nodes -o custom-columns=NODE:.metadata.name,\
POD_CIDR:.spec.podCIDR,IP:.status.addresses[0].address
kubectl create svc clusterip x --tcp=80:80 --clusterip=1.2.3.4 --dry-run=server 2>&1 | tail -1
kubectl get pods -n rutas-norte-pro -o custom-columns=POD:.metadata.name,IP:.status.podIP
kubectl get svc -n rutas-norte-pro -o custom-columns=SVC:.metadata.name,CLUSTERIP:.spec.clusterIPExpected classification:
| Address | Range | Does it exist? |
|---|---|---|
192.168.49.2 |
Node network | Yes, it is the node's eth0 |
10.244.0.31 |
podCIDR |
Yes, it is eth0 inside the pod |
10.96.140.22 |
serviceCIDR |
No: only rules in the kernel |
10.244.0.1 |
podCIDR |
Yes, it is the node bridge (the pods' gateway) |
Exercise 2
IP=$(kubectl get svc redis-cache -n rutas-norte-pro -o jsonpath='{.spec.clusterIP}')
# 1. It is on no interface
minikube -p rutas-norte ssh -- "ip -4 addr | grep $IP || echo 'NOT PRESENT: correct'"
# 2. But the rules are there
minikube -p rutas-norte ssh -- "sudo iptables-save -t nat | grep $IP"
# 3. And yet it connects
kubectl run t --rm -it --restart=Never -n rutas-norte-pro \
--image=nicolaka/netshoot -- nc -zv redis-cache 6379The translation (DNAT) is done by the kernel of the node where the client pod runs, using the rules kube-proxy programmed from the EndpointSlice. The packet leaves the node already carrying the real IP of the destination pod.
Exercise 3
Sequence from the lowest to the highest level, with the conclusion for each failure:
kubectl exec -it deploy/bookings-api -n rutas-norte-pro -- sh
# Plane 1: is the local process listening?
nc -zv localhost 3000 # fails -> the app did not start: not a network issue
# Plane 2: direct pod IP
PGIP=$(kubectl get pod -l app=bookings-postgres -n rutas-norte-pro \
-o jsonpath='{.items[0].status.podIP}')
nc -zv $PGIP 5432 # fails -> CNI or NetworkPolicy
# Plane 3: ClusterIP
nc -zv 10.96.140.22 5432 # fails (and the previous one OK) -> kube-proxy or empty endpoints
kubectl get endpointslices -n rutas-norte-pro -l kubernetes.io/service-name=bookings-postgres
# DNS plane
nslookup bookings-postgres # fails (and the previous one OK) -> CoreDNS or resolv.conf
# Application plane
psql -h bookings-postgres -U bookings -c 'select 1' # fails -> credentials or databaseEvery rung that works rules out a whole layer. The first one that fails names the culprit.
Conclusion
You now know what lies beneath the Kubernetes network. The model rests on four mandatory rules —one IP per pod, pod-to-pod communication without NAT, node agents reaching their own pods, and the IP a pod sees for itself being the one everybody else sees— and it is that uniformity that lets applications remain unaware they are in a cluster and lets ports stop being a resource to administer. You can tell apart the four planes of communication and you know which one to diagnose first when something fails.
You have seen that Kubernetes delegates the implementation to a CNI plugin, which the kubelet invokes through the runtime when creating each pod, and which does a very specific job: create a veth pair, ask IPAM for an IP within the node's podCIDR and programme the routes. You know the real differences between Flannel, Calico, Cilium and Weave Net, and in particular that Flannel does not implement NetworkPolicy, a detail that will shape lesson 04-06. You understand the trade-off between encapsulating with VXLAN (works anywhere, costs 50 bytes and some MTU) and routing natively with BGP (free, but requires cooperation from the physical network), and why eBPF is displacing both approaches in demanding clusters.
And above all, you have dismantled the most useful fiction in Kubernetes: the ClusterIP does not exist. It is on no interface, it does not answer ping and there is no process behind it. There are only iptables rules or IPVS entries that kube-proxy keeps synchronised with the EndpointSlices, doing DNAT on the source node and undoing it in the reply thanks to conntrack. You have followed a packet from bookings-api to bookings-postgres from start to finish and you have verified it on your own minikube.
With the plumbing understood, it is time to go up a floor. In 04-02 we will see that ClusterIP is only one of the four Service types, that each one is built on top of the previous one, and how NodePort, LoadBalancer and ExternalName start opening the platform to the outside: the first step towards letting a customer who wants to buy a coach ticket finally reach web-store.
Kubernetes Course
Module 1: Introduction to Kubernetes
- What Is Kubernetes?
- Kubernetes Architecture
- Key Concepts and Terminology
- Setting Up a Kubernetes Cluster
- The Kubernetes CLI: kubectl
- Objects, YAML Manifests and the Declarative Model
- The Course Project: the Rutas Norte Platform
Module 2: Core Kubernetes Components
- Pods
- ReplicaSets
- Deployments
- Updates, Rollbacks and Deployment Strategies
- Services
- Namespaces
- Labels, Selectors and Annotations
Module 3: Configuration and Secret Management
- ConfigMaps
- Secrets
- Environment Variables
- Resource Quotas and Limits
- LimitRanges and Quality of Service (QoS) Classes
- ServiceAccounts and API Access from Pods
Module 4: Networking in Kubernetes
- Cluster Networking
- Service Types
- Internal DNS and Service Discovery
- Ingress Controllers
- TLS and Certificate Management with cert-manager
- Network Policies
Module 5: Storage in Kubernetes
- Volumes
- Persistent Volumes
- Persistent Volume Claims
- Storage Classes
- Dynamic Provisioning, Expansion and Snapshots
- Backup and Restore of Persistent Data
Module 6: Advanced Kubernetes Concepts
- StatefulSets
- DaemonSets
- Jobs and CronJobs
- Init Containers, Sidecars and Multi-Container Patterns
- Scheduling: Affinity, Taints and Tolerations
- Custom Resource Definitions (CRDs)
- Operators and the Controller Pattern
Module 7: Monitoring and Logging
- Health Checks and Probes
- Metrics Server and kubectl top
- Monitoring with Prometheus
- Visualization and Alerting with Grafana and Alertmanager
- Centralized Logging with Elasticsearch, Fluentd and Kibana (EFK)
- Application Debugging and Cluster Events
Module 8: Kubernetes Security
- Role-Based Access Control (RBAC)
- Security Contexts and Container Hardening
- Pod Security Policies and Pod Security Standards
- Network Security
- Image Security
- Auditing, Scanning and Vulnerability Management
Module 9: Scaling and Performance
- Horizontal Pod Autoscaling
- Vertical Pod Autoscaling
- Cluster Autoscaling
- Event-Driven and Custom-Metric Scaling with KEDA
- High Availability: PodDisruptionBudgets and Topology
- Performance Tuning
Module 10: Kubernetes Ecosystem and Tooling
- Minikube and Local Environments with kind
- Kubeadm
- Helm
- Kustomize
- GitOps with Argo CD and Flux
- Managed Kubernetes: EKS, AKS and GKE
Module 11: Case Studies and Real-World Applications
- Deploying a Web Application
- Running Stateful Applications
- CI/CD with Kubernetes
- Deployment Strategies: Blue-Green and Canary
- Multi-Cluster Management
- Production Operations: Incidents, Runbooks and Costs
