CKS: upgrade Kubernetes — version skew, drain, and the eviction that waits
One minor at a time, API server first, node drained before the kubelet moves.
<!-- hal:authoritative:yaml -->
An unpatched cluster is a known vulnerability with a version number. The exam wants the upgrade done in order, one minor at a time, with the node drained first.
§I. Frame
Cluster Hardening is about 15% of the CKS, and Kubestronaut lists four bullets under it: RBAC, service accounts, API access, and "Upgrade Kubernetes to Avoid Vulnerabilities". The first three have lessons in this stack. This one covers the fourth.
The kubeadm upgrade page gives the security reason in one line: "The Kubernetes project recommends upgrading to the latest patch releases promptly, and to ensure that you are running a supported minor release of Kubernetes. Following this recommendation helps you to stay secure." Patch releases carry the CVE fixes. A cluster that falls off the supported minors stops receiving them.
Bootcamp question 15 frames the task the way the exam does: the control plane is already on the new minor, node01 is one behind, and you write or run the exact commands. Its grader greps for five things: kubectl drain, kubeadm upgrade node, an apt-get install of kubelet, systemctl restart kubelet, and kubectl uncordon. Miss one and the task fails.
§II. The Skew Ledger
Skew Ledger (named technique). Before touching a node, write down each component's minor version and check it against the policy. The version skew page uses 1.37 as its example, and so does this lesson (kubectl on the lab Mac reports v1.37.0):
| Component | Rule against kube-apiserver 1.37 | Allowed |
|---|---|---|
| kubelet | never newer; up to three minors older | 1.37, 1.36, 1.35, 1.34 |
| kube-proxy | never newer; up to three minors older | 1.37, 1.36, 1.35, 1.34 |
| controller-manager, scheduler | never newer; may be one minor older | 1.37, 1.36 |
| kubectl | one minor either way | 1.38, 1.37, 1.36 |
Three rules come out of the table:
- The upgrade order follows from "never newer". The API server goes first, then controller-manager and scheduler, then kubelet, then kube-proxy. A kubelet upgraded before its API server breaks the first row.
- One minor at a time. kubeadm says it directly: "Skipping MINOR versions when upgrading is unsupported." Going from 1.35 to 1.37 is two full passes.
- Old kubelets block the control plane. The skew page warns that kubelets "persistently three minor versions behind kube-apiserver" must be upgraded before the control plane can move again. Letting nodes lag builds a wall in front of the next patch.
KUAR 3e (p. 52) describes the skew as "three versions" for tools and cluster together. That was a looser, older description. On the exam, use the per-component table.
§III. The worker procedure
Run from the control plane unless the step says otherwise. Target is 1.37.x, and the node is node01:
kubectl drain node01 --ignore-daemonsets
ssh node01
sudo -i
# point /etc/apt/sources.list.d/kubernetes.list at the v1.37 pkgs.k8s.io repository
apt-mark unhold kubeadm && apt-get update && apt-get install -y kubeadm='1.37.x-*' && apt-mark hold kubeadm
kubeadm upgrade node
apt-mark unhold kubelet kubectl && apt-get install -y kubelet='1.37.x-*' kubectl='1.37.x-*' && apt-mark hold kubelet kubectl
systemctl daemon-reload
systemctl restart kubelet
exit; exit
kubectl uncordon node01
kubectl get nodes # node01 Ready, VERSION v1.37.x
Four details fail candidates:
- Repository first. The community
pkgs.k8s.iorepositories are split per minor. Without the repo change,apt-get install kubeadm='1.37.x-*'finds nothing. - **kubeadm before
kubeadm upgrade node.** The old binary cannot perform the new upgrade. - **
applyversusnode.**kubeadm upgrade apply v1.37.xruns once, on the first control plane node. Every other control plane node and every worker runskubeadm upgrade node. - **
daemon-reloadbefore restart.** It makes systemd pick up any unit or drop-in files the new kubelet package installed.
The skew page adds the reason drain is not optional: "Before performing a minor version kubelet upgrade, drain pods from that node. In-place minor version kubelet upgrades are not supported."
§IV. The eviction that waits
kubectl drain evicts through the Eviction API, so every PodDisruptionBudget on the node gets a vote. When a budget has no headroom, the API answers 429 and drain retries. From kubectl drain --help on the lab Mac (kubectl 1.37.0):
--timeout=0s: "zero means infinite". A drain against a zero-headroom budget waits forever by default.--disable-eviction: "Force drain to use delete, even if eviction is supported. This will bypass checking PodDisruptionBudgets, use with caution."--force: continue past pods "that do not declare a controller". Those pods are deleted and nothing recreates them.--delete-emptydir-data: continue past pods using emptyDir, whose data is deleted with the pod.
Exam rule: on a stuck drain, find the budget first (kubectl get pdb -A and read ALLOWED DISRUPTIONS). Reach for --disable-eviction only if the task tells you to. It solves the drain by breaking the availability promise someone else wrote. GKE makes the same trade on a timer: the Ops lesson in this trio shows it respects PDBs for up to an hour and then evicts anyway.
§V. Drills
- The API server is at 1.37 and a node's kubelet is at 1.33. Is that supported, and what must happen before the control plane moves to 1.38? (No. 1.33 is four minors behind 1.37. Upgrade that kubelet in single-minor steps before touching the control plane.)
- You ran
kubeadm upgrade apply v1.37.2on the second control plane node. What should you have run? *(kubeadm upgrade node.)* - A drain has printed "Cannot evict pod as it would violate the pod's disruption budget" for ten minutes. Name the first command you run. *(
kubectl get pdb -n <ns>, then read that budget's allowed disruptions and selector.)* - The drain refuses because of a bare pod named
debug-shell. What flag gets past it, and what happens to the pod? *(--force. The pod is deleted and does not come back.)* - A cluster is at 1.35 and the task says "upgrade to 1.37". How many
kubeadm upgrade applyruns on the first control plane node? (Two: 1.36.x, then 1.37.x.)
§VI. Traps
- Upgrading the kubelet before the API server.
- Skipping a minor because the package repository offers it.
- Running
kubeadm upgrade applyon more than one node. - Forgetting
apt-mark hold, so a laterapt-get upgrademoves the kubelet past the API server. - Treating
--disable-evictionas the normal way through a stuck drain.
§VII. Close instruction
On a practice cluster, set a PDB with minAvailable equal to its Deployment's replica count, start kubectl drain --ignore-daemonsets --timeout=90s on a node hosting one of its pods, and record the exit status and message. Then fix the budget, rerun the drain, finish the 1.36 to 1.37 worker procedure from memory, and check your commands against the five things the Bootcamp grader looks for.
Related
- Ops: GKE surge upgrades and PDB stalls (same trio)
- Dev: drain preflight over k8s-openapi types (same trio)
- Prior CKS: SA hardening with pod identity
- Prior CKS: anonymous auth and NodeRestriction