Hedronite · Ops Lesson · 01-Earth-DevOps / Kubernetes · Sun 2026-09-27

GKE surge upgrades — the PDB that holds a node for an hour

A PodDisruptionBudget can make GKE wait. It cannot make GKE stop.

Lesson Class: Ops (DevOps + Kubernetes + GKE node pool upgrades)
Topic: T2 Ops scripting
Cloud Referent: GKE surge upgrade · maxSurge / maxUnavailable · one-hour PDB window
Automation: cargo script · kube 2.0.1 · k8s-openapi 0.26.1 · Nix writeBashBin
Paired Dev: Rust drain preflight over k8s-openapi
Paired Cert: CKS upgrade, version skew, drain
Budget Headroom
Read disruptionsAllowed, not minAvailable.
One hour
GKE respects the PDB for sixty minutes, then evicts.
Server Selector
Let the API server match the PDB selector.
Give the budget headroom, or the upgrade takes it.

<!-- hal:authoritative:yaml -->

A PodDisruptionBudget can make GKE wait. It cannot make GKE stop. Find the budgets with no headroom before the upgrade finds them for you.

§I. Frame

A GKE Standard node pool upgrades by surge unless you chose blue-green. The GKE upgrade-strategies page lists the steps for one node with the default surge settings: provision a new node, wait for it to be ready, cordon the old node, then drain it, "respecting PodDisruptionBudget and GracefulTerminationPeriod settings for up to one hour. After one hour, any remaining Pods are forcefully evicted so that the upgrade can proceed." Then GKE deletes the old node.

That sentence sets the real contract. A PodDisruptionBudget (PDB) turns each eviction into a request that the API server can refuse with 429 Too Many Requests. GKE keeps asking. After sixty minutes it stops asking, and the pods leave anyway.

So a PDB with no headroom does two things during a surge upgrade. It stretches the upgrade by up to an hour for every node that runs a covered pod. Then it lets through the exact disruption it was written to block, only later and all at once.

The problem for today: set the surge knobs on a GKE pool deliberately, then build a read-only Rust census that names every PDB that will hold a node in that pool before anyone starts the upgrade.

§II. Budget Headroom

Budget Headroom (named technique). Read a PDB's status, not its spec. The disruption controller publishes four numbers on every PDB:

  1. expectedPods, the pods the budget covers.
  2. currentHealthy, how many of those are Ready now.
  3. desiredHealthy, the floor the spec implies.
  4. disruptionsAllowed, the headroom: current healthy minus the floor, never below zero.

An eviction succeeds only while disruptionsAllowed is above zero. The common mistake is minAvailable equal to the replica count:

apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: api-pdb
  namespace: prod
spec:
  minAvailable: 3   # the StatefulSet also runs 3 replicas
  selector:
    matchLabels:
      app: api

With three healthy replicas the floor is three, so the headroom is zero, and it stays zero while every pod is healthy. In practice the budget forbids every eviction, and on GKE that ban lasts one hour per node.

The fix is headroom, stated one of two ways: maxUnavailable: 1, or minAvailable one below the replica count. Either lets the drain move one pod, wait for the replacement to become Ready, and move the next.

§III. The surge knobs

Surge behavior is set per node pool with two numbers:

gcloud container node-pools update pool-a \
  --cluster=prod \
  --max-surge-upgrade=1 \
  --max-unavailable-upgrade=0

maxSurge is how many new nodes GKE may create ahead of removing old ones. maxUnavailable is how many existing nodes may be down at once without a replacement. GKE's own example is a five-node pool at maxSurge=2;maxUnavailable=1: GKE creates two upgraded nodes, disrupts at most one existing node at a time, and the pool holds between four and seven nodes during the upgrade. GKE upgrades maxSurge + maxUnavailable nodes at once, capped at 100 for Standard and 20 for Autopilot.

Three consequences follow:

  1. Parallelism multiplies the stall. Raising maxSurge drains more nodes at once. If a covered pod sits on each of them, each drain waits on the same zero-headroom budget.
  2. **maxUnavailable above zero removes capacity before replacements exist.** Evicted pods may go Pending, which keeps currentHealthy low and the headroom at zero for longer.
  3. Blue-green is not an escape hatch. Its delete-blue-pool phase "does not use eviction and instead attempts to delete the Pods. Unlike eviction, deletion doesn't respect PDBs." A longer soak gives pods more time to leave on their own, but the last phase still deletes whatever remains.

GKE also warns that parallel drains do not work with externalTrafficPolicy: Local. Check that before raising either number.

§IV. The census, in one cargo script

gke-pdb-upgrade-census.rs sits in this bundle. It takes a node pool name, finds that pool's nodes, and reports every zero-headroom PDB that covers a pod on them:

#!/usr/bin/env -S cargo +nightly -Zscript
---
[package]
edition = "2024"

[dependencies]
kube = "2"
k8s-openapi = { version = "0.26", features = ["v1_34"] }
tokio = { version = "1", features = ["macros", "rt-multi-thread"] }
anyhow = "1"
---

use k8s_openapi::api::core::v1::{Node, Pod};
use k8s_openapi::api::policy::v1::PodDisruptionBudget;
use kube::api::{Api, ListParams};
use kube::core::Selector;
use kube::{Client, ResourceExt};
use std::collections::BTreeSet;

#[tokio::main]
async fn main() -> anyhow::Result<()> {
    let pool = std::env::args()
        .nth(1)
        .ok_or_else(|| anyhow::anyhow!("usage: gke-pdb-upgrade-census NODE_POOL"))?;
    let client = Client::try_default().await?;

    // 1. Nodes in the pool. GKE labels every node with its pool name.
    let nodes: Api<Node> = Api::all(client.clone());
    let lp = ListParams::default().labels(&format!("cloud.google.com/gke-nodepool={pool}"));
    let pool_nodes: BTreeSet<String> = nodes.list(&lp).await?.iter().map(|n| n.name_any()).collect();
    println!("pool={pool} nodes={}", pool_nodes.len());

    // 2. Every PDB, then only the ones that cannot absorb a disruption right now.
    let pdbs: Api<PodDisruptionBudget> = Api::all(client.clone());
    let mut stalls = 0;
    for pdb in pdbs.list(&ListParams::default()).await? {
        let status = pdb.status.clone().unwrap_or_default();
        if status.disruptions_allowed > 0 {
            continue;
        }
        let ns = pdb.namespace().unwrap_or_default();
        let Some(sel) = pdb.spec.as_ref().and_then(|s| s.selector.clone()) else {
            continue;
        };
        // 3. Let the API server evaluate the PDB's own selector.
        let selector = Selector::try_from(sel)?;
        let pods: Api<Pod> = Api::namespaced(client.clone(), &ns);
        let covered = pods.list(&ListParams::default().labels_from(&selector)).await?;
        let on_pool: Vec<String> = covered
            .iter()
            .filter(|p| {
                p.spec
                    .as_ref()
                    .and_then(|s| s.node_name.as_ref())
                    .is_some_and(|n| pool_nodes.contains(n))
            })
            .map(|p| p.name_any())
            .collect();
        if on_pool.is_empty() {
            continue;
        }
        stalls += 1;
        println!(
            "STALL {ns}/{} healthy={}/{} expected={} pods_on_pool=[{}]",
            pdb.name_any(),
            status.current_healthy,
            status.desired_healthy,
            status.expected_pods,
            on_pool.join(",")
        );
    }
    println!("stalling_pdbs={stalls} (each can hold a node for up to 60 min, then GKE evicts anyway)");
    Ok(())
}

Server Selector (named technique). The census never re-implements label matching. Selector::try_from converts the PDB's LabelSelector, including matchExpressions, into kube's selector type, and labels_from sends it as the list call's labelSelector query. The API server applies the same matching rules the disruption controller uses, so the census and the controller cannot disagree about which pods a budget covers. The Dev lesson in this trio does the matching by hand on offline JSON, and it has to test each operator to earn the same trust.

Three smaller mechanics:

What was checked, and what was not. A scratch crate built from the frontmatter manifest and body passed cargo check on stable cargo 1.96.0 with zero warnings, resolving kube 2.0.1, k8s-openapi 0.26.1 and tokio 1.53.1. The -Zscript entry also ran on the lab Mac under nightly cargo 1.101.0, and the first version failed in a way worth keeping. With kube = { default-features = false, features = ["client", "rustls-tls"] } it compiled cleanly and then panicked at startup: rustls "Could not automatically determine the process-level CryptoProvider". Trimming default features dropped kube's ring feature, and cargo check cannot see a missing runtime provider. Plain kube = "2" restores it. The fixed script loaded ~/.kube/config, found the local container runtime context at https://127.0.0.1:26443, and stopped at "Connection refused" because no cluster was running. It was not run against GKE. Output from a pool where one StatefulSet has zero headroom would read like this (illustrative):

pool=pool-a nodes=6
STALL prod/api-pdb healthy=3/3 expected=3 pods_on_pool=[api-0,api-1,api-2]
stalling_pdbs=1 (each can hold a node for up to 60 min, then GKE evicts anyway)

With maxSurge=1 and maxUnavailable=0, nodes drain one at a time, so three pods on three nodes can mean up to three hours of waiting, followed by three forced evictions.

§V. Wrap it with Nix

gke-pdb-upgrade-census.nix gives the script a stable command name:

{ pkgs ? import <nixpkgs> { } }:
pkgs.writers.writeBashBin "gke-pdb-upgrade-census" ''
  exec cargo +nightly -Zscript ${./gke-pdb-upgrade-census.rs} "$@"
''

Built against nixpkgs d54020a6, it produced /nix/store/37d72h6v…-gke-pdb-upgrade-census. The generated bin/gke-pdb-upgrade-census is two lines: a store-path bash shebang and exec cargo +nightly -Zscript /nix/store/l4bivvq7…-gke-pdb-upgrade-census.rs "$@". The wrapper still does not pin cargo. On the lab Mac the first cargo +nightly call this fire made rustup install a nightly toolchain on its own (13:30 ET), which is exactly the kind of unpinned drift a rust-bin overlay in the wrapper would prevent.

§VI. What not to do

  1. Writing minAvailable equal to the replica count and calling the workload protected. On GKE it buys one hour per node, then nothing.
  2. Raising maxSurge to speed up an upgrade without running the census first.
  3. Switching to blue-green to get around a PDB. The final phase deletes pods and ignores budgets.
  4. Trimming kube's default features without running the binary once. The crypto provider is a runtime choice.
  5. Giving the census write verbs, or an evict permission "for later".

§VII. Close instruction

Change api-pdb so the drain can move one pod at a time, and state the new desiredHealthy and disruptionsAllowed for three healthy replicas. Then pick maxSurge and maxUnavailable for a six-node pool that must never drop below six schedulable nodes, and say how many nodes upgrade at once.

Related