Lesson 13 of 18 · Part 3 — Operate
Upgrades: release channels, maintenance windows & node upgrade strategies
Run GKE upgrades on your terms: how release channels and auto-upgrades work, maintenance windows and exclusions, surge versus blue-green node pool upgrades, deprecated API checks, upgrade notifications, and sequencing upgrades across environments.
How GKE upgrades work
- Each release channel has a target version. GKE upgrades the control plane automatically (minor and patch versions), and with node auto-upgrade (on by default, required on release channels) the node pools follow.
- Control plane first, then nodes; nodes may lag the control plane by a limited number of minor versions (Kubernetes' version-skew policy), never be newer.
- Regional clusters upgrade control-plane replicas one at a time, so the API stays available.
- Autopilot handles node upgrades entirely.
You don't stop upgrades; you decide when they happen and how nodes roll.
GKE upgrades are like a train timetable you can't cancel but can choose: the trains (new versions) will come. Maintenance windows say which hours the station accepts trains, exclusions close the station for a holiday, and the release channel decides whether you take the early or the later train.
When: windows and exclusions
| Control | Effect |
|---|---|
| Maintenance window | Recurring times when automatic maintenance may run (for example Tue–Thu 02:00–06:00 UTC); GKE needs enough window time per period |
| Maintenance exclusion | A period with no automatic upgrades, with a scope |
Scope no_upgrades |
Nothing at all; limited in length (up to about 30 days), for freezes |
Scope no_minor_upgrades |
Patches still arrive; no new minor version |
Scope no_minor_or_node_upgrades |
Control-plane patches only; nodes untouched |
The longer-scoped exclusions can last until the version's end of support. Check the current limits for your channel; Extended channel has its own rules.
How: node pool upgrade strategies
| Strategy | How it rolls | Needs | Rollback |
|---|---|---|---|
| Surge (default) | Adds up to maxSurge new nodes, drains old ones (respecting PDBs for up to an hour), maxUnavailable can speed it up |
Quota for the surge nodes | None built in: roll forward or start a new pool |
| Blue-green | Creates a full set of new nodes (green), cordons the old (blue), drains in batches, soaks, then deletes blue | Quota for a second set of nodes | During the soak: roll back to blue |
| Autoscaled blue-green (newer versions) | Like blue-green, but scales green with demand | Less spare capacity | During the soak |
Good settings for production surge upgrades: maxSurge 1–2 per zone, maxUnavailable 0, PDBs on every important workload, and enough Pod IP space for the extra nodes (lesson 04).
Before a minor upgrade
- Deprecated APIs: check the GKE deprecation insights and recommendations; GKE pauses automatic minor upgrades when it sees calls to APIs removed in the next version. Fix the callers (often old Helm charts or controllers).
- Read the GKE and Kubernetes release notes for the new minor (changed defaults, removed features).
- Check third-party controllers (cert-manager, Argo CD, service mesh, operators) for support of the new version; upgrade them first.
- PDBs present, not blocking (
minAvailablelower than replicas). - Let it happen in dev first: keep dev on a faster channel or upgrade it manually ahead of prod.
Know what's coming
- Upgrade notifications through Pub/Sub (
--notification-config): available upgrades, scheduled upgrades, security bulletins, end of support. - The console's upgrade view shows the next scheduled automatic upgrade per cluster.
- You can upgrade manually ahead of time (
gcloud container clusters upgrade --master) when you want to control the exact moment, for example right after a dev soak.
Sequencing environments
A common rollout:
| Environment | Channel | Upgrade timing |
|---|---|---|
| dev | Regular (or Rapid) | Automatic, or manual as soon as a version is available |
| staging | Regular | Automatic, a week or two later (manual trigger after dev soak) |
| prod | Regular or Stable | Automatic in narrow windows, with exclusions around business peaks |
Rollout sequencing across a fleet (GKE can upgrade groups of clusters in a defined order with soak time between them) automates this pattern; check its availability for your setup.
Try it: an upgrade on your terms (lab project)
- Create a Standard cluster on the Regular channel at an older available minor version (pick one from
get-server-config). - Set a maintenance window for tomorrow night and a
no_minor_upgradesexclusion for this week; read the result withdescribe. - Deploy an app with 3 replicas, a PDB and a curl loop counting failures.
- Upgrade the control plane by hand, then the node pool with surge (maxSurge 1, maxUnavailable 0) and count failed requests.
- Switch the pool to blue-green with a 10-minute soak, trigger another node upgrade, and roll back during the soak.
Going deeper: upgrades as a product
- Track version lag per cluster and days to patch after a security bulletin as platform metrics.
- Keep cluster settings, windows and exclusions in Terraform; changes are reviewed like code.
- For risky changes (several minors behind, network data plane change), build a new cluster and move workloads with GitOps and multi-cluster gateways, instead of an in-place jump.
Recap
- GKE upgrades automatically within the release channel: control plane first, then nodes.
- Decide when with maintenance windows and exclusions (scopes matter).
- Decide how nodes roll: surge (cheap, no rollback) or blue-green (soak and rollback, needs capacity).
- Before minors: deprecated APIs, release notes, third-party controllers, PDBs; dev first.
- Get notifications and sequence environments deliberately.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.