Production GKE Platform — From Zero to Production›13 · Upgrades: release channels, maintenance windows & node upgrade strategies

Lesson 13 of 18 · Part 3 — Operate

Upgrades: release channels, maintenance windows & node upgrade strategies

Run GKE upgrades on your terms: how release channels and auto-upgrades work, maintenance windows and exclusions, surge versus blue-green node pool upgrades, deprecated API checks, upgrade notifications, and sequencing upgrades across environments.

Advanced
Key wordsGKE upgradesrelease channelsauto-upgrademaintenance windowsmaintenance exclusionssurge upgradeblue-green upgradenode pool upgradedeprecated APIsdeprecation insightsversion skewupgrade notificationsrollout sequencing

How GKE upgrades work

  • Each release channel has a target version. GKE upgrades the control plane automatically (minor and patch versions), and with node auto-upgrade (on by default, required on release channels) the node pools follow.
  • Control plane first, then nodes; nodes may lag the control plane by a limited number of minor versions (Kubernetes' version-skew policy), never be newer.
  • Regional clusters upgrade control-plane replicas one at a time, so the API stays available.
  • Autopilot handles node upgrades entirely.

You don't stop upgrades; you decide when they happen and how nodes roll.

GKE upgrades are like a train timetable you can't cancel but can choose: the trains (new versions) will come. Maintenance windows say which hours the station accepts trains, exclusions close the station for a holiday, and the release channel decides whether you take the early or the later train.

When: windows and exclusions

Control Effect
Maintenance window Recurring times when automatic maintenance may run (for example Tue–Thu 02:00–06:00 UTC); GKE needs enough window time per period
Maintenance exclusion A period with no automatic upgrades, with a scope
Scope no_upgrades Nothing at all; limited in length (up to about 30 days), for freezes
Scope no_minor_upgrades Patches still arrive; no new minor version
Scope no_minor_or_node_upgrades Control-plane patches only; nodes untouched

The longer-scoped exclusions can last until the version's end of support. Check the current limits for your channel; Extended channel has its own rules.

How: node pool upgrade strategies

Strategy How it rolls Needs Rollback
Surge (default) Adds up to maxSurge new nodes, drains old ones (respecting PDBs for up to an hour), maxUnavailable can speed it up Quota for the surge nodes None built in: roll forward or start a new pool
Blue-green Creates a full set of new nodes (green), cordons the old (blue), drains in batches, soaks, then deletes blue Quota for a second set of nodes During the soak: roll back to blue
Autoscaled blue-green (newer versions) Like blue-green, but scales green with demand Less spare capacity During the soak

Good settings for production surge upgrades: maxSurge 1–2 per zone, maxUnavailable 0, PDBs on every important workload, and enough Pod IP space for the extra nodes (lesson 04).

Before a minor upgrade

  1. Deprecated APIs: check the GKE deprecation insights and recommendations; GKE pauses automatic minor upgrades when it sees calls to APIs removed in the next version. Fix the callers (often old Helm charts or controllers).
  2. Read the GKE and Kubernetes release notes for the new minor (changed defaults, removed features).
  3. Check third-party controllers (cert-manager, Argo CD, service mesh, operators) for support of the new version; upgrade them first.
  4. PDBs present, not blocking (minAvailable lower than replicas).
  5. Let it happen in dev first: keep dev on a faster channel or upgrade it manually ahead of prod.

Know what's coming

  • Upgrade notifications through Pub/Sub (--notification-config): available upgrades, scheduled upgrades, security bulletins, end of support.
  • The console's upgrade view shows the next scheduled automatic upgrade per cluster.
  • You can upgrade manually ahead of time (gcloud container clusters upgrade --master) when you want to control the exact moment, for example right after a dev soak.

Sequencing environments

A common rollout:

Environment Channel Upgrade timing
dev Regular (or Rapid) Automatic, or manual as soon as a version is available
staging Regular Automatic, a week or two later (manual trigger after dev soak)
prod Regular or Stable Automatic in narrow windows, with exclusions around business peaks

Rollout sequencing across a fleet (GKE can upgrade groups of clusters in a defined order with soak time between them) automates this pattern; check its availability for your setup.

Try it: an upgrade on your terms (lab project)

  1. Create a Standard cluster on the Regular channel at an older available minor version (pick one from get-server-config).
  2. Set a maintenance window for tomorrow night and a no_minor_upgrades exclusion for this week; read the result with describe.
  3. Deploy an app with 3 replicas, a PDB and a curl loop counting failures.
  4. Upgrade the control plane by hand, then the node pool with surge (maxSurge 1, maxUnavailable 0) and count failed requests.
  5. Switch the pool to blue-green with a 10-minute soak, trigger another node upgrade, and roll back during the soak.

Going deeper: upgrades as a product

  • Track version lag per cluster and days to patch after a security bulletin as platform metrics.
  • Keep cluster settings, windows and exclusions in Terraform; changes are reviewed like code.
  • For risky changes (several minors behind, network data plane change), build a new cluster and move workloads with GitOps and multi-cluster gateways, instead of an in-place jump.

Recap

  • GKE upgrades automatically within the release channel: control plane first, then nodes.
  • Decide when with maintenance windows and exclusions (scopes matter).
  • Decide how nodes roll: surge (cheap, no rollback) or blue-green (soak and rollback, needs capacity).
  • Before minors: deprecated APIs, release notes, third-party controllers, PDBs; dev first.
  • Get notifications and sequence environments deliberately.

This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.