Lesson 02 of 18 · Part 1 — Foundations
GKE architecture: Autopilot, Standard, regional and release channels
What Google runs and what you run in GKE, and the four decisions every cluster starts with: Autopilot or Standard, regional or zonal, which release channel, and how the control plane is reached.
What Google runs, what you run
| You | |
|---|---|
| Control plane: API server, scheduler, controllers, etcd, their upgrades and backups | Node pools (Standard): machine types, scaling, node images within GKE's choices |
| Control-plane availability (regional clusters: replicas across zones) | Workloads, their requests and limits, PodDisruptionBudgets |
| Node image and patches, automatic upgrades (if you allow them) | When upgrades may happen: windows, exclusions, channel |
| Managed components: CSI drivers, DNS, logging agents, load-balancer integration | Network design: projects, VPC, IP ranges, firewall rules |
| In Autopilot also: nodes, their sizing, scaling and security hardening | Identity: IAM, RBAC, Workload Identity, policies |
GKE is like renting a flat with a building manager. In Standard, you get the flat and choose and look after the furniture (nodes) yourself. In Autopilot, the flat comes furnished and maintained; you just say how many people are living there (Pods) and pay per person.
Decision 1: Autopilot or Standard
| Autopilot | Standard | |
|---|---|---|
| Nodes | Google creates, sizes, upgrades and secures them | You define node pools |
| Billing | Pod resource requests (plus the cluster fee) | VMs, whether or not they're full |
| Security defaults | Hardened: no privileged Pods, no host access, Workload Identity on | You configure them |
| Flexibility | Restricted: some DaemonSets, host access, kernel settings not allowed | Full control (GPUs, special kernels, privileged agents) |
| Good for | Most stateless and standard stateful apps, small platform teams | Specialised workloads, third-party agents needing host access, tight bin-packing |
Requests matter more in Autopilot: they are what you pay for, and Autopilot adjusts requests that are too small or outside allowed ratios. Recent GKE versions also let Standard clusters run some Pods the Autopilot way through compute classes; check what your version supports.
Decision 2: regional or zonal
- Regional: control-plane replicas in three zones of a region, nodes spread across zones (the default node count is per zone:
--num-nodes 1means three nodes). Survives a zone outage; the API stays available during control-plane upgrades. Higher SLA. - Zonal: one control-plane replica in one zone. Cheaper, fine for development; production loses the API during control-plane upgrades and zone failures.
Autopilot clusters are always regional.
Decision 3: release channel
| Channel | For |
|---|---|
| Rapid | Trying new features early; not for production |
| Regular | Most clusters: a balance of new features and validation (the default) |
| Stable | Change-averse production; versions arrive later |
| Extended | Staying on a minor version longer than standard support, typically at extra cost |
GKE upgrades the control plane and (with node auto-upgrade) the nodes automatically within the channel. You decide when with maintenance windows and exclusions (lesson 13). A common pattern: dev clusters on Regular a little ahead, production on Regular or Stable, so problems surface in dev first.
Decision 4: how the control plane is reached
The control plane has a DNS-based endpoint (access controlled by IAM, reachable without network plumbing) and IP-based endpoints (external and/or internal, optionally restricted by authorised networks). Production clusters usually keep nodes private and limit or disable the external IP endpoint (lesson 05).
Pricing and editions, in short
- A cluster management fee applies per cluster per hour, with a free-tier credit that covers one zonal or Autopilot cluster per billing account.
- Standard: you pay for the node VMs. Autopilot: you pay for Pod requests.
- Fleet and multi-cluster features (fleets, Config Sync, Policy Controller, multi-cluster gateways) have been packaged differently over time (GKE Enterprise); check the current packaging before you design around them.
Always check the current pricing page; numbers and packaging change.
Try it: two clusters, two modes (free trial project)
- Create an Autopilot cluster with
gcloud container clusters create-autoand a zonal Standard cluster with one small node pool. - Deploy the same app (2 replicas, requests 250m CPU / 256Mi) to both. In Autopilot, watch nodes appear only after the Pods are pending.
- Run
kubectl get nodes -L topology.kubernetes.io/zoneon both and compare. - Try a privileged Pod on Autopilot and read the rejection message.
- Look up each cluster's release channel and current version with
gcloud container clusters describe. Delete both clusters when done.
Going deeper: choosing per workload
- Many platforms use both: Autopilot for most application clusters, Standard for clusters that need GPUs, special agents or kernel tuning.
- In Autopilot, rightsizing requests is cost management. Use VPA recommendations before go-live.
- Treat the release channel as part of the environment's definition (in Terraform), and keep a version gap between dev and prod on purpose.
Recap
- Google runs the control plane; you own network, identity, workloads and, in Standard, node pools.
- Autopilot bills per Pod request and manages nodes; Standard gives full control and bills per VM.
- Production clusters are regional; zonal is for development.
- A release channel plus maintenance windows decides which versions arrive and when.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.