Production GKE Platform — From Zero to Production›02 · GKE architecture: Autopilot, Standard, regional and release channels

Lesson 02 of 18 · Part 1 — Foundations

GKE architecture: Autopilot, Standard, regional and release channels

What Google runs and what you run in GKE, and the four decisions every cluster starts with: Autopilot or Standard, regional or zonal, which release channel, and how the control plane is reached.

Practitioner
Key wordsGKE architectureAutopilotStandardregional clusterzonal clusterrelease channelsRapidRegularStableExtendedcontrol planeSLAshared responsibilityGKE editions
Google-managed project (you don't see it) GKE control plane regional: replicas in 3 zones etcd / cluster state managed and backed up by Google Control-plane endpoints DNS-based and/or IP-based Your project and VPC (you own it) Node pools Compute Engine VMs (Standard mode) Autopilot Google runs nodes, you pay per Pod Pods get VPC IPs from alias ranges VPC-native, Dataplane V2 (eBPF) Google APIs for pods Workload Identity Managed components DNS, CSI, logging, LB private
Google runs the control plane in a project you never see. Your project holds the nodes (or, in Autopilot, the Pods) and your VPC.

What Google runs, what you run

Google You
Control plane: API server, scheduler, controllers, etcd, their upgrades and backups Node pools (Standard): machine types, scaling, node images within GKE's choices
Control-plane availability (regional clusters: replicas across zones) Workloads, their requests and limits, PodDisruptionBudgets
Node image and patches, automatic upgrades (if you allow them) When upgrades may happen: windows, exclusions, channel
Managed components: CSI drivers, DNS, logging agents, load-balancer integration Network design: projects, VPC, IP ranges, firewall rules
In Autopilot also: nodes, their sizing, scaling and security hardening Identity: IAM, RBAC, Workload Identity, policies

GKE is like renting a flat with a building manager. In Standard, you get the flat and choose and look after the furniture (nodes) yourself. In Autopilot, the flat comes furnished and maintained; you just say how many people are living there (Pods) and pay per person.

Decision 1: Autopilot or Standard

Autopilot Standard
Nodes Google creates, sizes, upgrades and secures them You define node pools
Billing Pod resource requests (plus the cluster fee) VMs, whether or not they're full
Security defaults Hardened: no privileged Pods, no host access, Workload Identity on You configure them
Flexibility Restricted: some DaemonSets, host access, kernel settings not allowed Full control (GPUs, special kernels, privileged agents)
Good for Most stateless and standard stateful apps, small platform teams Specialised workloads, third-party agents needing host access, tight bin-packing

Requests matter more in Autopilot: they are what you pay for, and Autopilot adjusts requests that are too small or outside allowed ratios. Recent GKE versions also let Standard clusters run some Pods the Autopilot way through compute classes; check what your version supports.

Decision 2: regional or zonal

  • Regional: control-plane replicas in three zones of a region, nodes spread across zones (the default node count is per zone: --num-nodes 1 means three nodes). Survives a zone outage; the API stays available during control-plane upgrades. Higher SLA.
  • Zonal: one control-plane replica in one zone. Cheaper, fine for development; production loses the API during control-plane upgrades and zone failures.

Autopilot clusters are always regional.

Decision 3: release channel

Channel For
Rapid Trying new features early; not for production
Regular Most clusters: a balance of new features and validation (the default)
Stable Change-averse production; versions arrive later
Extended Staying on a minor version longer than standard support, typically at extra cost

GKE upgrades the control plane and (with node auto-upgrade) the nodes automatically within the channel. You decide when with maintenance windows and exclusions (lesson 13). A common pattern: dev clusters on Regular a little ahead, production on Regular or Stable, so problems surface in dev first.

Decision 4: how the control plane is reached

The control plane has a DNS-based endpoint (access controlled by IAM, reachable without network plumbing) and IP-based endpoints (external and/or internal, optionally restricted by authorised networks). Production clusters usually keep nodes private and limit or disable the external IP endpoint (lesson 05).

Pricing and editions, in short

  • A cluster management fee applies per cluster per hour, with a free-tier credit that covers one zonal or Autopilot cluster per billing account.
  • Standard: you pay for the node VMs. Autopilot: you pay for Pod requests.
  • Fleet and multi-cluster features (fleets, Config Sync, Policy Controller, multi-cluster gateways) have been packaged differently over time (GKE Enterprise); check the current packaging before you design around them.

Always check the current pricing page; numbers and packaging change.

Try it: two clusters, two modes (free trial project)

  1. Create an Autopilot cluster with gcloud container clusters create-auto and a zonal Standard cluster with one small node pool.
  2. Deploy the same app (2 replicas, requests 250m CPU / 256Mi) to both. In Autopilot, watch nodes appear only after the Pods are pending.
  3. Run kubectl get nodes -L topology.kubernetes.io/zone on both and compare.
  4. Try a privileged Pod on Autopilot and read the rejection message.
  5. Look up each cluster's release channel and current version with gcloud container clusters describe. Delete both clusters when done.

Going deeper: choosing per workload

  • Many platforms use both: Autopilot for most application clusters, Standard for clusters that need GPUs, special agents or kernel tuning.
  • In Autopilot, rightsizing requests is cost management. Use VPA recommendations before go-live.
  • Treat the release channel as part of the environment's definition (in Terraform), and keep a version gap between dev and prod on purpose.

Recap

  • Google runs the control plane; you own network, identity, workloads and, in Standard, node pools.
  • Autopilot bills per Pod request and manages nodes; Standard gives full control and bills per VM.
  • Production clusters are regional; zonal is for development.
  • A release channel plus maintenance windows decides which versions arrive and when.

This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.