Lesson 10 of 18 · Part 2 — Build the platform
Autoscaling: Pods, nodes and compute classes
Make GKE follow demand: Pods with the HPA and VPA, nodes with the cluster autoscaler and node auto-provisioning, preferences with compute classes, and Autopilot's per-Pod model, plus the settings that decide cost and how fast scale-up really is.
Three levels
| Level | Scales | Tool |
|---|---|---|
| Pods (count) | Replicas from load | HPA (CPU, memory, custom or external metrics) |
| Pods (size) | Requests from observed usage | VPA (built into GKE), or multidimensional Pod autoscaling |
| Nodes | Capacity from pending Pods | Cluster autoscaler, node auto-provisioning, compute classes; automatic in Autopilot |
Think of a restaurant. The HPA calls in more waiters when the room fills up. The VPA notices a waiter carries more plates than expected and gives them a bigger tray. The cluster autoscaler opens another dining room when there are no tables left, and closes it again when it has been empty for a while.
Pods: HPA and VPA
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: shop
namespace: shop
spec:
scaleTargetRef: { apiVersion: apps/v1, kind: Deployment, name: shop }
minReplicas: 3
maxReplicas: 30
metrics:
- type: Resource
resource:
name: cpu
target: { type: Utilization, averageUtilization: 70 }
behavior:
scaleDown:
stabilizationWindowSeconds: 300
- The HPA can also scale on Cloud Monitoring metrics (queue depth, requests per second from the load balancer) through the custom metrics adapter, or on Prometheus metrics from Managed Service for Prometheus.
- The VPA in
updateMode: "Off"gives request recommendations without restarting anything: the safest way to rightsize. InAutoit recreates Pods with new requests. - Don't let the HPA and VPA both act on CPU for the same workload.
Nodes: cluster autoscaler
The cluster autoscaler adds nodes when Pods are unschedulable because of their requests, and removes nodes whose Pods fit elsewhere:
| Setting | Effect |
|---|---|
| Per-pool min/max | Per zone in regional clusters (--total-min-nodes/--total-max-nodes set totals instead) |
| Location policy | BALANCED spreads across zones; ANY takes capacity wherever it exists (good for Spot) |
| Autoscaling profile | balanced (default) or optimize-utilization (removes nodes sooner, packs tighter) |
Pod annotation cluster-autoscaler.kubernetes.io/safe-to-evict: "false" |
Blocks scale-down of that Pod's node; use sparingly |
Things that block scale-down: PDBs with no room, Pods with local storage, Pods without a controller, the safe-to-evict annotation, system Pods without PDBs.
Node auto-provisioning and compute classes
With node auto-provisioning (NAP), GKE creates node pools sized for the pending Pods (CPU, memory, GPUs, Spot, zones) within cluster-wide limits, and deletes them when empty. Compute classes describe preferences, for example:
apiVersion: cloud.google.com/v1
kind: ComputeClass
metadata:
name: cost-first
spec:
priorities:
- machineFamily: n4
spot: true
- machineFamily: n4
spot: false
nodePoolAutoCreation:
enabled: true
whenUnsatisfiable: ScaleUpAnyway
Workloads pick it with nodeSelector: { cloud.google.com/compute-class: cost-first }: Spot first, on-demand when Spot isn't available. The CRD's fields have changed across GKE versions; use the reference for yours.
Autopilot
In Autopilot there are no pools to size. Pods are scheduled, GKE adds the capacity and bills for the Pods' requests. Your levers are requests (rightsize them with VPA recommendations), the compute class (general purpose, Scale-Out, Balanced, Spot, GPU) and the HPA.
How fast is scale-up, really?
Pending Pod → autoscaler decision (seconds) → VM created and joined (often 1–2 minutes) → image pulled (seconds to minutes) → readiness. To shorten it:
- Image streaming and small images.
- Overprovisioning: a low-priority "balloon" Deployment that reserves headroom and gets evicted first, so real Pods start instantly while new nodes come up.
- Scale on leading metrics (queue depth, requests) instead of CPU, which reacts late.
Cost
- In Standard, cost follows nodes: bin-packing, the
optimize-utilizationprofile, Spot pools and rightsized requests. - In Autopilot, cost follows requests: rightsizing is the main lever.
- Committed use discounts for the steady baseline, Spot for the bursty part.
- Use GKE cost allocation (cost per namespace and label in billing exports) so teams see what they spend.
Try it: watch every layer move
- On a Standard lab cluster, autoscale a pool from 1 to 5 nodes and deploy an app with an HPA at 50% CPU.
- Generate load (a
busyboxloop orhey) and watchkubectl get hpa -w, then pending Pods, then new nodes. - Stop the load and time how long until nodes are removed; switch to
optimize-utilizationand repeat. - Enable the VPA in
Offmode for the app and read its recommendations after the load test. - Add a balloon Deployment with a negative-priority PriorityClass and measure how much faster new Pods start.
Going deeper: scaling as a design input
- Set max limits everywhere (HPA, pools, NAP limits): an autoscaler without a ceiling turns a bug or an attack into a bill.
- Check quotas (CPUs per region, IPs, Pod range) before big scale events; the autoscaler can't create what quota or the IP plan forbids (lesson 04).
- Use multidimensional Pod autoscaling where you want CPU-based horizontal and memory-based vertical scaling together.
Recap
- HPA scales replicas, VPA rightsizes requests; don't let both act on CPU.
- The cluster autoscaler reacts to pending Pods, not CPU; NAP and compute classes choose machine shapes.
- Autopilot scales automatically and bills per request.
- Speed up scale-up with image streaming, overprovisioning and leading metrics; control cost with Spot, profiles, commitments and rightsizing.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.