Production GKE Platform — From Zero to Production›10 · Autoscaling: Pods, nodes and compute classes

Lesson 10 of 18 · Part 2 — Build the platform

Autoscaling: Pods, nodes and compute classes

Make GKE follow demand: Pods with the HPA and VPA, nodes with the cluster autoscaler and node auto-provisioning, preferences with compute classes, and Autopilot's per-Pod model, plus the settings that decide cost and how fast scale-up really is.

Practitioner → Advanced
Key wordsGKE autoscalingHPAVPAmultidimensional Pod autoscalingcluster autoscalerautoscaling profileoptimize-utilizationnode auto-provisioningcompute classesAutopilot scalingSpot VMslocation policycost optimisationimage streaming

Three levels

Level Scales Tool
Pods (count) Replicas from load HPA (CPU, memory, custom or external metrics)
Pods (size) Requests from observed usage VPA (built into GKE), or multidimensional Pod autoscaling
Nodes Capacity from pending Pods Cluster autoscaler, node auto-provisioning, compute classes; automatic in Autopilot

Think of a restaurant. The HPA calls in more waiters when the room fills up. The VPA notices a waiter carries more plates than expected and gives them a bigger tray. The cluster autoscaler opens another dining room when there are no tables left, and closes it again when it has been empty for a while.

Pods: HPA and VPA

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: shop
  namespace: shop
spec:
  scaleTargetRef: { apiVersion: apps/v1, kind: Deployment, name: shop }
  minReplicas: 3
  maxReplicas: 30
  metrics:
    - type: Resource
      resource:
        name: cpu
        target: { type: Utilization, averageUtilization: 70 }
  behavior:
    scaleDown:
      stabilizationWindowSeconds: 300
  • The HPA can also scale on Cloud Monitoring metrics (queue depth, requests per second from the load balancer) through the custom metrics adapter, or on Prometheus metrics from Managed Service for Prometheus.
  • The VPA in updateMode: "Off" gives request recommendations without restarting anything: the safest way to rightsize. In Auto it recreates Pods with new requests.
  • Don't let the HPA and VPA both act on CPU for the same workload.

Nodes: cluster autoscaler

The cluster autoscaler adds nodes when Pods are unschedulable because of their requests, and removes nodes whose Pods fit elsewhere:

Setting Effect
Per-pool min/max Per zone in regional clusters (--total-min-nodes/--total-max-nodes set totals instead)
Location policy BALANCED spreads across zones; ANY takes capacity wherever it exists (good for Spot)
Autoscaling profile balanced (default) or optimize-utilization (removes nodes sooner, packs tighter)
Pod annotation cluster-autoscaler.kubernetes.io/safe-to-evict: "false" Blocks scale-down of that Pod's node; use sparingly

Things that block scale-down: PDBs with no room, Pods with local storage, Pods without a controller, the safe-to-evict annotation, system Pods without PDBs.

Node auto-provisioning and compute classes

With node auto-provisioning (NAP), GKE creates node pools sized for the pending Pods (CPU, memory, GPUs, Spot, zones) within cluster-wide limits, and deletes them when empty. Compute classes describe preferences, for example:

apiVersion: cloud.google.com/v1
kind: ComputeClass
metadata:
  name: cost-first
spec:
  priorities:
    - machineFamily: n4
      spot: true
    - machineFamily: n4
      spot: false
  nodePoolAutoCreation:
    enabled: true
  whenUnsatisfiable: ScaleUpAnyway

Workloads pick it with nodeSelector: { cloud.google.com/compute-class: cost-first }: Spot first, on-demand when Spot isn't available. The CRD's fields have changed across GKE versions; use the reference for yours.

Autopilot

In Autopilot there are no pools to size. Pods are scheduled, GKE adds the capacity and bills for the Pods' requests. Your levers are requests (rightsize them with VPA recommendations), the compute class (general purpose, Scale-Out, Balanced, Spot, GPU) and the HPA.

How fast is scale-up, really?

Pending Pod → autoscaler decision (seconds) → VM created and joined (often 1–2 minutes) → image pulled (seconds to minutes) → readiness. To shorten it:

  • Image streaming and small images.
  • Overprovisioning: a low-priority "balloon" Deployment that reserves headroom and gets evicted first, so real Pods start instantly while new nodes come up.
  • Scale on leading metrics (queue depth, requests) instead of CPU, which reacts late.

Cost

  • In Standard, cost follows nodes: bin-packing, the optimize-utilization profile, Spot pools and rightsized requests.
  • In Autopilot, cost follows requests: rightsizing is the main lever.
  • Committed use discounts for the steady baseline, Spot for the bursty part.
  • Use GKE cost allocation (cost per namespace and label in billing exports) so teams see what they spend.

Try it: watch every layer move

  1. On a Standard lab cluster, autoscale a pool from 1 to 5 nodes and deploy an app with an HPA at 50% CPU.
  2. Generate load (a busybox loop or hey) and watch kubectl get hpa -w, then pending Pods, then new nodes.
  3. Stop the load and time how long until nodes are removed; switch to optimize-utilization and repeat.
  4. Enable the VPA in Off mode for the app and read its recommendations after the load test.
  5. Add a balloon Deployment with a negative-priority PriorityClass and measure how much faster new Pods start.

Going deeper: scaling as a design input

  • Set max limits everywhere (HPA, pools, NAP limits): an autoscaler without a ceiling turns a bug or an attack into a bill.
  • Check quotas (CPUs per region, IPs, Pod range) before big scale events; the autoscaler can't create what quota or the IP plan forbids (lesson 04).
  • Use multidimensional Pod autoscaling where you want CPU-based horizontal and memory-based vertical scaling together.

Recap

  • HPA scales replicas, VPA rightsizes requests; don't let both act on CPU.
  • The cluster autoscaler reacts to pending Pods, not CPU; NAP and compute classes choose machine shapes.
  • Autopilot scales automatically and bills per request.
  • Speed up scale-up with image streaming, overprovisioning and leading metrics; control cost with Spot, profiles, commitments and rightsizing.

This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.