Production GKE Platform — From Zero to Production›05 · The cluster: control plane access, node pools & compute

Lesson 05 of 18 · Part 2 — Build the platform

The cluster: control plane access, node pools & compute

Create a production GKE cluster: private nodes, how people and pipelines reach the control plane, node pools with the right machine types, images and labels, a minimal node service account, Spot VMs, and compute classes that tell workloads where to run.

Practitioner → Advanced
Key wordsprivate clusterprivate nodescontrol plane endpointDNS-based endpointauthorized networksnode poolsmachine typesContainer-Optimized OSSpot VMsnode service accounttaintscompute classesGPU node pools

The control plane: how it's reached

Endpoint Access controlled by Typical use
DNS-based IAM (the caller needs permission on the cluster), from anywhere that can reach Google APIs People and CI without VPN or bastions
IP-based, internal Network reachability inside the VPC (and hybrid links), authorised networks Access from on-prem or a jump host
IP-based, external Authorised networks (CIDR allowlist) Legacy setups; disable it in production if you can

All of them still require authentication (IAM) and authorisation (IAM + RBAC). Older guides talk about "private clusters" with a master-ipv4-cidr and VPC peering; newer clusters use Private Service Connect for the control-plane connection and treat node privacy and endpoint access as separate settings. Use the flags documented for your GKE version.

The control plane is the head office. The DNS-based endpoint is a front desk that checks your staff card wherever you come from; the internal IP endpoint is a side door you can only reach from inside the campus. Private nodes are workers with no street door at all.

The cluster

A production Standard cluster, on the Shared VPC from lesson 03:

gcloud container clusters create prod \
  --project my-prod-project --region europe-west1 \
  --release-channel regular \
  --network projects/net-host-prod/global/networks/prod-vpc \
  --subnetwork projects/net-host-prod/regions/europe-west1/subnetworks/gke-prod-ew1 \
  --cluster-secondary-range-name pods --services-secondary-range-name services \
  --enable-private-nodes --enable-dns-access \
  --enable-dataplane-v2 \
  --workload-pool my-prod-project.svc.id.goog \
  --service-account gke-nodes@my-prod-project.iam.gserviceaccount.com \
  --enable-shielded-nodes \
  --num-nodes 1

In practice you'll write this in Terraform (lesson 14), create the cluster with a small default pool, and then replace it with your own node pools.

Node pools

Pool Machine type (example) Purpose
system e2/n2-standard-4 Cluster add-ons, ingress gateways, monitoring agents (tainted for them if you like)
apps n2/n4-standard-8 General workloads, autoscaled
spot e2-standard-4, Spot Batch, CI runners, stateless workers that tolerate interruption
gpu g2 / a2 with GPUs ML inference or training, tainted

Per pool, decide:

  • Machine family and size: bigger nodes pack better and waste less on per-node overhead; smaller nodes limit the blast radius of one node. 8–16 vCPUs is a common middle.
  • Image: Container-Optimized OS with containerd (default, minimal, auto-updated) or Ubuntu (when you need kernel modules or tooling COS lacks).
  • Disk: type (pd-balanced is a good default) and size (image cache, emptyDir, logs).
  • Labels and taints: how workloads find (or avoid) the pool.
  • Autoscaling min/max (per zone in regional clusters) and the upgrade strategy (lesson 13).
  • Location: which zones; spread across all zones of the region for production.

The node service account

Create a dedicated, minimal service account for nodes instead of the Compute Engine default (which commonly has the broad Editor role):

gcloud iam service-accounts create gke-nodes --display-name "GKE nodes"
gcloud projects add-iam-policy-binding my-prod-project \
  --member serviceAccount:gke-nodes@my-prod-project.iam.gserviceaccount.com \
  --role roles/container.defaultNodeServiceAccount

That role bundles what nodes need (writing logs and metrics, reading metadata); add Artifact Registry Reader on the repositories nodes pull from. On older setups you'll see the individual logging, monitoring and storage roles instead. Workloads never use this account: they get their own identities (lesson 06).

Spot VMs

Spot VMs cost a fraction of the on-demand price but can be reclaimed with about 30 seconds' notice. Use them for workloads that tolerate interruption, spread across zones and machine types, and always keep an on-demand pool for critical services:

spec:
  nodeSelector:
    cloud.google.com/gke-spot: "true"
  tolerations:
    - key: cloud.google.com/gke-spot
      operator: Equal
      value: "true"
      effect: NoSchedule
  terminationGracePeriodSeconds: 25

Compute classes

A compute class names a set of node preferences: machine families in priority order, Spot first with on-demand fallback, minimum sizes. Workloads select it with a node selector (cloud.google.com/compute-class: <name>), and GKE (with node auto-provisioning or Autopilot) creates nodes that match. They replace a lot of hand-built node pools and affinity rules. Autopilot ships built-in classes (for example Balanced, Scale-Out, and GPU classes); custom compute classes are a CRD in recent GKE versions. Check the version and mode you run before relying on a specific field.

Try it: a private cluster you can still reach

  1. Create a Standard cluster with private nodes, the DNS-based endpoint enabled and the external IP endpoint disabled (flags per your GKE version).
  2. Run gcloud container clusters get-credentials --dns-endpoint from your laptop, then kubectl get nodes. No VPN needed: IAM decides.
  3. Remove your IAM role on the project and try again; read the error.
  4. Add an apps pool and a tainted spot pool; deploy a Job with the Spot toleration and node selector; delete the default pool.
  5. Check a node's service account with gcloud compute instances describe <node> --format='value(serviceAccounts[].email)'.

Going deeper: fewer, bigger, simpler

  • Prefer few node pools with clear purposes over one per team; use namespaces, quotas and compute classes for team separation.
  • Reserve capacity for system components (or a system pool) so a noisy app can't starve DNS or the ingress.
  • Use sole-tenant nodes or separate clusters only for hard isolation requirements; they cost more and complicate operations.

Recap

  • Nodes are private; reach the control plane through the DNS-based endpoint (IAM) or internal endpoints, and disable broad external access.
  • Build clusters on the Shared VPC subnet and its secondary ranges, with Dataplane V2, Workload Identity and Shielded nodes from day one.
  • Design node pools by purpose: machine type, image, labels, taints, autoscaling.
  • Give nodes a minimal service account, never the default one.
  • Use Spot for interruption-tolerant work and compute classes to express node preferences.

This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.