Lesson 05 of 18 · Part 2 — Build the platform
The cluster: control plane access, node pools & compute
Create a production GKE cluster: private nodes, how people and pipelines reach the control plane, node pools with the right machine types, images and labels, a minimal node service account, Spot VMs, and compute classes that tell workloads where to run.
The control plane: how it's reached
| Endpoint | Access controlled by | Typical use |
|---|---|---|
| DNS-based | IAM (the caller needs permission on the cluster), from anywhere that can reach Google APIs | People and CI without VPN or bastions |
| IP-based, internal | Network reachability inside the VPC (and hybrid links), authorised networks | Access from on-prem or a jump host |
| IP-based, external | Authorised networks (CIDR allowlist) | Legacy setups; disable it in production if you can |
All of them still require authentication (IAM) and authorisation (IAM + RBAC). Older guides talk about "private clusters" with a master-ipv4-cidr and VPC peering; newer clusters use Private Service Connect for the control-plane connection and treat node privacy and endpoint access as separate settings. Use the flags documented for your GKE version.
The control plane is the head office. The DNS-based endpoint is a front desk that checks your staff card wherever you come from; the internal IP endpoint is a side door you can only reach from inside the campus. Private nodes are workers with no street door at all.
The cluster
A production Standard cluster, on the Shared VPC from lesson 03:
gcloud container clusters create prod \
--project my-prod-project --region europe-west1 \
--release-channel regular \
--network projects/net-host-prod/global/networks/prod-vpc \
--subnetwork projects/net-host-prod/regions/europe-west1/subnetworks/gke-prod-ew1 \
--cluster-secondary-range-name pods --services-secondary-range-name services \
--enable-private-nodes --enable-dns-access \
--enable-dataplane-v2 \
--workload-pool my-prod-project.svc.id.goog \
--service-account gke-nodes@my-prod-project.iam.gserviceaccount.com \
--enable-shielded-nodes \
--num-nodes 1
In practice you'll write this in Terraform (lesson 14), create the cluster with a small default pool, and then replace it with your own node pools.
Node pools
| Pool | Machine type (example) | Purpose |
|---|---|---|
system |
e2/n2-standard-4 | Cluster add-ons, ingress gateways, monitoring agents (tainted for them if you like) |
apps |
n2/n4-standard-8 | General workloads, autoscaled |
spot |
e2-standard-4, Spot | Batch, CI runners, stateless workers that tolerate interruption |
gpu |
g2 / a2 with GPUs | ML inference or training, tainted |
Per pool, decide:
- Machine family and size: bigger nodes pack better and waste less on per-node overhead; smaller nodes limit the blast radius of one node. 8–16 vCPUs is a common middle.
- Image: Container-Optimized OS with containerd (default, minimal, auto-updated) or Ubuntu (when you need kernel modules or tooling COS lacks).
- Disk: type (pd-balanced is a good default) and size (image cache, emptyDir, logs).
- Labels and taints: how workloads find (or avoid) the pool.
- Autoscaling min/max (per zone in regional clusters) and the upgrade strategy (lesson 13).
- Location: which zones; spread across all zones of the region for production.
The node service account
Create a dedicated, minimal service account for nodes instead of the Compute Engine default (which commonly has the broad Editor role):
gcloud iam service-accounts create gke-nodes --display-name "GKE nodes"
gcloud projects add-iam-policy-binding my-prod-project \
--member serviceAccount:gke-nodes@my-prod-project.iam.gserviceaccount.com \
--role roles/container.defaultNodeServiceAccount
That role bundles what nodes need (writing logs and metrics, reading metadata); add Artifact Registry Reader on the repositories nodes pull from. On older setups you'll see the individual logging, monitoring and storage roles instead. Workloads never use this account: they get their own identities (lesson 06).
Spot VMs
Spot VMs cost a fraction of the on-demand price but can be reclaimed with about 30 seconds' notice. Use them for workloads that tolerate interruption, spread across zones and machine types, and always keep an on-demand pool for critical services:
spec:
nodeSelector:
cloud.google.com/gke-spot: "true"
tolerations:
- key: cloud.google.com/gke-spot
operator: Equal
value: "true"
effect: NoSchedule
terminationGracePeriodSeconds: 25
Compute classes
A compute class names a set of node preferences: machine families in priority order, Spot first with on-demand fallback, minimum sizes. Workloads select it with a node selector (cloud.google.com/compute-class: <name>), and GKE (with node auto-provisioning or Autopilot) creates nodes that match. They replace a lot of hand-built node pools and affinity rules. Autopilot ships built-in classes (for example Balanced, Scale-Out, and GPU classes); custom compute classes are a CRD in recent GKE versions. Check the version and mode you run before relying on a specific field.
Try it: a private cluster you can still reach
- Create a Standard cluster with private nodes, the DNS-based endpoint enabled and the external IP endpoint disabled (flags per your GKE version).
- Run
gcloud container clusters get-credentials --dns-endpointfrom your laptop, thenkubectl get nodes. No VPN needed: IAM decides. - Remove your IAM role on the project and try again; read the error.
- Add an
appspool and a taintedspotpool; deploy a Job with the Spot toleration and node selector; delete the default pool. - Check a node's service account with
gcloud compute instances describe <node> --format='value(serviceAccounts[].email)'.
Going deeper: fewer, bigger, simpler
- Prefer few node pools with clear purposes over one per team; use namespaces, quotas and compute classes for team separation.
- Reserve capacity for system components (or a
systempool) so a noisy app can't starve DNS or the ingress. - Use sole-tenant nodes or separate clusters only for hard isolation requirements; they cost more and complicate operations.
Recap
- Nodes are private; reach the control plane through the DNS-based endpoint (IAM) or internal endpoints, and disable broad external access.
- Build clusters on the Shared VPC subnet and its secondary ranges, with Dataplane V2, Workload Identity and Shielded nodes from day one.
- Design node pools by purpose: machine type, image, labels, taints, autoscaling.
- Give nodes a minimal service account, never the default one.
- Use Spot for interruption-tolerant work and compute classes to express node preferences.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.