Production GKE Platform — From Zero to Production
An end-to-end Google Kubernetes Engine (GKE) course: from an empty Google Cloud organisation to a running, observable, secured and upgradeable platform. Every Google Cloud building block, why it exists and how the pieces fit together: projects and Shared VPC, VPC-native IP planning, Autopilot and Standard, Workload Identity, Gateway API, storage, release channels, Terraform, GitOps and fleets across cloud and on-prem, ending with an architecture review and design and operations questions with detailed answers.
What you'll be able to do
- Explain every Google Cloud service a GKE platform depends on, and what Google manages versus what you own
- Design the organisation, projects, Shared VPC and IP plan for production clusters
- Choose between Autopilot and Standard, regional and zonal, and the right release channel
- Control access end to end: people via IAM and RBAC, pods via Workload Identity Federation, guard-rails via organisation policies
- Expose, store, scale, observe and secure workloads with the GKE-native building blocks
- Run upgrades, Terraform, GitOps and multi-cluster fleets across cloud and on-prem, and recover from failures with a tested plan
Before you start
Kubernetes Administration (Level 1–2). Knowing the Production EKS Platform track helps but isn't required: lesson 01 maps every EKS term to its GKE equivalent.
How it works
Each lesson: plain-language idea → how it really works → hands-on. Each section ends with a cheat sheet & self-check.
Curriculum
Lessons marked “Read” are ready; the rest are on the way.
Part 1 — Foundations
- 01Google Cloud & GKE services and termsEvery service and term a GKE platform uses, and how each maps to its AWS and EKS equivalentRead →
- 02GKE architecture: Autopilot, Standard, regional and release channelsWhat Google runs, what you run, and the four decisions every cluster starts withRead →
- 📋Cheat sheet & self-checkEvery command from this section on one page, then 7 questions to check yourself.Open →
Part 2 — Build the platform
- 03Organisation, projects & Shared VPCFolders, projects per environment, Shared VPC, subnets, Cloud NAT, Private Google Access and firewall rulesRead →
- 04VPC-native networking & IP planningAlias IPs, Pod and Service ranges, max Pods per node, adding ranges later, Dataplane V2 and network policyRead →
- 05The cluster: control plane access, node pools & computePrivate nodes, control-plane endpoints, node pools, machine types and images, Spot VMs, compute classesRead →
- 06Identity & access: IAM, RBAC and Workload Identity FederationWho can reach the cluster, what pods can do in Google Cloud, and the guard-rails around bothRead →
- 07Load balancing: Services, Ingress & Gateway APIPassthrough load balancers, container-native load balancing, the GKE Gateway controller, certificates and Cloud ArmorRead →
- 08Artifact Registry & application deliveryPrivate registries, scanning, image streaming, CI without keys, deploying and rolling out an appRead →
- 09Storage: Persistent Disk, Hyperdisk, Filestore & Cloud StorageCSI drivers, StorageClasses, regional disks, snapshots, StatefulSets and Backup for GKERead →
- 10Autoscaling: Pods, nodes and compute classesHPA, VPA, cluster autoscaler, node auto-provisioning, Autopilot scaling, Spot and costRead →
- 📋Cheat sheet & self-checkEvery command from this section on one page, then 24 questions to check yourself.Open →
Part 3 — Operate
- 11Observability: Cloud Logging, Cloud Monitoring & Managed PrometheusMetrics, logs, traces and alerts; control-plane metrics; managed versus self-run; log costRead →
- 12Security: nodes, secrets, policy and supply chainShielded nodes, COS, KMS secret encryption, Pod Security, security posture, Binary Authorization, GKE SandboxRead →
- 13Upgrades: release channels, maintenance windows & node upgrade strategiesAuto-upgrades you control, exclusions, surge versus blue-green, deprecated APIsRead →
- 14Terraform, GitOps & fleets across cloud and on-premGKE as code, Argo CD or Config Sync, fleets, Connect gateway, multi-cluster Gateway and hybrid clustersRead →
- 15Disaster recovery & reliabilityRegional design, Backup for GKE, multi-region failover, and recovery you can proveRead →
- 📋Cheat sheet & self-checkEvery command from this section on one page, then 15 questions to check yourself.Open →
Part 4 — Architect
Part 5 — Design & operations Q&A
- 17Q&A: platform design, networking & identityCommon questions with detailed answers: projects, Shared VPC, IP ranges, cluster modes, access, Workload IdentityRead →
- 18Q&A: operations, storage, security & costCommon questions with detailed answers: upgrades, scaling, storage and DR, observability, security, cost, fleetsRead →
- 📋Cheat sheet & self-checkEvery command from this section on one page, then 4 questions to check yourself.Open →
Real-world scenarios
Work through each one: symptom → misleading signal → evidence → root cause → prevention.
The Pod range ran out of per-node blocks. Find it, and fix it with an additional Pod range.
A workload called Google APIs as the node service account. Close the gap with Workload Identity Federation and a minimal node account.
A control-plane upgrade landed at the wrong time. Put maintenance windows, exclusions and release channels to work.
Verbose workloads and default log routing. Find the cost drivers and fix them with exclusions and sinks.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.