Production GKE Platform — From Zero to Production
Learning Hub / Cloud — OpenStack, AWS, EKS & GKE

Production GKE Platform — From Zero to Production

Intermediate → Architect18 lessonsAvailable

An end-to-end Google Kubernetes Engine (GKE) course: from an empty Google Cloud organisation to a running, observable, secured and upgradeable platform. Every Google Cloud building block, why it exists and how the pieces fit together: projects and Shared VPC, VPC-native IP planning, Autopilot and Standard, Workload Identity, Gateway API, storage, release channels, Terraform, GitOps and fleets across cloud and on-prem, ending with an architecture review and design and operations questions with detailed answers.

You'll meetGKEGoogle Kubernetes EngineGoogle CloudGCPAutopilotStandardrelease channelsShared VPCVPC-nativealias IPDataplane V2Workload Identity FederationGateway APIArtifact RegistryPersistent DiskFilestoreBackup for GKEManaged Service for PrometheusfleetsConfig SyncTerraform
Start lesson 01 →

What you'll be able to do

  • Explain every Google Cloud service a GKE platform depends on, and what Google manages versus what you own
  • Design the organisation, projects, Shared VPC and IP plan for production clusters
  • Choose between Autopilot and Standard, regional and zonal, and the right release channel
  • Control access end to end: people via IAM and RBAC, pods via Workload Identity Federation, guard-rails via organisation policies
  • Expose, store, scale, observe and secure workloads with the GKE-native building blocks
  • Run upgrades, Terraform, GitOps and multi-cluster fleets across cloud and on-prem, and recover from failures with a tested plan

Before you start

Kubernetes Administration (Level 1–2). Knowing the Production EKS Platform track helps but isn't required: lesson 01 maps every EKS term to its GKE equivalent.

How it works

Each lesson: plain-language idea → how it really works → hands-on. Each section ends with a cheat sheet & self-check.

Curriculum

Lessons marked “Read” are ready; the rest are on the way.

Real-world scenarios

Work through each one: symptom → misleading signal → evidence → root cause → prevention.

New nodes won't join a half-empty cluster

The Pod range ran out of per-node blocks. Find it, and fix it with an additional Pod range.

A pod used the node's identity

A workload called Google APIs as the node service account. Close the gap with Workload Identity Federation and a minimal node account.

The weekend auto-upgrade

A control-plane upgrade landed at the wrong time. Put maintenance windows, exclusions and release channels to work.

The logging bill doubled

Verbose workloads and default log routing. Find the cost drivers and fix them with exclusions and sinks.

This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.