Production GKE Platform — From Zero to Production›16 · Architecture & design review

Lesson 16 of 18 · Part 4 — Architect

Architecture & design review

Tie the track together: a reference GKE platform architecture, the key decisions with their trade-offs, and a review checklist to run before a cluster carries production traffic.

Architect
Key wordsGKE reference architecturedesign reviewarchitecture decisionstrade-offsreview checklistmulti-tenancyplatform engineeringGoogle Cloud landing zone

A reference architecture

Organization
├── platform folder
│   ├── net-host-prod / net-host-nonprod    Shared VPCs, Cloud NAT, firewall policies, Interconnect
│   ├── platform-shared                     Artifact Registry, CI identities, Certificate Manager, DNS
│   └── platform-fleet                      fleet host project, monitoring metrics scope
├── prod folder
│   └── gke-prod                            regional clusters per region (europe-west1, europe-west4)
└── nonprod folder
    └── gke-dev, gke-staging                regional (or zonal for dev) clusters

Each cluster
  private nodes · DNS-based control-plane endpoint · Dataplane V2 · Workload Identity
  release channel Regular/Stable · maintenance windows + exclusions · secrets encrypted with KMS
  node pools: system, apps (autoscaled), spot (tainted)  — or Autopilot
  Gateway API for ingress · Cloud Armor · Certificate Manager
  Managed Prometheus · Cloud Logging with exclusions · SLO alerts
  Backup for GKE (cross-region) · GitOps bootstrap · Policy Controller / Kyverno

Teams
  namespaces + quotas + RBAC (Google Groups) + network policies + Pod Security restricted
  apps delivered by GitOps; images built and signed in CI with Workload Identity Federation

An architecture review is like checking a house plan before building: does every room have a door, will the wiring take the load, can people get out in a fire? It's far cheaper to move a wall on paper than after the concrete is poured.

Key decisions and trade-offs

Decision Default When to choose differently
Autopilot or Standard Autopilot for app clusters Standard for GPUs, privileged agents, kernel tuning, tight bin-packing
Cluster count Few multi-tenant clusters per env and region Hard isolation (compliance, noisy GPU work, customer-dedicated)
Regional or zonal Regional for prod and staging Zonal for throwaway dev to save cost
Release channel Regular (dev ahead of prod) Stable for change-averse prod; Extended when you can't upgrade often
Network Shared VPC, one subnet per cluster, planned Pod ranges Separate VPCs for strong isolation; Class E for crowded IP space
Control-plane access DNS-based endpoint with IAM, no external IP endpoint Internal endpoint only, for strict network-perimeter rules
Ingress Gateway API with a shared, platform-owned Gateway Service LoadBalancer for L4; Ingress for existing setups
Identity Workload Identity Federation, groups for RBAC, no keys —
GitOps Argo CD (multi-cluster UI) or Config Sync (fleet-native) Choose one and standardise
Data Managed databases; StatefulSets only when justified Self-run when the managed service can't meet a requirement
DR Cross-region Backup for GKE + rebuild from code Warm standby or active-active for low-RTO services

The review checklist

Requirements

  • [ ] Availability target, RTO and RPO per service tier written down
  • [ ] Compliance scope (data residency, encryption keys, audit) known
  • [ ] Expected scale: nodes, Pods, Services, requests, growth

Network

  • [ ] IP plan for primary, Pod, Service and proxy-only ranges, with growth and surge headroom (lesson 04)
  • [ ] Private nodes, Cloud NAT sized and monitored, Private Google Access on
  • [ ] Firewall policies and network policies (default-deny per namespace)

Identity and security

  • [ ] No service-account keys; Workload Identity for pods; WIF for CI
  • [ ] Minimal node service account; IAM via groups; RBAC per namespace
  • [ ] Secrets encrypted with KMS; Secret Manager for app secrets
  • [ ] Pod Security restricted (exceptions documented); policy engine rules; image signing/Binary Authorization where required

Operations

  • [ ] Release channel, maintenance windows, exclusions for business peaks, upgrade notifications
  • [ ] Node pool upgrade strategy chosen; PDBs on critical workloads
  • [ ] Autoscaling limits set; quotas checked for peak and for the DR region
  • [ ] Observability: managed Prometheus, alerts on symptoms and SLOs, log cost controls
  • [ ] Backup for GKE plans with tested restores; game-day schedule

Delivery

  • [ ] Everything in Terraform and Git; no manual changes in production
  • [ ] GitOps bootstrap can rebuild a cluster; tested
  • [ ] Self-service path for teams (namespace, quota, RBAC, route) through pull requests

Try it: review your own design

  1. Write a one-page design for a GKE platform for a fictional company (3 environments, 2 regions, 20 teams) using the reference above.
  2. Fill in the decisions table with your choice and one sentence of reasoning per row.
  3. Run the checklist against it and list the gaps.
  4. Swap documents with someone (or come back a week later) and challenge each decision with "what requirement does this serve?".
  5. Turn the top three gaps into tasks with owners.

Going deeper: decision records

  • Record each major decision as an ADR (context, options, decision, consequences) in the platform repository; future reviews then argue with reasons, not memories.
  • Revisit decisions on a schedule: Google Cloud features change quickly (Autopilot capabilities, fleet packaging, Gateway features), and yesterday's constraint may be gone.

Recap

  • A reference GKE platform: organisation and Shared VPC, few multi-tenant regional clusters, private by default, Workload Identity, Gateway API, GitOps, managed observability and Backup for GKE.
  • Each major decision has a default and known reasons to differ.
  • Run the checklist before production, and keep decisions as ADRs.

This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.