Learning Hub / Cloud — OpenStack, AWS, EKS & GKE / Production GKE Platform — From Zero to Production
Lesson 16 of 18 · Part 4 — Architect
Architecture & design review
Tie the track together: a reference GKE platform architecture, the key decisions with their trade-offs, and a review checklist to run before a cluster carries production traffic.
Architect
Key wordsGKE reference architecturedesign reviewarchitecture decisionstrade-offsreview checklistmulti-tenancyplatform engineeringGoogle Cloud landing zone
A reference architecture
Organization
├── platform folder
│ ├── net-host-prod / net-host-nonprod Shared VPCs, Cloud NAT, firewall policies, Interconnect
│ ├── platform-shared Artifact Registry, CI identities, Certificate Manager, DNS
│ └── platform-fleet fleet host project, monitoring metrics scope
├── prod folder
│ └── gke-prod regional clusters per region (europe-west1, europe-west4)
└── nonprod folder
└── gke-dev, gke-staging regional (or zonal for dev) clusters
Each cluster
private nodes · DNS-based control-plane endpoint · Dataplane V2 · Workload Identity
release channel Regular/Stable · maintenance windows + exclusions · secrets encrypted with KMS
node pools: system, apps (autoscaled), spot (tainted) — or Autopilot
Gateway API for ingress · Cloud Armor · Certificate Manager
Managed Prometheus · Cloud Logging with exclusions · SLO alerts
Backup for GKE (cross-region) · GitOps bootstrap · Policy Controller / Kyverno
Teams
namespaces + quotas + RBAC (Google Groups) + network policies + Pod Security restricted
apps delivered by GitOps; images built and signed in CI with Workload Identity Federation
An architecture review is like checking a house plan before building: does every room have a door, will the wiring take the load, can people get out in a fire? It's far cheaper to move a wall on paper than after the concrete is poured.
Key decisions and trade-offs
| Decision | Default | When to choose differently |
|---|---|---|
| Autopilot or Standard | Autopilot for app clusters | Standard for GPUs, privileged agents, kernel tuning, tight bin-packing |
| Cluster count | Few multi-tenant clusters per env and region | Hard isolation (compliance, noisy GPU work, customer-dedicated) |
| Regional or zonal | Regional for prod and staging | Zonal for throwaway dev to save cost |
| Release channel | Regular (dev ahead of prod) | Stable for change-averse prod; Extended when you can't upgrade often |
| Network | Shared VPC, one subnet per cluster, planned Pod ranges | Separate VPCs for strong isolation; Class E for crowded IP space |
| Control-plane access | DNS-based endpoint with IAM, no external IP endpoint | Internal endpoint only, for strict network-perimeter rules |
| Ingress | Gateway API with a shared, platform-owned Gateway | Service LoadBalancer for L4; Ingress for existing setups |
| Identity | Workload Identity Federation, groups for RBAC, no keys | — |
| GitOps | Argo CD (multi-cluster UI) or Config Sync (fleet-native) | Choose one and standardise |
| Data | Managed databases; StatefulSets only when justified | Self-run when the managed service can't meet a requirement |
| DR | Cross-region Backup for GKE + rebuild from code | Warm standby or active-active for low-RTO services |
The review checklist
Requirements
- [ ] Availability target, RTO and RPO per service tier written down
- [ ] Compliance scope (data residency, encryption keys, audit) known
- [ ] Expected scale: nodes, Pods, Services, requests, growth
Network
- [ ] IP plan for primary, Pod, Service and proxy-only ranges, with growth and surge headroom (lesson 04)
- [ ] Private nodes, Cloud NAT sized and monitored, Private Google Access on
- [ ] Firewall policies and network policies (default-deny per namespace)
Identity and security
- [ ] No service-account keys; Workload Identity for pods; WIF for CI
- [ ] Minimal node service account; IAM via groups; RBAC per namespace
- [ ] Secrets encrypted with KMS; Secret Manager for app secrets
- [ ] Pod Security restricted (exceptions documented); policy engine rules; image signing/Binary Authorization where required
Operations
- [ ] Release channel, maintenance windows, exclusions for business peaks, upgrade notifications
- [ ] Node pool upgrade strategy chosen; PDBs on critical workloads
- [ ] Autoscaling limits set; quotas checked for peak and for the DR region
- [ ] Observability: managed Prometheus, alerts on symptoms and SLOs, log cost controls
- [ ] Backup for GKE plans with tested restores; game-day schedule
Delivery
- [ ] Everything in Terraform and Git; no manual changes in production
- [ ] GitOps bootstrap can rebuild a cluster; tested
- [ ] Self-service path for teams (namespace, quota, RBAC, route) through pull requests
Try it: review your own design
- Write a one-page design for a GKE platform for a fictional company (3 environments, 2 regions, 20 teams) using the reference above.
- Fill in the decisions table with your choice and one sentence of reasoning per row.
- Run the checklist against it and list the gaps.
- Swap documents with someone (or come back a week later) and challenge each decision with "what requirement does this serve?".
- Turn the top three gaps into tasks with owners.
Going deeper: decision records
- Record each major decision as an ADR (context, options, decision, consequences) in the platform repository; future reviews then argue with reasons, not memories.
- Revisit decisions on a schedule: Google Cloud features change quickly (Autopilot capabilities, fleet packaging, Gateway features), and yesterday's constraint may be gone.
Recap
- A reference GKE platform: organisation and Shared VPC, few multi-tenant regional clusters, private by default, Workload Identity, Gateway API, GitOps, managed observability and Backup for GKE.
- Each major decision has a default and known reasons to differ.
- Run the checklist before production, and keep decisions as ADRs.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.