Production GKE Platform — From Zero to Production›15 · Disaster recovery & reliability

Lesson 15 of 18 · Part 3 — Operate

Disaster recovery & reliability

Design GKE for failures you can actually recover from: what a regional cluster survives, rebuilding clusters from Terraform and GitOps, restoring state with Backup for GKE and managed databases, failing over between regions with multi-cluster Gateways, and proving it with game days.

Advanced
Key wordsGKE disaster recoveryRTORPOregional clusterzone failureregion failureBackup for GKEcross-region restoremulti-cluster GatewayGitOps rebuildPodDisruptionBudgetgame daymanaged databases

Failures and what covers them

Failure Covered by
Pod or node dies Replicas, PDBs, auto-repair, the scheduler
Zone outage Regional cluster, node pools in all zones, zone spread, regional disks or app replication
Bad deployment or config GitOps revert, progressive rollouts, Backup for GKE for deleted objects
Cluster lost or broken Rebuild from Terraform + GitOps, restore state from Backup for GKE
Region outage A second cluster in another region, data replicated there, traffic moved by multi-cluster Gateway or DNS
Data corruption or deletion Point-in-time backups of databases, versioned buckets, retained snapshots

Set RTO (how long you may be down) and RPO (how much data you may lose) per service tier first; they decide which rows you need.

Disaster recovery is like a school fire drill. Having exits (regions), a plan on the wall (runbooks) and a list of children (backups) matters, but only the drill shows whether everyone actually gets out in time.

Zone failures: regional by default

  • Regional clusters, node pools spread across all three zones, topologySpreadConstraints on zones for every important Deployment.
  • Enough spare capacity (or fast autoscaling) to run on two zones.
  • Stateful apps: application replication across zones, or regional Persistent Disks (lesson 09).
  • Managed databases with high availability (Cloud SQL HA, AlloyDB, Memorystore with replicas).

Cluster rebuild: cattle, not pets

If the cluster disappears tomorrow, recovery should be:

  1. terraform apply for the cluster and node pools (network and IAM already exist).
  2. Bootstrap GitOps (one root Application or Config Sync root), which installs add-ons, policies, namespaces and apps.
  3. Restore what isn't in Git: Backup for GKE for PVC data and objects created at runtime, database restores where needed.
  4. Point traffic at it (Gateway, DNS).

Measure how long each step takes in a rehearsal. Things that commonly slow it down: quota in the target region, IP ranges not prepared, secrets not reachable, image pulls from another region, DNS TTLs.

Backup for GKE

  • One backup plan per cluster (or per critical namespace), stored in another region than the cluster.
  • Include volume data and Secrets where needed, encrypt with your own key, keep enough retention to recover from slow-to-notice mistakes.
  • Restore plans for: the same cluster (deleted namespace), a new cluster in the same region (cluster loss), a cluster in another region (region loss).
  • Remember what it doesn't replace: managed databases have their own backups and replicas; apps in Git come back through GitOps.

Region failures: two regions

Pattern RTO Cost Notes
Backup and restore (cold) Hours Lowest Rebuild from code in region B, restore data from cross-region backups
Warm standby Minutes to an hour Medium Small cluster always running in region B, data replicated, scale up on failover
Active-active Near zero Highest Both regions serve traffic; multi-cluster Gateway routes and fails over automatically; data layer must support multi-region writes or a clear primary

A multi-cluster Gateway (global external class, lesson 07) puts one global address in front of Services in both regions and stops sending traffic to an unhealthy one. DNS failover with health checks is the simpler alternative for warm standby.

Proving it

  • Game days each quarter: drain a zone, delete a namespace and restore it, rebuild a cluster from code, fail over to the second region in staging.
  • Record the measured RTO and RPO and compare with the targets.
  • Check backups automatically: alert when a backup plan's last successful backup is older than expected, and restore a sample into a scratch cluster on a schedule.
  • Keep the runbook next to the code, with the exact commands, and update it after every rehearsal.

Try it: lose a namespace, then a cluster

  1. Enable Backup for GKE on a lab cluster and create a plan for namespace shop (with volume data), stored in another region.
  2. Deploy an app with a PVC, write data, take a backup.
  3. Delete the namespace and restore it with a restore plan; verify the data.
  4. Delete the cluster. Recreate it with your Terraform from lesson 14, bootstrap GitOps, and restore the backup into the new cluster.
  5. Write down how long each step took. That's your measured RTO for this setup.

Going deeper: the hard part is data

  • Prefer managed, replicated data services for critical state; the cluster is easier to recreate than a database.
  • Keep configuration, secrets references and IAM for the DR region in code, applied in advance, so failover doesn't wait on permissions.
  • Watch quotas and capacity in the DR region (CPUs, GPUs, IPs) and request them before you need them.

Recap

  • Match RTO/RPO per tier to the failures you must survive.
  • Regional clusters and zone spread cover zones; Terraform + GitOps rebuild clusters.
  • Backup for GKE restores objects and volumes, including into another region.
  • Second region: cold, warm or active-active, with multi-cluster Gateway or DNS failover.
  • Game days turn the plan into a measured, trusted procedure.

This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.