Part 3 — Operate · wrap-up
Cheat sheet & self-check
15 questions across 5 lessons. Each answer links back to the lesson it came from.
Pick an answer to see if you got it, and why.
Q1. What does Managed Service for Prometheus replace?
Show answer
B. Apps still expose /metrics; collection rules are PodMonitoring resources; storage and global querying are managed.
From lesson 11 · Observability: Cloud Logging, Cloud Monitoring & Managed PrometheusQ2. The logging bill doubled. What is the usual first fix?
Show answer
B. Ingestion volume drives cost. Exclusions and sinks control what is stored where; app log levels control what is produced.
From lesson 11 · Observability: Cloud Logging, Cloud Monitoring & Managed PrometheusQ3. Why alert on SLO burn rate rather than on every error?
Show answer
B. Burn-rate alerts (fast and slow windows) catch real incidents early with far fewer false pages.
From lesson 11 · Observability: Cloud Logging, Cloud Monitoring & Managed PrometheusQ4. What does application-layer secrets encryption add?
Show answer
B. Envelope encryption with a customer-managed key gives you control and auditability of the key; anyone with RBAC read on Secrets can still read them.
From lesson 12 · Security: nodes, secrets, policy and supply chainQ5. What does Binary Authorization enforce?
Show answer
B. It's an admission check against attestations, closing the gap between 'we scan images' and 'only scanned images run'.
From lesson 12 · Security: nodes, secrets, policy and supply chainQ6. When is GKE Sandbox (gVisor) worth using?
Show answer
B. gVisor intercepts system calls in a user-space kernel. It reduces the impact of container escapes but adds overhead and has compatibility limits.
From lesson 12 · Security: nodes, secrets, policy and supply chainQ7. A peak sales period is coming. How do you stop GKE upgrading the cluster then, without leaving the release channel?
Show answer
B. Exclusions block automatic upgrades for a period (scopes differ in what they block and how long they may last); windows say when upgrades may run otherwise.
From lesson 13 · Upgrades: release channels, maintenance windows & node upgrade strategiesQ8. When is a blue-green node pool upgrade better than a surge upgrade?
Show answer
B. Blue-green needs capacity for a second set of nodes but allows rollback during the soak. Surge is cheaper and faster but has no built-in rollback.
From lesson 13 · Upgrades: release channels, maintenance windows & node upgrade strategiesQ9. What can stop GKE from automatically upgrading a cluster to a new minor version?
Show answer
B. GKE checks deprecated API usage and pauses automatic minor upgrades when workloads still call removed APIs; fix the callers using the deprecation insights.
From lesson 13 · Upgrades: release channels, maintenance windows & node upgrade strategiesQ10. Why remove the default node pool and manage node pools as separate Terraform resources?
Show answer
B. Many node pool settings are immutable; inline pools tie those changes to the cluster resource. Separate google_container_node_pool resources keep changes small and safe.
From lesson 14 · Terraform, GitOps & fleets across cloud and on-premQ11. What should Terraform manage, and what GitOps?
Show answer
B. Each tool reconciles what it's good at: Terraform has plans and state for cloud APIs; GitOps controllers continuously reconcile Kubernetes objects and fix drift.
From lesson 14 · Terraform, GitOps & fleets across cloud and on-premQ12. What is a fleet?
Show answer
B. Fleets are the unit for multi-cluster management in Google Cloud, including clusters that don't run on Google Cloud.
From lesson 14 · Terraform, GitOps & fleets across cloud and on-premQ13. What does a regional GKE cluster protect against?
Show answer
B. Regional clusters cover zone failures. Region failures, bad changes and data loss need backups, GitOps rebuilds and a second region.
From lesson 15 · Disaster recovery & reliabilityQ14. What makes rebuilding a cluster in another region fast?
Show answer
B. When the cluster is cattle, recovery is apply + sync + restore data; the slow part is usually data, so plan RPO there.
From lesson 15 · Disaster recovery & reliabilityQ15. Why run game days?
Show answer
B. Rehearsals measure the real RTO and RPO and turn runbooks into something people trust.
From lesson 15 · Disaster recovery & reliability