Production GKE Platform — From Zero to Production›18 · Q&A: operations, storage, security & cost

Lesson 18 of 18 · Part 5 — Design & operations Q&A

Q&A: operations, storage, security & cost

Common questions about running GKE in production, with detailed answers: upgrades, scaling, storage and DR, observability, security, cost, and platform tooling.

Architect
Key wordsQ&AGKE operationsupgradesmaintenance windowsautoscalingstorageStatefulSetsBackup for GKEdisaster recoveryobservabilitylogging costsecurityBinary Authorizationcost optimisationfleetsGo controllers

How to use these questions

Same approach as the previous lesson: clarify the context, compare options, decide, name the failure modes and say how you'd verify. Open each answer after you've thought it through.

Running a platform is like running a kitchen that never closes: the questions are less about the recipe and more about how you keep cooking while replacing the oven, how you notice a burnt dish before a customer does, and how you keep the gas bill sane.

Upgrades and change

Q1. How should a production GKE fleet be upgraded with minimal risk?

  • Channels: dev on a faster channel (or manual early upgrades), staging and prod on Regular or Stable.
  • Timing: maintenance windows in low-traffic hours on working days; exclusions around business peaks; Pub/Sub upgrade notifications to the platform team.
  • Preparation: deprecation insights clean, third-party controllers compatible, PDBs on critical workloads, IP headroom for surge nodes.
  • Node strategy: surge with maxUnavailable: 0 for most pools; blue-green with soak for sensitive pools.
  • Sequencing: dev → staging → prod with soak time between; fleet rollout sequencing where available.
  • Verification: synthetic checks and SLO dashboards during the window; stop and investigate if error rates rise.

Measure version lag per cluster and time to patch after security bulletins.

Q2. A node upgrade is stuck for an hour. What are the likely causes?

  • A PodDisruptionBudget that can't be satisfied (minAvailable equal to replicas, or Pods not Ready), so drains wait. GKE respects PDBs for a limited time before forcing.
  • No capacity for surge nodes: quota, IP range exhaustion, or a machine type unavailable in the zone.
  • Pods with long terminationGracePeriodSeconds or stuck finalizers.
  • Webhooks failing (admission webhooks that block Pod creation on new nodes).

Check the cluster operation status, events in the affected namespaces, kubectl get pdb -A, and autoscaler and quota errors. Fix the PDB or capacity issue; avoid disabling PDBs globally.

Q3. How does a team handle a Kubernetes API removal in the next minor version?

  1. Find callers: GKE deprecation insights and recommendations list deprecated API usage with user agents; also scan Git (manifests, Helm charts) with tools such as pluto or kubent.
  2. Update manifests and charts to the new API versions; upgrade controllers that use old APIs.
  3. Re-check insights after the changes (they reflect recent usage, so allow time).
  4. Only then let the minor upgrade proceed (GKE pauses automatic minor upgrades while it detects removed-API usage).

Keep API versions current as routine work rather than a pre-upgrade scramble.

Scaling and capacity

Q4. Traffic spikes ten times within minutes. How should the platform keep up?

  • Scale Pods on a leading signal (requests per second, queue depth) rather than CPU alone, with HPA behaviour tuned for fast scale-up.
  • Keep headroom: an overprovisioning Deployment of low-priority placeholder Pods, so real Pods schedule instantly while nodes are added.
  • Make starts fast: small images, image streaming, quick readiness.
  • Ensure node scale-up can happen: autoscaler max values, compute classes with fallbacks, regional CPU quota, IP space.
  • Protect the system: rate limiting at the edge (Cloud Armor), load shedding in apps, and capacity for system components.

Rehearse with load tests that ramp as fast as real traffic does, and measure time from spike to healthy latency.

Q5. How is GKE cost brought under control without hurting reliability?

  • Visibility first: GKE cost allocation by namespace and label; idle and over-provisioned workload reports.
  • Rightsize requests with VPA recommendations (this is the main lever in Autopilot).
  • Standard: better bin-packing, the optimize-utilization profile, fewer and larger nodes, delete idle pools.
  • Spot for interruption-tolerant work; committed use discounts for the steady baseline.
  • Networking: watch cross-zone and cross-region traffic, Cloud NAT data processing, load balancer counts.
  • Observability: log exclusions and retention, metric cardinality.
  • Non-prod: scale dev clusters down at night, zonal clusters for throwaway environments.

Track cost per team and per request served, so savings don't come at the expense of SLOs.

Storage and recovery

Q6. Should a database run on GKE or on a managed service?

Prefer a managed service (Cloud SQL, AlloyDB, Spanner, Memorystore) for important data when it meets the requirements: backups, HA, patching, point-in-time recovery and replicas are handled, and clusters stay easy to rebuild.

Run it on GKE when a managed service can't meet a requirement (an engine or extension not offered, a specific version, portability across environments including on-prem), and then do it properly: a mature operator, StatefulSets with zone spread, application-level replication or regional disks, backups to Cloud Storage, tested restores, and capacity for maintenance.

Q7. What does a complete DR plan for a GKE platform include?

  • Targets: RTO and RPO per service tier.
  • Zone failure: regional clusters, zone spread, regional disks or replicated data.
  • Cluster loss: Terraform + GitOps rebuild, Backup for GKE for runtime objects and volumes, tested.
  • Region loss: a pattern per tier (cold, warm, active-active), data replicated to the DR region, quotas and IAM prepared there, traffic moved by multi-cluster Gateway or DNS.
  • Human error: GitOps reverts, deletion protection, backups with enough retention, database point-in-time recovery.
  • Proof: game days with measured RTO/RPO, automated backup checks, runbooks with exact commands.

Observability and incidents

Q8. What should a GKE platform team alert on?

Alert on what users feel and on platform conditions that will soon affect them:

  • Service symptoms: SLO burn rate (error rate, latency) per critical service.
  • Capacity: Pods Pending beyond a few minutes, autoscaler failures, quota and IP range headroom.
  • Health: nodes NotReady, crash-looping system components, DNS errors.
  • Edge: load balancer 5xx, certificate expiry, Cloud Armor blocks spiking.
  • Egress: Cloud NAT port exhaustion and dropped packets.
  • Data: PVC usage, backup age.
  • Change: upgrade and security bulletin notifications.

Everything else goes to dashboards, not pagers. Each alert has an owner and a runbook link.

Q9. Apps intermittently time out calling an external API. Where should the investigation start?

Common causes, from the platform side:

  • Cloud NAT port exhaustion: many connections to one destination use up a NAT IP's ports. Check NAT dropped-packet metrics; add NAT IPs, enable dynamic port allocation, or reuse connections.
  • DNS: kube-dns or Cloud DNS throttling or ndots-driven query storms; check DNS error rates and latency.
  • Network policy or firewall changes blocking some paths (Dataplane V2 policy logging helps).
  • Node pressure: CPU throttling from tight limits.

Correlate timeouts with these metrics in time, then confirm with a test Pod on an affected node.

Security and platform tooling

Q10. How can a platform guarantee that only reviewed, scanned images run in production?

  1. Images are built only by CI, from reviewed commits, into a restricted Artifact Registry repository (no human write access).
  2. CI runs tests and vulnerability scans, then signs the image or creates an attestation.
  3. Binary Authorization (or a policy engine verifying signatures) admits only attested images in production clusters, with logged break-glass.
  4. Manifests pin images by digest.
  5. Continuous scanning of running images (security posture dashboard) catches new vulnerabilities; a process rebuilds and redeploys.

Q11. When is a custom Kubernetes controller written in Go the right tool for a platform?

When the platform needs continuous reconciliation of something Kubernetes doesn't model yet, for example creating a namespace, quotas, RBAC, a Gateway route and cloud resources from one Team object, or keeping external systems in step with cluster state.

Good practice:

  • Use controller-runtime / Kubebuilder; one CRD with a clear spec and status; idempotent reconcile loops; owner references and finalizers for clean-up.
  • Prefer existing tools first (Config Connector for Google Cloud resources, GitOps with templates, policy engines) and write a controller only when they don't fit.
  • Test with envtest, ship with metrics and health endpoints, run with Workload Identity and least-privilege RBAC, and version the CRD carefully.

Q12. What makes a GKE platform "self-service" for application teams?

  • A golden path: a template repository (app, Dockerfile, CI workflow, Kubernetes manifests) that works on day one.
  • Pull-request onboarding: a team adds a file describing its namespace, quotas, groups and routes; review and GitOps do the rest.
  • Guard-rails instead of gates: policies enforce standards automatically; humans review exceptions.
  • Visibility: dashboards, logs, costs and SLOs per team without asking the platform team.
  • Documentation and support: runbooks, office hours, and a feedback loop that shapes the platform roadmap.

Measure it: time from "new service" to first production deploy, number of platform tickets per team, and adoption of the golden path.

Recap

  • Upgrades: channels, windows, exclusions, deprecation checks, surge or blue-green, dev first.
  • Scaling and cost: leading metrics, headroom, fast starts; rightsizing, Spot, commitments, cost per team.
  • Data and DR: managed databases where possible, tested restores, a pattern per tier for region loss.
  • Observability: alert on symptoms and capacity, not everything.
  • Security and tooling: signed images and admission checks; Go controllers only when existing tools don't fit; self-service through Git.

This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.