Lesson 17 of 18 · Part 5 — Design & operations Q&A
Q&A: platform design, networking & identity
Common questions about GKE platform design, networking and identity, with detailed answers: the approach, the trade-offs, the numbers that matter and how to verify the result.
How to use these questions
These questions come up again and again when designing and reviewing GKE platforms. Think each one through first, then open the answer. Good answers to design questions share a shape:
- Clarify requirements and constraints (scale, availability, compliance, team skills, budget).
- Options, at least two, with trade-offs.
- Decision and why, for this context.
- Failure modes: how they'd show up and how to recover.
- Evidence: numbers, tests, operating experience.
A good design answer is like a good travel plan: it asks where you're going and with whom, compares a couple of routes, picks one for good reasons, and says what to do if the train is cancelled.
Platform and project design
Q1. What does the top-level design of a GKE platform for 30 product teams look like?
Start with requirements: availability per tier, regulatory scope (data residency, customer-managed keys), expected scale, how independent teams need to be, and budget.
Then the layers:
- Resource hierarchy: folders for platform, prod and non-prod; a project per environment for clusters (
gke-prod,gke-staging,gke-dev), shared services in a platform project (Artifact Registry, CI identities, DNS, certificates), data in separate projects. - Network: a Shared VPC per environment tier, owned by the network team; one subnet per cluster with planned Pod and Service ranges; Cloud NAT; hybrid links in the host project.
- Clusters: a few regional, multi-tenant clusters per environment and region; Autopilot unless workloads need Standard; separate clusters only for hard isolation (regulated data, GPU-heavy work).
- Tenancy: namespaces per team with quotas, limit ranges, RBAC through Google Groups, default-deny network policies, Pod Security
restricted, and a policy engine for organisation rules. - Delivery: Terraform for cloud resources, GitOps for everything inside clusters, a self-service pull-request path for new namespaces and routes.
- Operations: release channels with windows and exclusions, managed Prometheus and SLO alerts, Backup for GKE across regions, game days.
Verify with a written decision record per major choice and a go-live checklist (lesson 16).
Q2. When should a platform use more clusters rather than more namespaces?
Namespaces share the control plane, the nodes (unless separated by pools), cluster-scoped objects (CRDs, admission webhooks) and the upgrade schedule. A separate cluster is justified when any of those can't be shared:
- Hard isolation requirements (regulated data, customer-dedicated environments, untrusted code that GKE Sandbox doesn't cover).
- Different lifecycles: a workload that can't take the shared upgrade cadence, or needs a different release channel.
- Conflicting cluster-scoped components: incompatible CRD versions or operators.
- Scale limits or blast radius: one cluster becoming too large to upgrade or reason about.
- Location: other regions, or on-prem and edge sites (fleet members).
Otherwise more clusters mean more upgrades, more monitoring, more cost (cluster fee, system overhead) and more places for drift. Fleets and GitOps reduce that cost but don't remove it.
Q3. How do Autopilot and Standard compare for a general application platform?
Autopilot: Google manages nodes, scaling, node upgrades and a hardened security baseline; billing is per Pod request. Less to operate, good defaults, predictable per-workload cost; restrictions on privileged Pods, host access and some DaemonSets.
Standard: full control of node pools (machine types, GPUs, kernel settings, privileged agents), billing per VM, so efficient bin-packing pays off, but the team owns pool design, scaling and upgrade settings.
A common answer: Autopilot for most application clusters, Standard for clusters with specialised needs. Check that third-party agents (security, observability, service mesh) are supported on Autopilot before committing. Compute classes in recent versions blur the line, so re-check the current capabilities.
Q4. What does moving a platform from EKS to GKE change, and what stays the same?
Stays the same: Kubernetes itself, manifests, Helm and Kustomize, GitOps tooling, most policies, observability with Prometheus and OpenTelemetry.
Changes:
- Identity: IAM roles and access entries become Google IAM plus RBAC; IRSA/Pod Identity become Workload Identity Federation; CI uses a Google workload identity pool instead of an AWS OIDC provider.
- Network: a global VPC with regional subnets; Pods use a secondary range sized up front, not node-subnet IPs; security groups become firewall rules and network policy; the AWS Load Balancer Controller becomes built-in Services, Ingress and Gateway.
- Upgrades: automatic within release channels; the work shifts from running upgrades to controlling windows, exclusions and deprecations.
- Registry and storage: ECR → Artifact Registry; EBS/EFS → Persistent Disk/Hyperdisk/Filestore with managed CSI drivers.
Plan the migration per workload: rebuild the platform as code in GKE, move apps through GitOps, migrate data with the data service's own replication, and shift traffic gradually with DNS.
Networking and IP planning
Q5. How should Pod and Service ranges be sized for a cluster expected to reach 400 nodes?
- Decide max Pods per node from real density, say 64 (most clusters run far fewer Pods per node than 110).
- Per-node block: twice 64, rounded to a power of two = /25 (128 addresses).
- Max nodes with headroom for surge upgrades and autoscaling: 400 × 1.2 ≈ 480.
- Pod range: 480 × 128 = 61,440 → a /16 (65,536).
- Service range: a /20 (4,096 Services) is plenty for most clusters.
- Primary range: 480 nodes + internal load balancers → a /23 (512) or /22 to be safe.
Record it in the organisation's IP plan, and avoid overlaps with on-prem and other clouds. If private space is scarce, use non-RFC 1918 or Class E ranges for Pods, after checking hybrid equipment.
Q6. New nodes fail to come up while existing nodes are half empty. What is likely happening, and how is it fixed?
The classic cause is Pod range exhaustion: each node takes a fixed Pod block (based on max Pods per node), so the range runs out of blocks long before it runs out of Pods. Symptoms: node creation errors in cluster operations, autoscaler events about IP space, Pending Pods.
Confirm: count nodes × block size against the Pod range (kubectl get nodes -o custom-columns=...podCIDR, gcloud container clusters describe).
Fix without rebuilding:
- Add a secondary range to the subnet and create new node pools using it (
--pod-ipv4-range), or add it cluster-wide on versions that support that. - Use a lower max Pods per node on new pools.
- Move workloads to the new pools and delete old ones if needed.
Prevent it: size ranges for peak plus surge (Q5), and alert on remaining capacity.
Q7. How should private clusters be reached by engineers and CI?
Separate node privacy from control-plane access:
- Nodes: always private (no external IPs), egress through Cloud NAT, Google APIs through Private Google Access.
- Control plane: the DNS-based endpoint lets anyone with the right IAM permission reach the API without VPNs or bastions, and every request is authenticated, authorised and logged. Disable the external IP-based endpoint, or restrict it with authorised networks if it must stay.
- CI outside Google Cloud: Workload Identity Federation for credentials, then the DNS endpoint; or no direct access at all when a GitOps agent in the cluster applies changes.
- On-prem tooling: the internal endpoint over Interconnect or VPN.
Verify by removing IAM permissions (access must stop) and by checking audit logs for every API call from CI.
Q8. When are VPC firewall rules enough, and when is Kubernetes network policy needed?
VPC firewall rules see nodes (by service account or tag) and subnets. They are right for: what may reach the nodes at all (load balancers, health checks, the control plane), egress to on-prem ranges, and environment separation.
They can't distinguish Pods of different teams on the same node pool. For that, use network policy (enforced by Dataplane V2): default-deny per namespace, then explicit allows per flow. Use FQDN policies (where available) for egress to external services.
Combine both: firewall policies at the organisation and VPC level for the coarse boundaries, network policies inside clusters for Pod-level isolation.
Identity and access
Q9. How should access be structured for platform engineers, app teams and CI?
- Platform engineers: a group with
roles/container.adminon the cluster projects, plus documented break-glass access. - App teams: a group with
container.clusterViewer(enough to get credentials) and RBAC RoleBindings toeditor a custom role in their namespaces; read-onlyviewcluster-wide if useful. - CI: Workload Identity Federation to a CI service account; ideally CI only pushes images and updates Git, and GitOps applies changes, so CI needs no cluster write access.
- Everyone: groups, not individuals; Google Groups for RBAC enabled, with team groups nested in
gke-security-groups@.
Verify with kubectl auth can-i --list --as <user> per role and periodic IAM recommender reviews.
Q10. A Pod is calling Google Cloud APIs with the node's service account. Why is that a problem, and how is it fixed?
Every Pod on that node shares the node's permissions, often far broader than any single app needs (the default Compute Engine account usually has Editor). A compromised Pod can then act on much of the project.
Fix:
- Give nodes a minimal service account (logging, monitoring, image pulls only).
- Enable Workload Identity Federation on the cluster and set node pools to the GKE metadata server, which hides the node's credentials from Pods.
- Give each workload its own Kubernetes ServiceAccount and grant IAM roles to that principal (or let it impersonate a dedicated Google service account).
- Remove any service-account keys stored in Secrets, and block key creation with an organization policy.
Verify from inside a Pod: the token's identity is the workload's, and the node identity isn't reachable.
Q11. How can a multi-tenant GKE cluster keep teams from affecting each other?
Layer the controls:
- Access: RBAC per namespace; no cluster-scoped write access for teams.
- Resources: ResourceQuota and LimitRange per namespace; PriorityClasses so platform components win under pressure.
- Network: default-deny network policies; explicit allows.
- Workload safety: Pod Security
restricted; policy engine rules (approved registries, required limits, no host access). - Identity: one workload identity per app; no shared service accounts.
- Scheduling: dedicated node pools or compute classes for noisy or sensitive workloads; GKE Sandbox for untrusted code.
- Cost visibility: GKE cost allocation by namespace and label.
Test it: from a Pod in team A, try to reach team B's Services, read its Secrets, and exhaust its quota; each attempt should fail.
Q12. How do on-prem clusters and GKE clusters fit into one platform?
Make the operating model the same, even when the infrastructure differs:
- Register all clusters to a fleet (Google Distributed Cloud on-prem, GKE in the cloud, attached clusters elsewhere) or manage them from one Argo CD hub.
- One Git source of truth for configuration and policies, applied by GitOps everywhere.
- Connect gateway (or an equivalent) for access through central identity, without per-site VPN credentials.
- Consistent observability: the same metrics and logs pipeline, one place to look.
- Plan IP ranges across all sites and clouds together, and decide which Services must be reachable across environments.
- Respect real differences: air-gapped sites need local registries and a way to deliver updates without internet; edge sites need tolerance for disconnection.
Recap
- Start every design answer from requirements, then options, decision, failure modes and evidence.
- Few multi-tenant regional clusters, Shared VPC, Autopilot by default, separate clusters for real isolation needs.
- Size Pod ranges from max Pods per node and peak nodes; fix exhaustion with additional ranges.
- Private nodes, DNS-based control-plane access, IAM plus RBAC by group, Workload Identity for every Pod.
- Hybrid platforms share one operating model: fleet or hub, Git, policies, access and observability.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.