Lesson 17 of 18 · Part 5 — Design & operations Q&A
Q&A: platform design, networking & identity
Common questions about EKS platform design, networking, compute and identity, with detailed answers: the approach, the trade-offs, the numbers that matter and how to verify the result.
How to use these questions
These are the questions that come up again and again when designing and reviewing EKS platforms. Think each one through first, then open the answer. Good answers to design questions usually follow the same shape:
- Clarify requirements and constraints (scale, availability target, compliance, team skills, budget).
- Options (at least two), with trade-offs.
- Decision and why, for this context.
- Failure modes, how you'd detect them, and how you'd recover.
- Evidence: numbers, tests, operating experience.
A good design decision works like a good doctor's advice: asks questions before prescribing, explains the options and side effects, picks one for this patient, and says what to watch for afterwards.
Platform and account design
Q1. What does the top-level design of an EKS platform for 40 product teams look like?
Start with requirements: availability targets per tier, regulatory scope (PCI, data residency), expected scale (services, pods, requests), team autonomy, and budget.
Then the layers:
- Accounts: an Organization with separate prod and non-prod workload accounts (possibly per business unit), plus shared-services (ECR, CI, DNS, Transit Gateway) and security/log-archive accounts. SCPs as outer guard-rails.
- Clusters: one cluster per environment per region to start, split further only for hard isolation needs (a PCI cluster, a latency-critical cluster). 40 teams fit comfortably in a few multi-tenant clusters.
- Tenancy: a namespace (or namespaces) per team; access entries scoped to namespaces; ResourceQuota and LimitRange; Pod Security
restricted; default-deny network policies; Kyverno policies for registries, labels and limits. - Platform services owned by the platform team and delivered by GitOps: ingress (LB Controller), DNS, certificates, secrets sync, observability, Karpenter, policy engine.
- Paved road: a service template (Helm chart, pipeline, dashboards, alerts) so teams get production defaults without asking.
Close with how you'd measure success: lead time to first deploy for a new team, upgrade cadence achieved, incidents caused by the platform, cost per team.
Q2. Many small clusters or a few large ones: which fits, and why?
It's a trade-off between isolation and operating cost.
- More clusters: hard blast-radius and compliance boundaries, independent upgrade timing, simpler tenancy. Cost: every cluster multiplies upgrades, add-ons, monitoring, access management and control-plane fees; fleet tooling (GitOps ApplicationSets, Cluster API or Terraform per cluster) becomes mandatory.
- Fewer clusters: better bin-packing and lower cost, one place to operate. Cost: you must invest in multi-tenancy controls, and one bad upgrade affects everyone.
A sensible default: few clusters per environment and region, with strong tenancy controls, plus dedicated clusters only when a concrete driver exists (compliance scope, a workload with conflicting requirements such as GPUs or very different upgrade needs, or a blast-radius requirement from the business). Write the drivers into an ADR so the next split is a decision, not a drift.
Q3. What does AWS manage in EKS, and what remains your responsibility?
AWS runs the control plane (API servers, etcd, controllers) across AZs, patches it, scales it and backs up etcd. With Fargate and Auto Mode it also manages node lifecycle.
You own: the data plane (node AMIs and their rollout on managed node groups, capacity), add-on versions and configuration, Kubernetes version upgrades (AWS provides them, you trigger and validate), access to the API, workload security (pod security, network policy, secrets, images), IAM for pods, observability of your workloads, backups of your workloads and data, and cost. A strong answer adds that "managed control plane" still means you must stay on supported versions: standard support ends and extended support costs extra.
Q4. A business unit wants its own cluster "for autonomy". What's the right response?
Ask what autonomy they actually need: deploy independently, choose their own tools, or different compliance? Most of that is achievable in a shared cluster with namespaces, delegated access, quotas and a self-service pipeline. Offer that first, with clear SLOs from the platform team.
If the driver is real (compliance, conflicting requirements, very different change cadence), a dedicated cluster is fine, but built from the same Terraform modules and GitOps baseline, owned under the same upgrade and security policy. Autonomy over workloads, consistency over platform. Unmanaged snowflake clusters are how fleets become unpatchable.
Networking and IP planning
Q5. How should the VPC be sized for a cluster expected to run 3,000 pods?
With the VPC CNI each pod needs a VPC IP, plus headroom for rolling updates (surge), node churn, load balancers and endpoints. 3,000 pods → plan for roughly 2× (6,000+) addresses.
- Primary VPC
/16(non-overlapping with corporate and peered ranges), three AZs. - Private node subnets of
/19per AZ (~8,000 IPs each) or smaller node subnets plus a secondary CIDR (100.64.0.0/16, carrier-grade NAT range) dedicated to pods via custom networking, which keeps routable corporate space for nodes and load balancers. - Prefix delegation to raise pods-per-node beyond ENI limits and reduce IP allocation calls.
- Public
/22subnets per AZ for ALBs/NLBs and NAT.
Mention the alternative: IPv6 removes pod IP exhaustion but requires IPv6-capable dependencies; decided at cluster creation.
Q6. Pods are stuck Pending with IP-assignment errors, but nodes have spare CPU and memory. What's the cause and the fix?
Two possible limits:
- Per-node: each instance type has a maximum number of ENIs and IPs per ENI, so a node can run out of pod IPs before CPU. Check
max-podsand the CNI'sipamdlogs. Fix with prefix delegation (each ENI slot gets a /28), or larger instance types. - Per-subnet: the subnet itself has no free addresses (
AvailableIpAddressCountnear zero). Fix with a secondary CIDR and custom networking for pods, or larger/additional subnets.
Also check WARM_IP_TARGET / WARM_ENI_TARGET: aggressive warm pools hoard IPs on idle nodes. Prevent recurrence with a dashboard and alert on free IPs per subnet.
Q7. How do you build a fully private EKS cluster with no internet access?
- Endpoint: private only.
- No NAT or IGW routes from private subnets.
- VPC interface endpoints for everything the nodes and controllers call: ECR (api and dkr), S3 (gateway), STS, EC2, EKS, EKS Auth (Pod Identity), CloudWatch Logs, SSM, Elastic Load Balancing, plus any service apps use.
- Images only from ECR (use pull-through cache or replication to bring in upstream images).
- Admin and CI access over VPN, Direct Connect, or runners inside the VPC.
- Controllers configured with regional endpoints; Helm charts and OS packages mirrored internally.
Call out the operational cost: every new dependency needs an endpoint or a mirror, so document the allowed list and test cluster creation end to end in the private setup.
Q8. How do pods talk to an RDS database securely, and how do you limit which pods can?
Network: RDS in private subnets; its security group allows the database port only from a specific source. Options for that source: the node security group (coarse: every pod on every node), or security groups for pods (the VPC CNI attaches a security group to selected pods, so only those pods are allowed). Add Kubernetes network policies for east-west control inside the cluster.
Identity and credentials: prefer IAM database authentication with the pod's own role (Pod Identity), or credentials from Secrets Manager with rotation, read with the pod's role. Encrypt in transit (require TLS on the DB). Audit with CloudTrail and database logs.
Q9. NAT costs doubled. How do you find and fix the cause?
Measure first: VPC Flow Logs or NAT gateway metrics (BytesOutToDestination) show which sources and destinations dominate. Typical culprits: image pulls through NAT, S3 traffic without the gateway endpoint, cross-AZ NAT usage (single NAT gateway), chatty external APIs, log shipping to external services.
Fixes: S3 gateway endpoint (free), ECR/STS/CloudWatch interface endpoints where volume justifies the hourly cost, per-AZ NAT to avoid cross-AZ charges, image caching (smaller images, pull-through cache, fewer node churns), and sending logs through endpoints. Put a cost alarm on NAT data processing.
Cluster and compute
Q10. Managed node groups, Karpenter, Fargate or Auto Mode? How do you choose?
- Managed node groups: predictable, simple, one instance-type family per group; good for steady and system workloads.
- Karpenter: provisions right-sized instances per pending pod, diversifies across instance types and Spot, consolidates; best for mixed and spiky workloads, at the cost of running Karpenter and tuning disruption.
- Fargate: per-pod isolation, no node management, per-pod billing; no DaemonSets, privileged pods or GPUs; good for small, isolated or bursty jobs.
- Auto Mode: AWS runs the Karpenter-based compute plus storage and load-balancing components; minimal operations for an extra per-instance fee and less low-level control.
A common recommendation: a small managed node group for critical system components and Karpenter itself, Karpenter for workloads with On-Demand and Spot pools, Fargate for isolated batch where it fits; Auto Mode for teams without platform engineering capacity. Justify with workload profile, Spot appetite and team capacity.
Q11. How do you run Spot safely for production workloads?
- Only for interruption-tolerant workloads: stateless, multiple replicas, graceful shutdown under two minutes.
- Diversify widely across instance families, sizes and AZs (Karpenter does this when the NodePool allows many types), so one capacity pool running dry doesn't take out the service.
- Handle interruption notices (Karpenter's interruption queue) to drain early.
- PDBs, topology spread and enough replicas that losing a node or an AZ's Spot pool is survivable.
- A fallback to On-Demand (weighted NodePools) and a floor of On-Demand capacity for critical paths.
- Measure: interruption rate, rescheduling time, and savings achieved.
Q12. What would you pin down at cluster creation because it's hard to change later?
IP family (IPv4/IPv6), VPC and subnet design (and whether pods use a secondary CIDR), service CIDR, the cluster name (referenced by tags and roles), secrets encryption approach, authentication mode direction (API), endpoint access model, and the account/region placement. Everything else (versions, add-ons, node groups, access) should be changeable through code without replacement, and you should know which Terraform attributes force a new cluster.
Identity and access
Q13. How does a developer's kubectl command get authorised, end to end?
- The developer signs in through IAM Identity Center and gets temporary credentials for a permission-set role.
aws eks update-kubeconfigwrote a kubeconfig whose user runsaws eks get-token, which produces a signed STS token.- The EKS API server's authenticator validates the token and finds the IAM principal.
- The access entry for that role maps it to access policies (scoped to namespaces) and/or Kubernetes groups.
- Kubernetes RBAC (or the access policy's permission set) authorises the verb, resource and namespace.
- Admission (Pod Security, Kyverno) can still reject the request.
- The call is recorded in the audit log.
A strong answer mentions that no long-lived credentials exist anywhere in that chain.
Q14. Pod Identity or IRSA? When would you still use IRSA?
Pod Identity for new EKS workloads: the trust policy uses one fixed service principal (pods.eks.amazonaws.com), so roles are reusable across clusters without editing trust for each cluster's OIDC provider; associations live in the EKS API (easy to audit); and session tags enable ABAC.
IRSA still fits where Pod Identity doesn't: Fargate pods (at the time of writing), self-managed Kubernetes on EC2, some older SDKs, or existing estates where migration isn't worth it yet. Migration path: add Pod Identity associations alongside, verify with sts get-caller-identity in the pods, then remove annotations and old trust statements.
Q15. A pod accessed an S3 bucket it shouldn't have. How could that happen, and how do you prevent it?
Likely paths: the pod used the node role through the instance metadata service (IMDS reachable, node role too broad); a shared or over-broad pod role; a wildcard in a policy; credentials in a Secret or environment variable; or a bucket policy that grants too widely.
Investigate with CloudTrail (which principal made the call), then prevent: IMDSv2 with hop limit 1; minimal node role with the CNI on its own role; one role per workload with least-privilege policies (or ABAC by namespace tag); permissions boundaries for delegated roles; bucket policies that restrict to expected principals or VPC endpoints; IAM Access Analyzer for over-broad access; and an alert on access-denied spikes and unusual principals.
Q16. How do you migrate a cluster from the aws-auth ConfigMap to access entries without locking anyone out?
- Switch authentication mode to
API_AND_CONFIG_MAP(both work). - Create access entries for every role and user mapping in
aws-auth, including node roles (EKS creates entries for managed node groups) and the break-glass role. - Verify each principal's access with
kubectl auth can-iwhile assuming it. - Remove the entries from
aws-auth, watch authenticator logs for failures, then switch toAPI(a one-way change). - Manage access entries only in Terraform from then on, with CloudTrail alerts on out-of-band changes.
Q17. How do you let 40 teams create their own IAM roles for pods without the platform team becoming a bottleneck, and without privilege escalation?
Self-service through a pipeline or module with guard-rails: every created role must have a platform-owned permissions boundary and a naming prefix; SCPs deny creating roles without the boundary; policies are templated per resource type (the team's own S3 prefix, SQS queues tagged with their team); Pod Identity associations are only allowed for ServiceAccounts in the team's own namespaces (enforced in the pipeline and with an admission policy); ABAC session tags let one policy template serve all namespaces. Review with Access Analyzer's unused-access findings and quarterly access reviews.
Q18. What does a good break-glass design for the cluster look like?
A dedicated IAM role with a cluster-admin access entry, assumable only with MFA by a small group, usable when SSO or the normal path is down. Every use triggers an alert (CloudTrail AssumeRole event → notification), is recorded in the incident, and is followed by a review. Test it on a schedule (for example twice a year) so it works when needed. Keep a second path for when the API endpoint itself is unreachable from outside: a runner or bastion inside the VPC. Never make break-glass the everyday admin route.
Recap
- Work design questions through requirements → options → trade-offs → decision → failure modes → evidence.
- Design questions revolve around isolation vs cost, IP planning, compute strategy and identity chains.
- Continue with the second set: operations, security, cost and reliability.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.