Lesson 18 of 18 · Part 5 — Design & operations Q&A
Q&A: operations, security, cost & reliability
Common questions about operating EKS in production (upgrades, scaling, observability, security, disaster recovery, cost and multi-tenancy), plus troubleshooting scenarios, with detailed answers.
How to use these questions
Same approach as part 1: think each one through, then open the answer. Operations questions reward three things: a routine (how it's done every time), evidence (numbers, drills) and lessons learned from failures.
Anyone can fly a plane on a sunny day. What matters is how the crew handles the engine warning at night, and the checklist habits that make those nights rare.
Upgrades and lifecycle
Q1. What does a safe upgrade process look like for a fleet of 15 EKS clusters?
Upgrades are a routine, run every release cycle, not a project:
- Prepare: read the Kubernetes and EKS release notes; check cluster insights and scan manifests and Helm charts for removed APIs (e.g.
pluto,kubent); check add-on and controller compatibility (LB Controller, Karpenter, CSI drivers, policy engine, service mesh). - Order of clusters: sandbox → dev → staging → a low-risk prod → the rest, with soak time between waves.
- Per cluster: control plane → managed add-ons to compatible versions → nodes (Karpenter drift or node group rolling update) with PDBs protecting workloads.
- Validate: smoke tests, SLO dashboards, error budgets; automatic halt of the wave on regression.
- Everything in code: version bumps are pull requests through Terraform and GitOps.
Metrics that show maturity: time from a new EKS version to fleet-wide adoption, and whether any cluster ever entered extended support unintentionally.
Q2. What's EKS extended support, and how do you avoid paying for it by accident?
Each Kubernetes minor version has a standard support period on EKS, followed by an extended support period at a higher per-cluster hourly price. At the end of extended support, AWS upgrades the cluster automatically. Avoid surprises by tracking version end dates in the fleet inventory, alerting well in advance, keeping upgrades routine (Q1), and using the cluster's upgrade policy setting to decide explicitly whether a cluster may enter extended support. Budget for it only as a deliberate exception.
Q3. In-place upgrade or blue/green cluster? When do you choose each?
In place is the default: cheaper, and fine when upgrades are routine and workloads tolerate rolling nodes.
Blue/green clusters (build a new cluster at the new version, move traffic gradually, keep the old one for rollback) when: the jump is large (several versions behind), a major platform change coincides (CNI swap, new node OS), risk tolerance is very low, or rollback must be instant. The cost: double capacity during the move, traffic-shifting capability (Route 53 weighted records or a global load balancer), and a plan for state: stateful workloads and data must be decoupled or migrated. That's easy with managed databases and much harder with in-cluster state.
Q4. An upgrade stopped halfway: control plane on the new version, half the nodes old, and an add-on failing. What do you do?
Stabilise first: the version skew policy allows kubelets to lag, so the cluster is still supported; don't rush. Identify the failing add-on's error (describe the add-on, its pods' logs); move it to a version compatible with the new control plane, or roll it back to its previous version if that's still compatible. Pause the node rollout (PDBs and Karpenter disruption budgets), check workloads on both node versions, then resume. Control-plane downgrades aren't possible, so the path is forward. Afterwards: add the missed compatibility check to the pre-flight list.
Scaling and capacity
Q5. Traffic spikes 10× within two minutes every morning. How do you make the platform keep up?
Scaling has latency at every layer: metric collection → HPA decision → pod scheduling → node launch → image pull → app warm-up. Reduce each:
- Predictable spike: scheduled scaling (KEDA cron scaler or scheduled HPA min replicas) ahead of time.
- Headroom: low-priority "balloon" pods that reserve capacity and get pre-empted instantly.
- Faster nodes: Karpenter with broad instance choices; smaller images, pre-pulled or cached images; fast-booting AMIs (Bottlerocket).
- Faster pods: readiness that reflects real warm-up, tuned HPA behaviour (scale up fast, down slowly).
- Upstream: ALB pre-warming is rarely needed now, but check quotas (EC2 vCPU, IPs) so scaling isn't blocked.
Validate with a load test that reproduces the curve.
Q6. Karpenter keeps churning nodes and pods restart too often. How do you tune it?
Look at disruption reasons (consolidation, drift, expiry, interruption). Then: set disruption budgets on NodePools (for example at most 10% of nodes at once, none during business peaks); use consolidateAfter to wait before consolidating; ensure PDBs exist for every multi-replica workload; mark pods that must not move with karpenter.sh/do-not-disrupt; check that requests are realistic (inflated requests cause oscillation); and separate NodePools for workloads with different tolerance. Measure node lifetime and pod restarts per day before and after.
Observability and operations
Q7. What would you monitor and alert on for the platform itself (not the apps)?
Page-worthy: API server error rate and latency; nodes NotReady across more than one AZ; pods Pending beyond a threshold (capacity, quota or IP exhaustion); CoreDNS errors and latency; free IPs per subnet; certificate expiry; failed Karpenter launches; add-on health; ALB target groups with no healthy targets; GuardDuty high-severity findings; break-glass use.
Dashboards (not pages): utilisation, cost, upgrade status, quota headroom. Every alert has an owner and a runbook, and the alert set is reviewed after every incident.
Q8. How do you define SLOs for a shared platform?
From the tenants' point of view: deploy success rate and time, API availability (kubectl and controllers), pod scheduling latency (time from created to running at normal load), ingress availability and latency, and DNS success rate. Set targets per environment tier, measure with synthetic checks and platform metrics, publish error budgets, and use them to decide between feature work and reliability work. Make it explicit what the platform does not guarantee (application availability depends on their replicas and probes).
Security
Q9. What does defence in depth look like for EKS?
Layers, each assuming the one before can fail:
- Account and network: separate accounts, SCPs, private subnets, private or restricted API endpoint.
- Identity: SSO roles via access entries, OIDC for CI, Pod Identity per workload, no long-lived keys, break-glass with alerts.
- Supply chain: scanned and signed images from ECR only, verified at admission.
- Admission: Pod Security
restricted, Kyverno policies (registries, limits, labels). - Runtime: IMDSv2 hop limit 1, minimal node role, Bottlerocket/AL2023, network policies default-deny, security groups for pods for data stores.
- Data: KMS encryption, Secrets Manager with rotation, TLS in transit.
- Detection and response: audit logs, GuardDuty (audit + runtime), CloudTrail, alerts on risky changes, practised incident runbooks.
Q10. GuardDuty reports crypto-mining in a pod. What should happen in the first 30 minutes?
- Contain: isolate the pod with a deny-all network policy (and cordon the node); don't delete it yet, evidence matters.
- Scope: which image, namespace, ServiceAccount and role? What did that role do (CloudTrail)? Are other pods running the same image?
- Preserve: capture pod spec, logs, and a node snapshot if the runtime could be compromised.
- Eradicate: revoke the role's sessions (deny policy with a time condition), rotate any secrets it could read, replace the node, redeploy from a known-good image.
- Recover and learn: how did it get in (vulnerable image, exposed endpoint, leaked credentials)? Add the missing control (admission rule, patch, network policy) and a detection test.
Q11. How do you manage secrets on EKS across 40 teams?
Source of truth in AWS Secrets Manager (or Parameter Store for non-sensitive config), organised by team path and encrypted with KMS. Delivery either through the External Secrets Operator (synced into namespaced Kubernetes Secrets, with the operator's access scoped by team) or read directly by apps with their own pod roles. Rotation handled by Secrets Manager where possible, apps reloading on change. No secrets in Git (sealed or encrypted values only if unavoidable), no secrets in environment variables of manifests in Git, audit access through CloudTrail.
Reliability and disaster recovery
Q12. The business wants RPO 5 minutes and RTO 1 hour for the order service, including a regional outage. How can that be designed?
Split the problem into platform and data:
- Platform: a standby cluster in a second region built from the same Terraform and GitOps (warm standby: small, scaled up on failover), images replicated with ECR replication, secrets replicated.
- Data: RPO 5 minutes rules out daily backups; use a managed database with cross-region replication (Aurora Global Database or DynamoDB global tables), queues and object storage replicated (S3 CRR).
- Traffic: Route 53 health-checked failover or a global accelerator.
- Process: a written, practised failover runbook; RTO measured in game days, not assumed.
Discuss the trade-off: active-active would lower RTO but adds write-conflict and cost complexity; warm standby is the usual balance for a one-hour RTO.
Q13. What does Velero protect, and what doesn't it?
Velero backs up Kubernetes objects (through the API) and can snapshot persistent volumes (EBS snapshots or file-level backup), and restores them into the same or another cluster. It's useful for namespace recovery, migrations and clusters with in-cluster state.
It doesn't replace: GitOps as the rebuild source for platform and app manifests, managed database backups and replication, or tested restore procedures. In a well-designed platform, most of the cluster is rebuilt from Git, and Velero covers the stateful gaps and accidental deletions.
Cost
Q14. The EKS bill grew 60% in a quarter without matching traffic growth. How do you investigate and fix it?
Break it down by driver with Cost Explorer and split cost allocation data (by cluster, namespace, label): compute (idle capacity, over-requested pods, On-Demand share), NAT and data transfer (cross-AZ traffic, pulls through NAT), load balancers (one ALB per app), EBS (orphaned volumes, oversized gp3), observability (log ingestion, metric cardinality), extended support fees.
Fixes, roughly by impact: right-size requests (VPA recommendations), Karpenter consolidation and Spot, Graviton instances where images support arm64, S3/ECR endpoints, shared ALBs with IngressGroups, log filtering and retention, cleaning orphaned volumes and snapshots, and Savings Plans for the stable baseline. Make it stick with per-team showback and budgets.
Q15. How do you do chargeback or showback for teams sharing clusters?
Label everything consistently (team, cost-centre, app) and enforce the labels at admission. Use EKS split cost allocation data in the Cost and Usage Report (CPU and memory cost per pod, based on requests) or OpenCost/Kubecost for real-time views. Allocate shared costs (control plane, system pods, idle capacity) by a published rule. Start with showback, move to chargeback once numbers are trusted, and pair it with guidance on right-sizing so teams can act on it.
Multi-tenancy and governance
Q16. How do you stop one team from harming others in a shared cluster?
Resource: ResourceQuota and LimitRange per namespace, PriorityClasses so platform components win under pressure, PDBs. Access: access entries scoped to their namespaces only. Security: Pod Security restricted, Kyverno policies, network policies default-deny. Capacity: separate NodePools (and taints) for noisy or special workloads. Blast radius: rate limits on shared services, separate ingress groups for critical apps. Observability: per-team dashboards and quota alerts. And a clear escalation path when limits bite.
Q17. How do you migrate 200 services from self-managed on-prem Kubernetes to EKS?
Assess and group services by complexity (stateless, stateful, special hardware, compliance). Build the target platform first with the same GitOps model, so workloads move as manifests plus configuration changes (storage classes, ingress annotations, secrets source, IAM via Pod Identity instead of static credentials). Replace in-cluster platform pieces with AWS equivalents where it reduces operations (load balancers, managed databases). Migrate in waves, starting with low-risk stateless services; use DNS-weighted cutover with rollback; move data with replication, not big-bang copies. Track progress and incidents per wave, and decommission on-prem capacity as waves complete.
Troubleshooting scenarios
Q18. "The API server is slow and kubectl times out." Where do you look?
Scope: all clients or one? From inside the VPC and outside? Check the EKS API metrics (request latency, 429s from API priority and fairness) and control-plane logs. Common causes: a controller or operator hammering the API (list/watch storms, too many objects), a huge number of objects (secrets, events, CRDs), admission webhooks that are slow or unreachable (a webhook with failurePolicy: Fail whose pods are down blocks writes), or network issues on the client side (endpoint access, DNS, proxy). Fix the noisy client, set webhook timeouts and failure policies appropriately, clean up object sprawl, and alert on API latency.
Q19. "Deployments succeed but the ALB returns 502/504 for a minute on every release." Why?
Pods are removed from the load balancer later than they stop serving, or added before they're ready. Fix: readiness probes that reflect real readiness; a preStop sleep longer than the target deregistration propagation; terminationGracePeriodSeconds above that; IP-mode targets with pod readiness gates from the Load Balancer Controller so rollouts wait for targets to be healthy in the ALB; maxUnavailable: 0. Check the target group's deregistration delay to match the app's connection behaviour.
Q20. "A whole AZ became unhealthy for 40 minutes." What should have happened on a well-designed platform, and what do you check afterwards?
Expected: replicas spread across three AZs keep serving (topology spread, PDBs); Karpenter launches replacement capacity in healthy AZs; NAT per AZ keeps the other AZs' egress working; ALBs stop routing to unhealthy targets; stateful workloads on EBS in that AZ are the weak point (they can't move), so they should be replicated at the application level or use managed services.
Afterwards: did any service lose all replicas (spread not enforced)? Did capacity or IPs run short in the remaining AZs? Were there single-AZ dependencies (one NAT, one-AZ volume, one-AZ database)? Turn every finding into a test for the next game day, where you drain an AZ on purpose.
Recap
- Good operations practice rests on routine, evidence and learning from failure.
- Know the upgrade order and support calendar, scaling latency chain, platform SLOs, defence in depth, DR split into platform and data, and the cost drivers.
- Work through the troubleshooting scenarios on a sandbox cluster where you can; they build the instincts that matter during incidents.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.