Production EKS Platform — From Zero to Production›16 · Architecture & design review

Lesson 16 of 18 · Part 4 — Architect

Architecture & design review

Tie the whole track together into one reference architecture, write down the decisions and their trade-offs as ADRs, and run a design review with a thorough checklist.

Architect
Key wordsreference architecturearchitecture decision recordADRdesign reviewtrade-offsmulti-accountcluster topologymulti-tenancyWell-Architectedcostreliability
AWS-managed (you don't see it) EKS control plane API servers across 3 AZs etcd managed, backed up by AWS EKS API endpoint public and/or private Your VPC (you own it) Node groups / Karpenter EC2 instances Fargate (optional) Pods get VPC IPs Amazon VPC CNI (ENIs) IAM for pods IRSA / Pod Identity Add-ons CoreDNS, kube-proxy, CSI cross-account ENIs
The reference platform: accounts and VPC, the EKS control plane, the data plane, and the services around them.

The reference architecture

Putting the previous lessons together, a production platform typically looks like this:

Area Choice (default) Lesson
Accounts Separate prod and non-prod accounts in an Organization; SCP guard-rails; SSO 03, 06
Network VPC across 3 AZs; public subnets for LBs/NAT, private for nodes; secondary pod CIDR; endpoints for S3/ECR/STS 03, 04
Cluster Recent Kubernetes version; private (or CIDR-restricted) endpoint; auth mode API; KMS; audit logs 05, 12
Compute Small managed node group for system pods + Karpenter (On-Demand + Spot) for workloads 05, 10
Identity SSO roles via access entries; CI via OIDC role; Pod Identity per ServiceAccount; break-glass role 06
Governance Namespace per team; quotas and limit ranges; Pod Security restricted; Kyverno policies; tags for cost 06, 12
Traffic AWS Load Balancer Controller; ALB for HTTP with ACM certificates; NLB for TCP 07
Delivery ECR with immutable tags and scanning; CI builds; Argo CD deploys 08, 15
State EBS gp3 for RWO, EFS for RWX, S3 for objects; managed databases (RDS) for critical data 09
Scaling HPA/KEDA for pods, Karpenter for nodes, PDBs everywhere 10
Observability Control-plane logs; Prometheus + Grafana (managed or self-run); Fluent Bit; OTel; SLO alerts 11
Security IMDSv2, Bottlerocket/AL2023, network policies, GuardDuty, Secrets Manager + External Secrets 12
Lifecycle Routine upgrades every release cycle; add-ons pinned and tested; blue/green cluster option 13
Recovery Everything as code; GitOps rebuild; Velero for cluster state; data backups; tested restores 14

A design review is like an MOT test for a car before a long trip. The mechanic doesn't just look at the shiny paint; they check the brakes, tyres, lights and what happens in an emergency stop, and they want to see the service history as proof.

Write the decisions down: ADRs

For each significant choice, a one-page Architecture Decision Record:

ADR-007: Karpenter for workload capacity
Status:     Accepted (2026-05-14)
Context:    Spiky, mixed workloads; 40% idle capacity on fixed node groups; Spot savings desired.
Options:    (1) Managed node groups + Cluster Autoscaler  (2) Karpenter  (3) EKS Auto Mode
Decision:   Karpenter for workloads; one small managed node group for system pods and Karpenter itself.
Consequences:
  + Right-sized instances, consolidation, Spot diversification
  - We operate and upgrade Karpenter; must set disruption budgets and PDBs carefully
  - Revisit Auto Mode if node operations become a bottleneck

Good ADRs make reviews faster (the reasoning is already there) and make it safe to change your mind later.

The review checklist

Walk through it with evidence, not opinions:

Pillar Questions to answer
Reliability What happens when an AZ fails? A node? The control plane is unreachable for 10 minutes? Where are the single points of failure (single NAT, single replica, one-AZ volume)?
Security Who can reach the API and how? How do pods get AWS credentials? Can a pod reach the node role? How are secrets stored and rotated? What's enforced at admission?
Operability How is an upgrade done, how long does it take, and when was the last one? What pages on-call, and is there a runbook?
Performance & scale IP headroom? Service quotas? Max pods per node? How fast can capacity be added?
Cost Top five cost drivers (compute, NAT, load balancers, logs, cross-AZ traffic)? Spot share? Idle capacity? Cost per team?
Recoverability RPO and RTO per workload, and the date of the last restore test? Can the platform be rebuilt from code in another region?
Tenancy How are teams isolated? What stops one team from exhausting the cluster?

Typical trade-offs to be ready for

Decision Option A Option B What tips it
Cluster count Few large clusters Many small clusters Isolation and compliance needs vs operating cost
Compute Managed node groups Karpenter / Auto Mode Workload variety, Spot appetite, appetite for operations
Pod IPs Secondary CIDR (IPv4) IPv6 cluster Scale, and whether dependencies speak IPv6
Metrics Self-run Prometheus Managed Prometheus / Container Insights Team skills, volume, cost model
Ingress One shared ALB (IngressGroup) ALB per app Cost and simplicity vs blast radius and ownership
State In-cluster databases Managed services (RDS, ElastiCache) Operational maturity; usually managed for critical data
Upgrades In place Blue/green clusters Risk tolerance, statefulness, traffic-shift capability

Try it: run your own review

  1. Fill in the reference-architecture table for a cluster you know (or the one built in this track), marking every row where you differ and why.
  2. Write two ADRs: one for compute (lesson 05/10) and one for identity (lesson 06).
  3. Answer the checklist questions with evidence (commands, dashboards, test dates). Every "we think so" is a finding.
  4. List your top three risks with an owner and a date.

Going deeper: presenting to a review board

Lead with the business context (availability target, compliance, growth), then the diagram, then the five decisions that matter most with their trade-offs, then risks and what you're doing about them. Keep detail in the appendix. Reviewers trust designs that show what was rejected and why, and that name their own weaknesses before being asked.

Recap

  • One reference architecture ties accounts, network, cluster, identity, governance, traffic, delivery, state, scaling, observability, security, lifecycle and recovery together.
  • Record decisions as ADRs with options and consequences.
  • Review against reliability, security, operability, scale, cost, recoverability and tenancy, with evidence.
  • Know the classic trade-offs and what tips each one; the Q&A lessons that follow practise exactly that.

This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.