Lesson 16 of 18 · Part 4 — Architect
Architecture & design review
Tie the whole track together into one reference architecture, write down the decisions and their trade-offs as ADRs, and run a design review with a thorough checklist.
The reference architecture
Putting the previous lessons together, a production platform typically looks like this:
| Area | Choice (default) | Lesson |
|---|---|---|
| Accounts | Separate prod and non-prod accounts in an Organization; SCP guard-rails; SSO | 03, 06 |
| Network | VPC across 3 AZs; public subnets for LBs/NAT, private for nodes; secondary pod CIDR; endpoints for S3/ECR/STS | 03, 04 |
| Cluster | Recent Kubernetes version; private (or CIDR-restricted) endpoint; auth mode API; KMS; audit logs |
05, 12 |
| Compute | Small managed node group for system pods + Karpenter (On-Demand + Spot) for workloads | 05, 10 |
| Identity | SSO roles via access entries; CI via OIDC role; Pod Identity per ServiceAccount; break-glass role | 06 |
| Governance | Namespace per team; quotas and limit ranges; Pod Security restricted; Kyverno policies; tags for cost |
06, 12 |
| Traffic | AWS Load Balancer Controller; ALB for HTTP with ACM certificates; NLB for TCP | 07 |
| Delivery | ECR with immutable tags and scanning; CI builds; Argo CD deploys | 08, 15 |
| State | EBS gp3 for RWO, EFS for RWX, S3 for objects; managed databases (RDS) for critical data | 09 |
| Scaling | HPA/KEDA for pods, Karpenter for nodes, PDBs everywhere | 10 |
| Observability | Control-plane logs; Prometheus + Grafana (managed or self-run); Fluent Bit; OTel; SLO alerts | 11 |
| Security | IMDSv2, Bottlerocket/AL2023, network policies, GuardDuty, Secrets Manager + External Secrets | 12 |
| Lifecycle | Routine upgrades every release cycle; add-ons pinned and tested; blue/green cluster option | 13 |
| Recovery | Everything as code; GitOps rebuild; Velero for cluster state; data backups; tested restores | 14 |
A design review is like an MOT test for a car before a long trip. The mechanic doesn't just look at the shiny paint; they check the brakes, tyres, lights and what happens in an emergency stop, and they want to see the service history as proof.
Write the decisions down: ADRs
For each significant choice, a one-page Architecture Decision Record:
ADR-007: Karpenter for workload capacity
Status: Accepted (2026-05-14)
Context: Spiky, mixed workloads; 40% idle capacity on fixed node groups; Spot savings desired.
Options: (1) Managed node groups + Cluster Autoscaler (2) Karpenter (3) EKS Auto Mode
Decision: Karpenter for workloads; one small managed node group for system pods and Karpenter itself.
Consequences:
+ Right-sized instances, consolidation, Spot diversification
- We operate and upgrade Karpenter; must set disruption budgets and PDBs carefully
- Revisit Auto Mode if node operations become a bottleneck
Good ADRs make reviews faster (the reasoning is already there) and make it safe to change your mind later.
The review checklist
Walk through it with evidence, not opinions:
| Pillar | Questions to answer |
|---|---|
| Reliability | What happens when an AZ fails? A node? The control plane is unreachable for 10 minutes? Where are the single points of failure (single NAT, single replica, one-AZ volume)? |
| Security | Who can reach the API and how? How do pods get AWS credentials? Can a pod reach the node role? How are secrets stored and rotated? What's enforced at admission? |
| Operability | How is an upgrade done, how long does it take, and when was the last one? What pages on-call, and is there a runbook? |
| Performance & scale | IP headroom? Service quotas? Max pods per node? How fast can capacity be added? |
| Cost | Top five cost drivers (compute, NAT, load balancers, logs, cross-AZ traffic)? Spot share? Idle capacity? Cost per team? |
| Recoverability | RPO and RTO per workload, and the date of the last restore test? Can the platform be rebuilt from code in another region? |
| Tenancy | How are teams isolated? What stops one team from exhausting the cluster? |
Typical trade-offs to be ready for
| Decision | Option A | Option B | What tips it |
|---|---|---|---|
| Cluster count | Few large clusters | Many small clusters | Isolation and compliance needs vs operating cost |
| Compute | Managed node groups | Karpenter / Auto Mode | Workload variety, Spot appetite, appetite for operations |
| Pod IPs | Secondary CIDR (IPv4) | IPv6 cluster | Scale, and whether dependencies speak IPv6 |
| Metrics | Self-run Prometheus | Managed Prometheus / Container Insights | Team skills, volume, cost model |
| Ingress | One shared ALB (IngressGroup) | ALB per app | Cost and simplicity vs blast radius and ownership |
| State | In-cluster databases | Managed services (RDS, ElastiCache) | Operational maturity; usually managed for critical data |
| Upgrades | In place | Blue/green clusters | Risk tolerance, statefulness, traffic-shift capability |
Try it: run your own review
- Fill in the reference-architecture table for a cluster you know (or the one built in this track), marking every row where you differ and why.
- Write two ADRs: one for compute (lesson 05/10) and one for identity (lesson 06).
- Answer the checklist questions with evidence (commands, dashboards, test dates). Every "we think so" is a finding.
- List your top three risks with an owner and a date.
Going deeper: presenting to a review board
Lead with the business context (availability target, compliance, growth), then the diagram, then the five decisions that matter most with their trade-offs, then risks and what you're doing about them. Keep detail in the appendix. Reviewers trust designs that show what was rejected and why, and that name their own weaknesses before being asked.
Recap
- One reference architecture ties accounts, network, cluster, identity, governance, traffic, delivery, state, scaling, observability, security, lifecycle and recovery together.
- Record decisions as ADRs with options and consequences.
- Review against reliability, security, operability, scale, cost, recoverability and tenancy, with evidence.
- Know the classic trade-offs and what tips each one; the Q&A lessons that follow practise exactly that.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.