Amazon EKS in Production with Terraform
An architect-level EKS track built with Terraform: networking with the VPC CNI and secondary CIDRs, identity with Pod Identity/IRSA and access entries, Karpenter for capacity, and a CI/CD flow from laptop to cluster.
What you'll be able to do
- Build a production EKS stack from Terraform with remote state
- Solve IP exhaustion with prefix delegation and secondary CIDRs
- Upgrade EKS clusters and add-ons safely
- Recover from Terraform state and cluster disasters
Before you start
AWS for Platform Engineers, Terraform & IaC.
How it works
Each lesson: plain-language idea → how it really works → hands-on. Each section ends with a cheat sheet & self-check.
Curriculum
Lessons marked “Read” are ready; the rest are on the way.
Build
- 01EKS architecture & mental modelWhat AWS runs, what you runRead →
- 02Remote state: S3, locking, env-per-tfvarsDrift and safe collaborationRead →
- 03The multi-provider chicken-and-eggBootstrapping providers that depend on the clusterRead →
- 04VPC CNI & IP planningPrefix delegation, secondary 100.64/16 pod CIDRRead →
- 📋Cheat sheet & self-checkEvery command from this section on one page, then 12 questions to check yourself.Open →
Operate
- 05IdentityAccess entries, Pod Identity, IRSARead →
- 06Capacity with KarpenterNodePools, consolidation, SpotRead →
- 07Storage tiersEBS and EFS CSI, hot/warm/coldRead →
- 08IngressAWS Load Balancer Controller, Gateway APIRead →
- 09Upgrades & add-onsVersion skew, managed add-ons, blue/green clustersRead →
- 📋Cheat sheet & self-checkEvery command from this section on one page, then 15 questions to check yourself.Open →
Automate & Recover
Playbooks, challenges & practice
- 13Playbook: creating the clusterPre-flight, run order by layer, verification gates, rollbackRead →
- 14Playbook: day-2 operationsNode pools, access, add-ons, upgrades, cost checks, safe teardownRead →
- 15The flow: laptop → cluster → kubeconfig → appHow humans and pipelines reach the cluster, the CI/CD wayRead →
- 16Challenges on this stackReal-world problems with symptoms, checks and fixesRead →
- 17Recovery playbookStep-by-step runbooks for state, access and cluster recoveryRead →
- 18Simulator: practise for $0terraform test mocks, LocalStack and kind before touching AWSRead →
- 📋Cheat sheet & self-checkEvery command from this section on one page, then 18 questions to check yourself.Open →
Real-world scenarios
Work through each one: symptom → misleading signal → evidence → root cause → prevention.
Subnet exhaustion, and why prefix delegation helps.
Recovery from state disasters.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.