Manage EKS with Terraform
The as-code companion to Production EKS Platform: lay out a Terraform repo for EKS, protect state, build the VPC, cluster, IAM and add-ons as code, run it through a CI/CD pipeline, upgrade and change it safely, handle drift and imports, and recover when state or the cluster goes wrong. Ends with runbooks, a free practice setup and a set of questions with detailed answers.
What you'll be able to do
- Structure an EKS Terraform repo in layers with safe, repeatable run order
- Build the VPC, cluster, node groups, access and pod identities as code
- Hand platform add-ons from Terraform to GitOps at the right boundary
- Run plan/apply from a pipeline with short-lived OIDC credentials
- Upgrade, refactor and import without replacing the cluster
- Recover from state and cluster disasters with tested runbooks
Before you start
Terraform & Infrastructure as Code, Production EKS Platform (Part 1–2).
How it works
Each lesson: plain-language idea → how it really works → hands-on. Each section ends with a cheat sheet & self-check.
Curriculum
Lessons marked “Read” are ready; the rest are on the way.
Part 1 — Foundations
- 01Repo layout & run orderLayers, environments, module pinning and the order things must be appliedRead →
- 02Remote state: S3, locking, env-per-tfvarsProtecting the source of truth; drift and safe collaborationRead →
- 03The multi-provider chicken-and-eggProviders that depend on a cluster that doesn't exist yetRead →
- 📋Cheat sheet & self-checkEvery command from this section on one page, then 9 questions to check yourself.Open →
Part 2 — Build as code
- 04The VPC as codeSubnets, NAT, endpoints, load-balancer tags and a secondary pod CIDRRead →
- 05The EKS cluster as codeCluster, authentication mode, managed node groups, add-ons, with raw resources and the community moduleRead →
- 06IAM as code: access entries, Pod Identity & IRSARoles, trust and permissions policies, associations, least privilege in TerraformRead →
- 07Add-ons, Karpenter & the GitOps hand-offWhat Terraform installs, what Argo CD owns, and the line between themRead →
- 📋Cheat sheet & self-checkEvery command from this section on one page, then 12 questions to check yourself.Open →
Part 3 — Operate as code
- 08CI/CD for the stackPlan/apply pipelines with OIDC to AWS and approval gatesRead →
- 09Upgrades with TerraformControl plane, add-ons and node groups in the right order, without surprisesRead →
- 10Drift, import, moved & safe destroyBring existing resources under code, refactor without replacing, tear down safelyRead →
- 11Playbook: creating the clusterPre-flight, run order by layer, verification gates, rollbackRead →
- 12Playbook: day-2 operationsNode pools, access, add-ons, upgrades, cost checks, safe teardownRead →
- 13Challenges on this stackReal-world problems with symptoms, checks and fixesRead →
- 14Recovery playbookStep-by-step runbooks for state, access and cluster recoveryRead →
- 15Simulator: practise for $0terraform test mocks, LocalStack and kind before touching AWSRead →
- 16Capstone: the whole platform from an empty accountBuild, prove, hand overRead →
- 📋Cheat sheet & self-checkEvery command from this section on one page, then 25 questions to check yourself.Open →
Part 4 — Q&A
Real-world scenarios
Work through each one: symptom → misleading signal → evidence → root cause → prevention.
One innocent change forces a new cluster. Find the attribute and change it safely.
Recover state and resources after the worst command of the week.
The cluster is gone before the resources that live in it. Untangle the order.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.