Lesson 01 of 17 · Part 1 — Foundations
Repo layout & run order
Structure a Terraform repository for EKS so it stays safe as it grows: separate layers with their own state, environments as inputs, pinned modules and providers, a fixed run order, and clean ways to pass outputs between layers.
One big state is the first mistake
A single Terraform configuration that creates the VPC, the cluster, IAM, Helm releases and Kubernetes objects works for a demo and hurts in production:
- Blast radius: a typo in a Helm value can sit in the same plan as a VPC replacement.
- Speed: every plan refreshes hundreds of resources.
- Chicken-and-egg: Kubernetes and Helm providers need a cluster that the same plan is still creating (lesson 03).
- Ownership: network, platform and app teams all touch one state.
You don't keep your passport, your shopping list and the house deeds in one envelope. If you spill coffee on the shopping list, you don't want to be replacing the deeds too.
Layers
| Layer | Contains | Changes |
|---|---|---|
00-bootstrap |
State bucket, KMS key, CI OIDC role (applied once, often by hand) | Rarely |
10-network |
VPC, subnets, NAT, endpoints, tags | Rarely |
20-cluster |
EKS cluster, node groups, access entries, core managed add-ons, KMS | Upgrades, access changes |
30-platform |
IAM roles for controllers, Pod Identity associations, Karpenter, Argo CD bootstrap | Platform releases |
40-apps (optional) |
Per-team IAM roles, queues and buckets for apps | Often; can be team-owned |
Kubernetes workloads themselves (Deployments, Ingresses) are usually not in Terraform at all: Argo CD owns them (lesson 07).
Repository layout
eks-platform/
├── modules/ # reusable, versioned building blocks
│ ├── network/
│ ├── cluster/
│ └── platform/
└── live/ # one root module per environment and layer
├── nonprod/
│ ├── 10-network/ { main.tf, backend.tf, nonprod.tfvars }
│ ├── 20-cluster/
│ └── 30-platform/
└── prod/
├── 10-network/
├── 20-cluster/
└── 30-platform/
Each folder under live/ is a root module with its own backend key (prod/20-cluster.tfstate). Environments differ only by inputs:
# live/prod/20-cluster/prod.tfvars
cluster_name = "prod"
kubernetes_version = "1.33"
system_node_count = 3
admin_role_arns = ["arn:aws:iam::111122223333:role/aws-reserved/sso.amazonaws.com/AWSReservedSSO_PlatformAdmin_abc123"]
Some teams prefer one root module per layer with a workspace or tfvars file per environment; that works too, as long as each environment has its own state and prod can't be applied by accident (separate credentials per account help).
Pin everything
terraform {
required_version = ">= 1.10"
required_providers {
aws = { source = "hashicorp/aws", version = "~> 6.0" }
kubernetes = { source = "hashicorp/kubernetes", version = "~> 2.30" }
helm = { source = "hashicorp/helm", version = "~> 3.0" }
}
}
provider "aws" {
region = var.region
default_tags {
tags = { platform = "eks", environment = var.environment, managed-by = "terraform" }
}
}
module "cluster" {
source = "git::https://git.example.com/platform/eks-platform.git//modules/cluster?ref=v3.4.0"
# ...
}
- Commit
.terraform.lock.hclso provider versions and checksums are identical everywhere. - Reference internal modules by tag, community modules by exact or
~>version, and upgrade them in their own pull requests. default_tagsgive every AWS resource environment and ownership tags for cost and audit. (Use provider major versions your team has tested; the constraints above are illustrative.)
Passing outputs between layers
The cluster layer needs the VPC's subnet IDs; the platform layer needs the cluster name and OIDC issuer. Two good ways:
| Method | How | Trade-off |
|---|---|---|
terraform_remote_state |
Read the other layer's outputs from its state | Simple; readers need read access to that whole state |
| SSM Parameter Store (or tags + data sources) | The producing layer writes /platform/prod/vpc_id; consumers read it |
Decoupled, least privilege, readable by other tools |
# 10-network publishes
resource "aws_ssm_parameter" "private_subnets" {
name = "/platform/${var.environment}/private_subnet_ids"
type = "StringList"
value = join(",", module.vpc.private_subnets)
}
# 20-cluster consumes
data "aws_ssm_parameter" "private_subnets" {
name = "/platform/${var.environment}/private_subnet_ids"
}
locals { private_subnet_ids = split(",", data.aws_ssm_parameter.private_subnets.value) }
Run order
Create top-down and destroy bottom-up:
apply: 00-bootstrap → 10-network → 20-cluster → 30-platform → 40-apps
destroy: 40-apps → 30-platform → 20-cluster → 10-network (bootstrap last, if ever)
The pipeline (lesson 08) encodes this order, so nobody has to remember it at 2 a.m.
Try it: split a monolith
- Take any single-folder EKS example and list its resources by layer (network, cluster, platform).
- Create the
live/<env>/<layer>folders with separate backends. - Move resources between states with
moved/removedblocks orterraform state mvinto the new layers (lesson 10 shows how), untilterraform planshows no changes in every layer. - Replace one
terraform_remote_statewith an SSM parameter.
Going deeper: orchestration tools
When the number of environments and layers grows, tools such as Terragrunt, Terramate or a platform like Spacelift/HCP Terraform handle dependency order, DRY backends and parallel runs across layers. They're conveniences on top of the same principles: small states, pinned versions, inputs per environment, explicit order.
Recap
- Layers with separate state: bootstrap, network, cluster, platform, (apps). Workloads belong to GitOps.
- Same modules, different tfvars per environment; each environment in its own state (and ideally account).
- Pin Terraform, providers (lock file) and modules; tag everything with
default_tags. - Pass outputs through remote state or SSM parameters; apply top-down, destroy bottom-up.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.