Manage EKS with Terraform›01 · Repo layout & run order
Learning Hub / Cloud — OpenStack, AWS & EKS / Manage EKS with Terraform

Lesson 01 of 17 · Part 1 — Foundations

Repo layout & run order

Structure a Terraform repository for EKS so it stays safe as it grows: separate layers with their own state, environments as inputs, pinned modules and providers, a fixed run order, and clean ways to pass outputs between layers.

Advanced
Key wordsTerraformrepository layoutlayersroot modulesenvironmentstfvarsmodule versioningprovider pinning.terraform.lock.hcldefault_tagsrun orderremote state outputsSSM parameters
0 · Pre-flight access, quotas verify gate before next layer 1 · State S3 + locking verify gate before next layer 2 · Network VPC, subnets verify gate before next layer 3 · Cluster EKS, nodes verify gate before next layer 4 · Platform LB, Karpenter verify gate before next layer 5 · Apps GitOps, smoke verify hand-over checklist each layer has its own state; apply in order, destroy in reverse never start the next layer until the verification gate passes
Layers applied in order, each with its own state; destroy runs in reverse.

One big state is the first mistake

A single Terraform configuration that creates the VPC, the cluster, IAM, Helm releases and Kubernetes objects works for a demo and hurts in production:

  • Blast radius: a typo in a Helm value can sit in the same plan as a VPC replacement.
  • Speed: every plan refreshes hundreds of resources.
  • Chicken-and-egg: Kubernetes and Helm providers need a cluster that the same plan is still creating (lesson 03).
  • Ownership: network, platform and app teams all touch one state.

You don't keep your passport, your shopping list and the house deeds in one envelope. If you spill coffee on the shopping list, you don't want to be replacing the deeds too.

Layers

Layer Contains Changes
00-bootstrap State bucket, KMS key, CI OIDC role (applied once, often by hand) Rarely
10-network VPC, subnets, NAT, endpoints, tags Rarely
20-cluster EKS cluster, node groups, access entries, core managed add-ons, KMS Upgrades, access changes
30-platform IAM roles for controllers, Pod Identity associations, Karpenter, Argo CD bootstrap Platform releases
40-apps (optional) Per-team IAM roles, queues and buckets for apps Often; can be team-owned

Kubernetes workloads themselves (Deployments, Ingresses) are usually not in Terraform at all: Argo CD owns them (lesson 07).

Repository layout

eks-platform/
├── modules/                    # reusable, versioned building blocks
│   ├── network/
│   ├── cluster/
│   └── platform/
└── live/                       # one root module per environment and layer
    ├── nonprod/
    │   ├── 10-network/   { main.tf, backend.tf, nonprod.tfvars }
    │   ├── 20-cluster/
    │   └── 30-platform/
    └── prod/
        ├── 10-network/
        ├── 20-cluster/
        └── 30-platform/

Each folder under live/ is a root module with its own backend key (prod/20-cluster.tfstate). Environments differ only by inputs:

# live/prod/20-cluster/prod.tfvars
cluster_name       = "prod"
kubernetes_version = "1.33"
system_node_count  = 3
admin_role_arns    = ["arn:aws:iam::111122223333:role/aws-reserved/sso.amazonaws.com/AWSReservedSSO_PlatformAdmin_abc123"]

Some teams prefer one root module per layer with a workspace or tfvars file per environment; that works too, as long as each environment has its own state and prod can't be applied by accident (separate credentials per account help).

Pin everything

terraform {
  required_version = ">= 1.10"
  required_providers {
    aws        = { source = "hashicorp/aws",        version = "~> 6.0" }
    kubernetes = { source = "hashicorp/kubernetes", version = "~> 2.30" }
    helm       = { source = "hashicorp/helm",       version = "~> 3.0" }
  }
}

provider "aws" {
  region = var.region
  default_tags {
    tags = { platform = "eks", environment = var.environment, managed-by = "terraform" }
  }
}

module "cluster" {
  source = "git::https://git.example.com/platform/eks-platform.git//modules/cluster?ref=v3.4.0"
  # ...
}
  • Commit .terraform.lock.hcl so provider versions and checksums are identical everywhere.
  • Reference internal modules by tag, community modules by exact or ~> version, and upgrade them in their own pull requests.
  • default_tags give every AWS resource environment and ownership tags for cost and audit. (Use provider major versions your team has tested; the constraints above are illustrative.)

Passing outputs between layers

The cluster layer needs the VPC's subnet IDs; the platform layer needs the cluster name and OIDC issuer. Two good ways:

Method How Trade-off
terraform_remote_state Read the other layer's outputs from its state Simple; readers need read access to that whole state
SSM Parameter Store (or tags + data sources) The producing layer writes /platform/prod/vpc_id; consumers read it Decoupled, least privilege, readable by other tools
# 10-network publishes
resource "aws_ssm_parameter" "private_subnets" {
  name  = "/platform/${var.environment}/private_subnet_ids"
  type  = "StringList"
  value = join(",", module.vpc.private_subnets)
}

# 20-cluster consumes
data "aws_ssm_parameter" "private_subnets" {
  name = "/platform/${var.environment}/private_subnet_ids"
}
locals { private_subnet_ids = split(",", data.aws_ssm_parameter.private_subnets.value) }

Run order

Create top-down and destroy bottom-up:

apply:    00-bootstrap → 10-network → 20-cluster → 30-platform → 40-apps
destroy:  40-apps → 30-platform → 20-cluster → 10-network   (bootstrap last, if ever)

The pipeline (lesson 08) encodes this order, so nobody has to remember it at 2 a.m.

Try it: split a monolith

  1. Take any single-folder EKS example and list its resources by layer (network, cluster, platform).
  2. Create the live/<env>/<layer> folders with separate backends.
  3. Move resources between states with moved/removed blocks or terraform state mv into the new layers (lesson 10 shows how), until terraform plan shows no changes in every layer.
  4. Replace one terraform_remote_state with an SSM parameter.

Going deeper: orchestration tools

When the number of environments and layers grows, tools such as Terragrunt, Terramate or a platform like Spacelift/HCP Terraform handle dependency order, DRY backends and parallel runs across layers. They're conveniences on top of the same principles: small states, pinned versions, inputs per environment, explicit order.

Recap

  • Layers with separate state: bootstrap, network, cluster, platform, (apps). Workloads belong to GitOps.
  • Same modules, different tfvars per environment; each environment in its own state (and ideally account).
  • Pin Terraform, providers (lock file) and modules; tag everything with default_tags.
  • Pass outputs through remote state or SSM parameters; apply top-down, destroy bottom-up.

This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.