Manage EKS with Terraform›05 · The EKS cluster as code
Learning Hub / Cloud — OpenStack, AWS & EKS / Manage EKS with Terraform

Lesson 05 of 17 · Part 2 — Build as code

The EKS cluster as code

Build the cluster layer in Terraform: the EKS cluster with a private endpoint, API authentication mode, KMS and logging; a managed node group with a hardened launch template; and managed add-ons, first with raw resources so you see every setting, then with the community module.

Advanced
Key wordsTerraformaws_eks_clusteraccess_configauthentication_modebootstrap_cluster_creator_admin_permissionsencryption_configupgrade_policyaws_eks_node_grouplaunch templateIMDSv2aws_eks_addonaws_eks_addon_versionterraform-aws-modules/eks

Raw resources first

The community module is convenient, but writing the cluster with plain aws_* resources once shows you every decision. (Arguments below match current AWS provider versions; check the provider documentation for the version you pin.)

Before using a food processor, chop an onion by hand once. You'll understand what the machine is doing, and you'll know what to check when it jams.

The cluster

resource "aws_eks_cluster" "this" {
  name     = var.cluster_name
  version  = var.kubernetes_version
  role_arn = aws_iam_role.cluster.arn

  access_config {
    authentication_mode                         = "API"
    bootstrap_cluster_creator_admin_permissions = false   # admins are declared below
  }

  vpc_config {
    subnet_ids              = local.private_subnet_ids
    endpoint_private_access = true
    endpoint_public_access  = var.public_endpoint            # false in prod
    public_access_cidrs     = var.public_endpoint ? var.admin_cidrs : null
  }

  encryption_config {
    resources = ["secrets"]
    provider { key_arn = aws_kms_key.eks.arn }
  }

  enabled_cluster_log_types = ["api", "audit", "authenticator"]

  upgrade_policy {
    support_type = "STANDARD"    # don't drift into paid extended support silently
  }

  depends_on = [aws_cloudwatch_log_group.cluster]

  lifecycle {
    prevent_destroy = true
  }
}

resource "aws_cloudwatch_log_group" "cluster" {
  name              = "/aws/eks/${var.cluster_name}/cluster"   # create it first to control retention
  retention_in_days = 90
}

The cluster role:

resource "aws_iam_role" "cluster" {
  name = "${var.cluster_name}-cluster"
  assume_role_policy = jsonencode({
    Version   = "2012-10-17"
    Statement = [{ Effect = "Allow", Principal = { Service = "eks.amazonaws.com" }, Action = ["sts:AssumeRole", "sts:TagSession"] }]
  })
}

resource "aws_iam_role_policy_attachment" "cluster" {
  role       = aws_iam_role.cluster.name
  policy_arn = "arn:aws:iam::aws:policy/AmazonEKSClusterPolicy"
}

Admin access as code

resource "aws_eks_access_entry" "admins" {
  for_each      = toset(var.admin_role_arns)
  cluster_name  = aws_eks_cluster.this.name
  principal_arn = each.value
}

resource "aws_eks_access_policy_association" "admins" {
  for_each      = toset(var.admin_role_arns)
  cluster_name  = aws_eks_cluster.this.name
  principal_arn = each.value
  policy_arn    = "arn:aws:eks::aws:cluster-access-policy/AmazonEKSClusterAdminPolicy"
  access_scope { type = "cluster" }
  depends_on    = [aws_eks_access_entry.admins]
}

Team access (namespace-scoped) follows the same pattern and is covered in lesson 06.

A system node group with a hardened launch template

resource "aws_launch_template" "system" {
  name_prefix = "${var.cluster_name}-system-"
  metadata_options {
    http_tokens                 = "required"   # IMDSv2 only
    http_put_response_hop_limit = 1            # pods can't reach the node role
  }
  block_device_mappings {
    device_name = "/dev/xvda"
    ebs {
      volume_size = 50
      volume_type = "gp3"
      encrypted   = true
    }
  }
}

resource "aws_eks_node_group" "system" {
  cluster_name    = aws_eks_cluster.this.name
  node_group_name = "system"
  node_role_arn   = aws_iam_role.node.arn
  subnet_ids      = local.private_subnet_ids
  ami_type        = "AL2023_x86_64_STANDARD"
  instance_types  = ["m6i.large"]

  scaling_config {
    min_size     = 3
    max_size     = 6
    desired_size = 3
  }
  update_config { max_unavailable = 1 }

  launch_template {
    id      = aws_launch_template.system.id
    version = aws_launch_template.system.latest_version
  }

  labels = { role = "system" }
  taint {
    key    = "CriticalAddonsOnly"
    value  = "true"
    effect = "NO_SCHEDULE"
  }

  lifecycle {
    ignore_changes = [scaling_config[0].desired_size]   # the autoscaler owns it
  }
}

The node role gets AmazonEKSWorkerNodePolicy, an ECR pull policy and AmazonSSMManagedInstanceCore; the CNI's policy goes on its own Pod Identity role (lesson 06), not here. Workload capacity comes from Karpenter (lesson 07).

Managed add-ons, pinned

locals {
  addons = {
    vpc-cni                = "v1.20.0-eksbuild.1"   # example versions: pick valid ones for your
    coredns                = "v1.12.1-eksbuild.2"   # Kubernetes version with
    kube-proxy             = "v1.33.0-eksbuild.2"   # aws eks describe-addon-versions
    eks-pod-identity-agent = "v1.3.7-eksbuild.2"
  }
}

resource "aws_eks_addon" "this" {
  for_each                    = local.addons
  cluster_name                = aws_eks_cluster.this.name
  addon_name                  = each.key
  addon_version               = each.value
  resolve_conflicts_on_update = "OVERWRITE"
  depends_on                  = [aws_eks_node_group.system]   # CoreDNS needs nodes to become healthy
}

Pinning makes add-on upgrades explicit pull requests. If you prefer "latest compatible", use data "aws_eks_addon_version" with most_recent = true, but then every plan may carry add-on upgrades you didn't ask for.

The same with the community module

terraform-aws-modules/eks/aws wraps all of the above (and the IAM roles, security groups and KMS key) behind inputs:

module "eks" {
  source  = "terraform-aws-modules/eks/aws"
  version = "~> 21.0"          # input names changed between majors: read the upgrade guide for yours

  name               = var.cluster_name
  kubernetes_version = var.kubernetes_version
  vpc_id             = local.vpc_id
  subnet_ids         = local.private_subnet_ids

  endpoint_public_access                   = false
  enable_cluster_creator_admin_permissions = false
  authentication_mode                      = "API"

  addons = {
    vpc-cni                = { before_compute = true }
    coredns                = {}
    kube-proxy             = {}
    eks-pod-identity-agent = { before_compute = true }
  }

  eks_managed_node_groups = {
    system = {
      ami_type       = "AL2023_x86_64_STANDARD"
      instance_types = ["m6i.large"]
      min_size       = 3
      max_size       = 6
      desired_size   = 3
      labels         = { role = "system" }
    }
  }

  access_entries = {
    platform_admins = {
      principal_arn = var.admin_role_arns[0]
      policy_associations = {
        admin = {
          policy_arn   = "arn:aws:eks::aws:cluster-access-policy/AmazonEKSClusterAdminPolicy"
          access_scope = { type = "cluster" }
        }
      }
    }
  }
}

Use the module for speed and community-tested defaults; use raw resources (or your own thin module) when you need full control or want fewer moving parts. Either way, pin the version and upgrade it in its own pull request with a careful plan review: major module versions have renamed inputs and moved resources.

Attributes that replace the cluster

Some changes can't be made in place and would destroy and recreate the cluster: for example the cluster name, the cluster IAM role_arn, and certain network and encryption settings (which ones has changed over time as EKS added in-place updates). Protect against surprises:

  • lifecycle { prevent_destroy = true } on the cluster.
  • A pipeline step that fails on must be replaced for aws_eks_cluster (cheat sheet).
  • Read the provider's resource documentation for "forces new resource" before touching an attribute.

Try it: build the cluster layer (sandbox)

  1. Using the network layer from lesson 04, apply the raw-resource cluster, node group and add-ons (pick valid add-on versions first).
  2. aws eks update-kubeconfig with an admin role from admin_role_arns, and confirm kubectl get nodes shows the three system nodes with the taint.
  3. Change name in a branch and run plan only: note the replacement warning, then revert.
  4. Bump one add-on version and apply; watch aws eks describe-addon move through UPDATING to ACTIVE.

Recap

  • Cluster as code: auth mode API, creator admin off, private endpoint, KMS, logs with retention, upgrade policy explicit, prevent_destroy.
  • Admins as access entries in code; a system node group with IMDSv2 hop limit 1, encrypted disks and a taint; ignore_changes on desired size.
  • Pin add-on versions; bump them in reviewed changes.
  • The community module packages the same; pin it and read upgrade guides. Watch for replacement attributes.

This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.