Lesson 05 of 17 · Part 2 — Build as code
The EKS cluster as code
Build the cluster layer in Terraform: the EKS cluster with a private endpoint, API authentication mode, KMS and logging; a managed node group with a hardened launch template; and managed add-ons, first with raw resources so you see every setting, then with the community module.
Raw resources first
The community module is convenient, but writing the cluster with plain aws_* resources once shows you every decision. (Arguments below match current AWS provider versions; check the provider documentation for the version you pin.)
Before using a food processor, chop an onion by hand once. You'll understand what the machine is doing, and you'll know what to check when it jams.
The cluster
resource "aws_eks_cluster" "this" {
name = var.cluster_name
version = var.kubernetes_version
role_arn = aws_iam_role.cluster.arn
access_config {
authentication_mode = "API"
bootstrap_cluster_creator_admin_permissions = false # admins are declared below
}
vpc_config {
subnet_ids = local.private_subnet_ids
endpoint_private_access = true
endpoint_public_access = var.public_endpoint # false in prod
public_access_cidrs = var.public_endpoint ? var.admin_cidrs : null
}
encryption_config {
resources = ["secrets"]
provider { key_arn = aws_kms_key.eks.arn }
}
enabled_cluster_log_types = ["api", "audit", "authenticator"]
upgrade_policy {
support_type = "STANDARD" # don't drift into paid extended support silently
}
depends_on = [aws_cloudwatch_log_group.cluster]
lifecycle {
prevent_destroy = true
}
}
resource "aws_cloudwatch_log_group" "cluster" {
name = "/aws/eks/${var.cluster_name}/cluster" # create it first to control retention
retention_in_days = 90
}
The cluster role:
resource "aws_iam_role" "cluster" {
name = "${var.cluster_name}-cluster"
assume_role_policy = jsonencode({
Version = "2012-10-17"
Statement = [{ Effect = "Allow", Principal = { Service = "eks.amazonaws.com" }, Action = ["sts:AssumeRole", "sts:TagSession"] }]
})
}
resource "aws_iam_role_policy_attachment" "cluster" {
role = aws_iam_role.cluster.name
policy_arn = "arn:aws:iam::aws:policy/AmazonEKSClusterPolicy"
}
Admin access as code
resource "aws_eks_access_entry" "admins" {
for_each = toset(var.admin_role_arns)
cluster_name = aws_eks_cluster.this.name
principal_arn = each.value
}
resource "aws_eks_access_policy_association" "admins" {
for_each = toset(var.admin_role_arns)
cluster_name = aws_eks_cluster.this.name
principal_arn = each.value
policy_arn = "arn:aws:eks::aws:cluster-access-policy/AmazonEKSClusterAdminPolicy"
access_scope { type = "cluster" }
depends_on = [aws_eks_access_entry.admins]
}
Team access (namespace-scoped) follows the same pattern and is covered in lesson 06.
A system node group with a hardened launch template
resource "aws_launch_template" "system" {
name_prefix = "${var.cluster_name}-system-"
metadata_options {
http_tokens = "required" # IMDSv2 only
http_put_response_hop_limit = 1 # pods can't reach the node role
}
block_device_mappings {
device_name = "/dev/xvda"
ebs {
volume_size = 50
volume_type = "gp3"
encrypted = true
}
}
}
resource "aws_eks_node_group" "system" {
cluster_name = aws_eks_cluster.this.name
node_group_name = "system"
node_role_arn = aws_iam_role.node.arn
subnet_ids = local.private_subnet_ids
ami_type = "AL2023_x86_64_STANDARD"
instance_types = ["m6i.large"]
scaling_config {
min_size = 3
max_size = 6
desired_size = 3
}
update_config { max_unavailable = 1 }
launch_template {
id = aws_launch_template.system.id
version = aws_launch_template.system.latest_version
}
labels = { role = "system" }
taint {
key = "CriticalAddonsOnly"
value = "true"
effect = "NO_SCHEDULE"
}
lifecycle {
ignore_changes = [scaling_config[0].desired_size] # the autoscaler owns it
}
}
The node role gets AmazonEKSWorkerNodePolicy, an ECR pull policy and AmazonSSMManagedInstanceCore; the CNI's policy goes on its own Pod Identity role (lesson 06), not here. Workload capacity comes from Karpenter (lesson 07).
Managed add-ons, pinned
locals {
addons = {
vpc-cni = "v1.20.0-eksbuild.1" # example versions: pick valid ones for your
coredns = "v1.12.1-eksbuild.2" # Kubernetes version with
kube-proxy = "v1.33.0-eksbuild.2" # aws eks describe-addon-versions
eks-pod-identity-agent = "v1.3.7-eksbuild.2"
}
}
resource "aws_eks_addon" "this" {
for_each = local.addons
cluster_name = aws_eks_cluster.this.name
addon_name = each.key
addon_version = each.value
resolve_conflicts_on_update = "OVERWRITE"
depends_on = [aws_eks_node_group.system] # CoreDNS needs nodes to become healthy
}
Pinning makes add-on upgrades explicit pull requests. If you prefer "latest compatible", use data "aws_eks_addon_version" with most_recent = true, but then every plan may carry add-on upgrades you didn't ask for.
The same with the community module
terraform-aws-modules/eks/aws wraps all of the above (and the IAM roles, security groups and KMS key) behind inputs:
module "eks" {
source = "terraform-aws-modules/eks/aws"
version = "~> 21.0" # input names changed between majors: read the upgrade guide for yours
name = var.cluster_name
kubernetes_version = var.kubernetes_version
vpc_id = local.vpc_id
subnet_ids = local.private_subnet_ids
endpoint_public_access = false
enable_cluster_creator_admin_permissions = false
authentication_mode = "API"
addons = {
vpc-cni = { before_compute = true }
coredns = {}
kube-proxy = {}
eks-pod-identity-agent = { before_compute = true }
}
eks_managed_node_groups = {
system = {
ami_type = "AL2023_x86_64_STANDARD"
instance_types = ["m6i.large"]
min_size = 3
max_size = 6
desired_size = 3
labels = { role = "system" }
}
}
access_entries = {
platform_admins = {
principal_arn = var.admin_role_arns[0]
policy_associations = {
admin = {
policy_arn = "arn:aws:eks::aws:cluster-access-policy/AmazonEKSClusterAdminPolicy"
access_scope = { type = "cluster" }
}
}
}
}
}
Use the module for speed and community-tested defaults; use raw resources (or your own thin module) when you need full control or want fewer moving parts. Either way, pin the version and upgrade it in its own pull request with a careful plan review: major module versions have renamed inputs and moved resources.
Attributes that replace the cluster
Some changes can't be made in place and would destroy and recreate the cluster: for example the cluster name, the cluster IAM role_arn, and certain network and encryption settings (which ones has changed over time as EKS added in-place updates). Protect against surprises:
lifecycle { prevent_destroy = true }on the cluster.- A pipeline step that fails on
must be replacedforaws_eks_cluster(cheat sheet). - Read the provider's resource documentation for "forces new resource" before touching an attribute.
Try it: build the cluster layer (sandbox)
- Using the network layer from lesson 04, apply the raw-resource cluster, node group and add-ons (pick valid add-on versions first).
aws eks update-kubeconfigwith an admin role fromadmin_role_arns, and confirmkubectl get nodesshows the three system nodes with the taint.- Change
namein a branch and runplanonly: note the replacement warning, then revert. - Bump one add-on version and apply; watch
aws eks describe-addonmove throughUPDATINGtoACTIVE.
Recap
- Cluster as code: auth mode
API, creator admin off, private endpoint, KMS, logs with retention, upgrade policy explicit,prevent_destroy. - Admins as access entries in code; a system node group with IMDSv2 hop limit 1, encrypted disks and a taint;
ignore_changeson desired size. - Pin add-on versions; bump them in reviewed changes.
- The community module packages the same; pin it and read upgrade guides. Watch for replacement attributes.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.