Lesson 09 of 17 · Part 3 — Operate as code
Upgrades with Terraform
Run EKS upgrades through Terraform as a sequence of small, reviewed pull requests: pre-checks, control plane, add-ons, system nodes, Karpenter nodes, and the module and provider upgrades that must never ride along with them.
Upgrade as a series of pull requests
The platform track (lesson "Upgrades & add-ons") explains the Kubernetes side. With Terraform, the same routine becomes a small series of pull requests, each with its own plan, review and verification:
| PR | Change | Verify before the next |
|---|---|---|
| 0 | Fix deprecated APIs in manifests; check insights; bump controllers that need a newer version first | Insights clear; apps healthy |
| 1 | kubernetes_version = "1.34" (one minor at a time) |
Control plane ACTIVE, API healthy |
| 2 | Add-on versions for 1.34 | All add-ons ACTIVE, DNS and networking fine |
| 3 | System node group version / AMI | Nodes on 1.34, PDBs held |
| 4 | Karpenter AMI selection (if pinned) → drift replaces workload nodes | All nodes on 1.34, SLOs green |
Changing the parts of a moving train one carriage at a time, checking the couplings after each one, instead of swapping the whole train at full speed and hoping it holds together.
PR 1: the control plane
# prod.tfvars
kubernetes_version = "1.34" # was 1.33
$ terraform plan -out=tfplan
# aws_eks_cluster.this will be updated in-place
~ resource "aws_eks_cluster" "this" {
~ version = "1.33" -> "1.34"
}
Plan: 0 to add, 1 to change, 0 to destroy.
That's the only acceptable shape: 1 to change, 0 to destroy. The update takes a while; Terraform waits for it. EKS upgrades one minor version at a time, so jumping two versions means two cycles.
PR 2: add-ons
Bump each pinned version to one listed as compatible with the new Kubernetes version:
locals {
addons = {
vpc-cni = "v1.20.x-eksbuild.y" # choose from describe-addon-versions for 1.34
coredns = "v1.12.x-eksbuild.y"
kube-proxy = "v1.34.x-eksbuild.y" # kube-proxy tracks the Kubernetes minor
eks-pod-identity-agent = "v1.3.x-eksbuild.y"
}
}
Apply, then check aws eks list-addons / describe-addon for ACTIVE, CoreDNS resolution from a test pod, and new pods getting IPs.
PR 3: managed node groups
Managed node groups follow the cluster's version when you set version on them (or let the AMI release follow). Two useful patterns:
resource "aws_eks_node_group" "system" {
# ...
version = var.kubernetes_version # nodes follow the cluster's minor version
release_version = var.system_ami_release # optional: pin the exact AMI release
update_config { max_unavailable = 1 }
}
EKS rolls the group: new nodes launch, old ones are cordoned and drained respecting PDBs. max_unavailable limits how many nodes change at once. If a PDB blocks draining for too long, the update fails rather than breaking the workload; fix the PDB or the workload, then re-apply.
PR 4: Karpenter nodes
If the EC2NodeClass selects AMIs by a moving alias (for example the latest AL2023 AMI), Karpenter sees the new AMI for the new version and drifts nodes automatically. If you pin AMI versions (recommended for production so AMI changes are deliberate), bump the pin in the GitOps repo:
spec:
amiSelectorTerms:
- alias: al2023@v20250715 # example pinned release; bump to one built for the new version
Watch kubectl get nodeclaims and your SLO dashboards; NodePool disruption budgets control the pace.
What must not ride along
Keep these in separate pull requests, never mixed into an upgrade:
- Terraform provider major upgrades.
- Community module major upgrades (renamed inputs, moved resources).
- Network or IAM refactors.
Each can produce replacements or large diffs that hide the one line you meant to change.
Extended support is a setting, not a surprise
With upgrade_policy { support_type = "STANDARD" }, a cluster doesn't stay on an old version at extended-support prices; AWS upgrades it automatically when standard support ends. Decide per cluster, in code, which behaviour you want, and track version end dates so upgrades happen on your schedule, not AWS's.
Try it: one full upgrade cycle (sandbox)
- Build a sandbox cluster one minor version behind the latest.
- Run the pre-checks from the cheat sheet and fix anything found.
- Raise PRs 1–4 (or four commits) and apply them one at a time, recording how long each step took and what you verified.
- Write a short upgrade report: durations, surprises, and one improvement for next time.
Recap
- Upgrades are a sequence of small PRs: pre-checks → control plane → add-ons → system nodes → Karpenter nodes.
- A version bump plan must be in-place only; any replacement means something else changed.
- Pin add-on and AMI versions and bump them deliberately; let PDBs and disruption budgets pace node rollouts.
- Never combine upgrades with provider/module major upgrades or refactors.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.