Manage EKS with Terraform cheat sheet
106 commands from every lesson of Manage EKS with Terraform, on one page.
Working in a layer
terraform -chdir=live/prod/20-cluster init | Initialise one layer of one environment |
terraform -chdir=live/prod/20-cluster plan -out=tfplan | Plan and save it |
terraform -chdir=live/prod/20-cluster apply tfplan | Apply exactly what was reviewed |
terraform providers lock -platform=linux_amd64 -platform=darwin_arm64 | Lock provider hashes for all platforms your team uses |
terraform fmt -recursive && terraform validate | Formatting and static checks |
Across layers
for l in 10-network 20-cluster 30-platform; do terraform -chdir=live/prod/$l apply; done | Create in order |
for l in 30-platform 20-cluster 10-network; do terraform -chdir=live/prod/$l destroy; done | Destroy in reverse order |
aws ssm get-parameter --name /platform/prod/vpc_id | Read an output another layer published |
State
terraform init -backend-config=backend-prod.hcl | Initialise against an environment's backend |
terraform state list | Resources tracked in state |
terraform state show module.eks.aws_eks_cluster.this[0] | One resource's recorded attributes |
terraform force-unlock <lock-id> | Remove a stale lock (only when sure nobody is running!) |
Plans & drift
terraform plan -var-file=prod.tfvars -out=prod.plan | Plan and save it |
terraform apply prod.plan | Apply exactly what was planned |
terraform plan -detailed-exitcode | Exit 0 = no changes, 2 = changes (drift), 1 = error |
Patterns
exec { command = "aws" args = ["eks", "get-token", …] } | Fresh tokens during long applies |
data "terraform_remote_state" "cluster" { … } | Read the cluster layer's outputs from another stack |
terraform apply -target=module.eks | Emergency only: bootstrap one part first |
Useful commands
aws eks get-token --cluster-name prod | jq -r .status.expirationTimestamp | When does this token expire? |
terraform providers | Which providers each module uses |
Network layer
terraform -chdir=live/prod/10-network plan | Review network changes |
terraform -chdir=live/prod/10-network state list | grep subnet | Subnets in state |
terraform -chdir=live/prod/10-network output private_subnets | Private subnet IDs |
aws ec2 describe-subnets --filters Name=tag:karpenter.sh/discovery,Values=prod --query 'Subnets[].SubnetId' | Subnets Karpenter will use |
Cluster layer
terraform -chdir=live/prod/20-cluster plan | grep -E 'must be replaced|forces replacement' | Catch replacements before they happen |
terraform -chdir=live/prod/20-cluster state show aws_eks_cluster.this | Everything Terraform knows about the cluster |
aws eks update-kubeconfig --name prod --alias prod | kubeconfig after apply |
aws eks describe-addon-versions --addon-name coredns --kubernetes-version 1.33 --query 'addons[].addonVersions[].addonVersion' | Valid add-on versions to pin |
Review identity in code and in AWS
terraform -chdir=live/prod/30-platform state list | grep -E 'access_entry|pod_identity|iam_role' | Identity resources Terraform owns |
aws eks list-access-entries --cluster-name prod | Compare with what exists (drift) |
aws eks list-pod-identity-associations --cluster-name prod --query 'associations[].[namespace,serviceAccount]' --output table | Which ServiceAccounts have roles |
aws iam simulate-principal-policy --policy-source-arn <role-arn> --action-names s3:GetObject --resource-arns arn:aws:s3:::bucket/shop/x | Test a role's permissions without calling the service |
Platform layer
terraform -chdir=live/prod/30-platform apply | IAM, queues, Karpenter and Argo CD bootstrap |
helm list -A | Helm releases in the cluster (who installed what) |
kubectl get applications -n argocd | Everything Argo CD owns |
kubectl get secret -n argocd -l argocd.argoproj.io/secret-type=cluster -o yaml | Cluster secret with Terraform-provided annotations |
kubectl get nodepools,ec2nodeclasses | Karpenter configuration (owned by GitOps) |
Pipeline stages
terraform fmt -check && terraform validate | Formatting and syntax |
tflint | Terraform linter (provider-aware rules) |
checkov -d . / trivy config . | Security misconfiguration scanning |
terraform plan -out=tfplan | Plan (posted to the PR) |
terraform apply tfplan | Apply the exact reviewed plan |
GitHub OIDC to AWS
permissions: id-token: write | Allow the job to request an OIDC token |
uses: aws-actions/configure-aws-credentials@v4 with role-to-assume | Exchange it for temporary AWS credentials |
Before
aws eks list-insights --cluster-name prod --filter kubernetesVersions=1.34 | Upgrade blockers EKS found (example target version) |
kubent | Deprecated/removed APIs in use (kube-no-trouble) |
aws eks describe-addon-versions --kubernetes-version 1.34 --addon-name vpc-cni --query 'addons[].addonVersions[0].addonVersion' | Add-on versions for the target |
During
terraform -chdir=live/prod/20-cluster plan | grep -E '~ version|must be replaced' | The plan should show in-place version changes only |
aws eks describe-update --name prod --update-id <id> | Progress of a control-plane update |
kubectl get nodes -L node.kubernetes.io/instance-type -o wide | Node versions rolling |
kubectl get nodeclaims -o wide | Karpenter replacing drifted nodes |
Drift
terraform plan -detailed-exitcode | Exit 0 = no changes, 2 = changes (drift or pending code), 1 = error |
terraform plan -refresh-only | Show only what changed outside Terraform |
terraform apply -refresh-only | Accept reality into state (no infrastructure changes) |
Import and refactor
terraform plan -generate-config-out=generated.tf | Write HCL for resources declared in import blocks |
terraform state list | Addresses currently in state |
terraform state mv aws_eks_node_group.ng module.cluster.aws_eks_node_group.system | Move an address by hand (prefer moved blocks) |
Destroy
aws elbv2 describe-load-balancers --query 'LoadBalancers[?VpcId==`<vpc>`].LoadBalancerName' | Load balancers still in the VPC |
aws ec2 describe-volumes --filters Name=tag-key,Values=kubernetes.io/cluster/prod --query 'Volumes[].VolumeId' | EBS volumes created for the cluster |
aws ec2 describe-network-interfaces --filters Name=vpc-id,Values=<vpc> --query 'NetworkInterfaces[].[NetworkInterfaceId,Description]' | ENIs that block subnet deletion |
Pre-flight
aws sts get-caller-identity | Right account and role? |
aws service-quotas get-service-quota --service-code ec2 --quota-code L-1216C47A | On-Demand standard vCPU quota |
terraform version / kubectl version --client / helm version | Tool versions match the repo's pins |
Per layer (network → cluster → platform)
terraform init -backend-config=backend-prod.hcl | Point at this layer's state |
terraform plan -var-file=prod.tfvars -out=prod.plan | Review: only expected creates |
terraform apply prod.plan | Apply exactly what was reviewed |
Verification gates
aws eks describe-cluster --name prod --query 'cluster.status' | ACTIVE |
aws eks update-kubeconfig --name prod --region eu-west-1 --alias prod | kubectl access |
aws eks list-addons --cluster-name prod / describe-addon … | Add-ons ACTIVE |
kubectl get nodes -o wide / kubectl get pods -A | Nodes Ready, system pods Running |
Everyday changes (through Git + pipeline)
NodePool / node group change in the platform or cluster layer | Add or reshape capacity |
aws_eks_access_entry + aws_eks_access_policy_association | Grant a team access |
aws eks describe-addon-versions --addon-name vpc-cni --kubernetes-version 1.32 | Compatible add-on versions |
Upgrade
pluto detect-helm -o wide / pluto detect-files -d manifests/ | Deprecated APIs before upgrading |
kubernetes_version = "1.33" → plan/apply (cluster layer) | Control plane first |
kubectl get nodes -L karpenter.sh/nodepool -o wide | Watch nodes roll to the new version |
Teardown (reverse order)
kubectl delete ingress,svc -A -l <app selector> (LB-backed) | Let controllers delete ALBs/NLBs |
kubectl delete nodepools --all | Let Karpenter terminate its nodes |
terraform destroy (platform → cluster → network) | Then Terraform |
Quick checks
terraform plan | grep -E 'must be replaced|forces replacement' | Dangerous replacements |
terraform plan -detailed-exitcode (2 = changes) | Scheduled drift detection |
kubectl get nodeclaims / kubectl -n kube-system logs deploy/karpenter | Why no new nodes? |
kubectl describe pvc <p> / kubectl get pv -o wide | Volume AZ and binding |
kubectl -n kube-system logs deploy/aws-load-balancer-controller | tail | Why no ALB? |
Guard rails
lifecycle { prevent_destroy = true } | On the cluster and state-critical resources |
volumeBindingMode: WaitForFirstConsumer | EBS volumes created in the pod's AZ |
VPC endpoints: ecr.api, ecr.dkr, s3 (gateway), sts | Cut NAT traffic |
State
terraform force-unlock <LOCK_ID> | Release a stale lock (after confirming nothing is running) |
aws s3api list-object-versions --bucket <b> --prefix eks/prod/cluster/terraform.tfstate | Find earlier state versions |
terraform state pull > before-repair.tfstate | Backup before any repair |
import { to = … id = … } then plan/apply | Adopt existing resources |
Network & cluster
aws ec2 describe-network-interfaces --filters Name=vpc-id,Values=<vpc> --query 'NetworkInterfaces[].[NetworkInterfaceId,Description,Status]' | Leftover ENIs blocking destroy |
aws eks list-access-entries --cluster-name prod | Who can get in |
velero restore create --from-backup <name> | Restore app data/objects to a rebuilt cluster |
Terraform without AWS
terraform fmt -check && terraform validate | Syntax and references |
mock_provider "aws" {} in tests/*.tftest.hcl | Fake provider for tests (Terraform 1.7+) |
terraform test | Run plan-based tests with assertions |
Local stand-ins
docker run -d -p 4566:4566 localstack/localstack | LocalStack (check feature coverage for your version) |
backend "s3" { endpoints = { s3 = "http://localhost:4566" } use_path_style = true … } | State backend on LocalStack |
kind create cluster --name eks-sim | Practise the platform layer |
AWS Budgets alert + destroy after each session | When you do use real AWS |
Readiness checks
aws eks describe-cluster --name prod --query 'cluster.resourcesVpcConfig' | Endpoint access and networking |
aws eks list-access-entries --cluster-name prod | Who can access the cluster |
aws eks list-insights --cluster-name prod | Upgrade readiness |
kubectl get nodepools,nodeclaims | Karpenter capacity |
velero backup get | Backups and their status |