Manage EKS with Terraform›Part 3 · Cheat sheet & self-check
Learning Hub / Cloud — OpenStack, AWS & EKS / Manage EKS with Terraform

Part 3 — Operate as code · wrap-up

Cheat sheet & self-check

Every command from this section on one page.

08 · CI/CD for the stack

Pipeline stages

terraform fmt -check && terraform validateFormatting and syntax
tflintTerraform linter (provider-aware rules)
checkov -d . / trivy config .Security misconfiguration scanning
terraform plan -out=tfplanPlan (posted to the PR)
terraform apply tfplanApply the exact reviewed plan

GitHub OIDC to AWS

permissions: id-token: writeAllow the job to request an OIDC token
uses: aws-actions/configure-aws-credentials@v4 with role-to-assumeExchange it for temporary AWS credentials

09 · Upgrades with Terraform

Before

aws eks list-insights --cluster-name prod --filter kubernetesVersions=1.34Upgrade blockers EKS found (example target version)
kubentDeprecated/removed APIs in use (kube-no-trouble)
aws eks describe-addon-versions --kubernetes-version 1.34 --addon-name vpc-cni --query 'addons[].addonVersions[0].addonVersion'Add-on versions for the target

During

terraform -chdir=live/prod/20-cluster plan | grep -E '~ version|must be replaced'The plan should show in-place version changes only
aws eks describe-update --name prod --update-id <id>Progress of a control-plane update
kubectl get nodes -L node.kubernetes.io/instance-type -o wideNode versions rolling
kubectl get nodeclaims -o wideKarpenter replacing drifted nodes

10 · Drift, import, moved & safe destroy

Drift

terraform plan -detailed-exitcodeExit 0 = no changes, 2 = changes (drift or pending code), 1 = error
terraform plan -refresh-onlyShow only what changed outside Terraform
terraform apply -refresh-onlyAccept reality into state (no infrastructure changes)

Import and refactor

terraform plan -generate-config-out=generated.tfWrite HCL for resources declared in import blocks
terraform state listAddresses currently in state
terraform state mv aws_eks_node_group.ng module.cluster.aws_eks_node_group.systemMove an address by hand (prefer moved blocks)

Destroy

aws elbv2 describe-load-balancers --query 'LoadBalancers[?VpcId==`<vpc>`].LoadBalancerName'Load balancers still in the VPC
aws ec2 describe-volumes --filters Name=tag-key,Values=kubernetes.io/cluster/prod --query 'Volumes[].VolumeId'EBS volumes created for the cluster
aws ec2 describe-network-interfaces --filters Name=vpc-id,Values=<vpc> --query 'NetworkInterfaces[].[NetworkInterfaceId,Description]'ENIs that block subnet deletion

11 · Playbook: creating the cluster

Pre-flight

aws sts get-caller-identityRight account and role?
aws service-quotas get-service-quota --service-code ec2 --quota-code L-1216C47AOn-Demand standard vCPU quota
terraform version / kubectl version --client / helm versionTool versions match the repo's pins

Per layer (network → cluster → platform)

terraform init -backend-config=backend-prod.hclPoint at this layer's state
terraform plan -var-file=prod.tfvars -out=prod.planReview: only expected creates
terraform apply prod.planApply exactly what was reviewed

Verification gates

aws eks describe-cluster --name prod --query 'cluster.status'ACTIVE
aws eks update-kubeconfig --name prod --region eu-west-1 --alias prodkubectl access
aws eks list-addons --cluster-name prod / describe-addon …Add-ons ACTIVE
kubectl get nodes -o wide / kubectl get pods -ANodes Ready, system pods Running

12 · Playbook: day-2 operations

Everyday changes (through Git + pipeline)

NodePool / node group change in the platform or cluster layerAdd or reshape capacity
aws_eks_access_entry + aws_eks_access_policy_associationGrant a team access
aws eks describe-addon-versions --addon-name vpc-cni --kubernetes-version 1.32Compatible add-on versions

Upgrade

pluto detect-helm -o wide / pluto detect-files -d manifests/Deprecated APIs before upgrading
kubernetes_version = "1.33" → plan/apply (cluster layer)Control plane first
kubectl get nodes -L karpenter.sh/nodepool -o wideWatch nodes roll to the new version

Teardown (reverse order)

kubectl delete ingress,svc -A -l <app selector> (LB-backed)Let controllers delete ALBs/NLBs
kubectl delete nodepools --allLet Karpenter terminate its nodes
terraform destroy (platform → cluster → network)Then Terraform

13 · Challenges on this stack

Quick checks

terraform plan | grep -E 'must be replaced|forces replacement'Dangerous replacements
terraform plan -detailed-exitcode (2 = changes)Scheduled drift detection
kubectl get nodeclaims / kubectl -n kube-system logs deploy/karpenterWhy no new nodes?
kubectl describe pvc <p> / kubectl get pv -o wideVolume AZ and binding
kubectl -n kube-system logs deploy/aws-load-balancer-controller | tailWhy no ALB?

Guard rails

lifecycle { prevent_destroy = true }On the cluster and state-critical resources
volumeBindingMode: WaitForFirstConsumerEBS volumes created in the pod's AZ
VPC endpoints: ecr.api, ecr.dkr, s3 (gateway), stsCut NAT traffic

14 · Recovery playbook

State

terraform force-unlock <LOCK_ID>Release a stale lock (after confirming nothing is running)
aws s3api list-object-versions --bucket <b> --prefix eks/prod/cluster/terraform.tfstateFind earlier state versions
terraform state pull > before-repair.tfstateBackup before any repair
import { to = … id = … } then plan/applyAdopt existing resources

Network & cluster

aws ec2 describe-network-interfaces --filters Name=vpc-id,Values=<vpc> --query 'NetworkInterfaces[].[NetworkInterfaceId,Description,Status]'Leftover ENIs blocking destroy
aws eks list-access-entries --cluster-name prodWho can get in
velero restore create --from-backup <name>Restore app data/objects to a rebuilt cluster

15 · Simulator: practise for $0

Terraform without AWS

terraform fmt -check && terraform validateSyntax and references
mock_provider "aws" {} in tests/*.tftest.hclFake provider for tests (Terraform 1.7+)
terraform testRun plan-based tests with assertions

Local stand-ins

docker run -d -p 4566:4566 localstack/localstackLocalStack (check feature coverage for your version)
backend "s3" { endpoints = { s3 = "http://localhost:4566" } use_path_style = true … }State backend on LocalStack
kind create cluster --name eks-simPractise the platform layer
AWS Budgets alert + destroy after each sessionWhen you do use real AWS

16 · Capstone: the whole platform from an empty account

Readiness checks

aws eks describe-cluster --name prod --query 'cluster.resourcesVpcConfig'Endpoint access and networking
aws eks list-access-entries --cluster-name prodWho can access the cluster
aws eks list-insights --cluster-name prodUpgrade readiness
kubectl get nodepools,nodeclaimsKarpenter capacity
velero backup getBackups and their status