Cheat Sheets / Cloud — OpenStack, AWS & EKS

Manage EKS with Terraform cheat sheet

106 commands from every lesson of Manage EKS with Terraform, on one page.

01 · Repo layout & run order

Working in a layer

terraform -chdir=live/prod/20-cluster initInitialise one layer of one environment
terraform -chdir=live/prod/20-cluster plan -out=tfplanPlan and save it
terraform -chdir=live/prod/20-cluster apply tfplanApply exactly what was reviewed
terraform providers lock -platform=linux_amd64 -platform=darwin_arm64Lock provider hashes for all platforms your team uses
terraform fmt -recursive && terraform validateFormatting and static checks

Across layers

for l in 10-network 20-cluster 30-platform; do terraform -chdir=live/prod/$l apply; doneCreate in order
for l in 30-platform 20-cluster 10-network; do terraform -chdir=live/prod/$l destroy; doneDestroy in reverse order
aws ssm get-parameter --name /platform/prod/vpc_idRead an output another layer published

02 · Remote state: S3, locking, env-per-tfvars

State

terraform init -backend-config=backend-prod.hclInitialise against an environment's backend
terraform state listResources tracked in state
terraform state show module.eks.aws_eks_cluster.this[0]One resource's recorded attributes
terraform force-unlock <lock-id>Remove a stale lock (only when sure nobody is running!)

Plans & drift

terraform plan -var-file=prod.tfvars -out=prod.planPlan and save it
terraform apply prod.planApply exactly what was planned
terraform plan -detailed-exitcodeExit 0 = no changes, 2 = changes (drift), 1 = error

03 · The multi-provider chicken-and-egg

Patterns

exec { command = "aws" args = ["eks", "get-token", …] }Fresh tokens during long applies
data "terraform_remote_state" "cluster" { … }Read the cluster layer's outputs from another stack
terraform apply -target=module.eksEmergency only: bootstrap one part first

Useful commands

aws eks get-token --cluster-name prod | jq -r .status.expirationTimestampWhen does this token expire?
terraform providersWhich providers each module uses

04 · The VPC as code

Network layer

terraform -chdir=live/prod/10-network planReview network changes
terraform -chdir=live/prod/10-network state list | grep subnetSubnets in state
terraform -chdir=live/prod/10-network output private_subnetsPrivate subnet IDs
aws ec2 describe-subnets --filters Name=tag:karpenter.sh/discovery,Values=prod --query 'Subnets[].SubnetId'Subnets Karpenter will use

05 · The EKS cluster as code

Cluster layer

terraform -chdir=live/prod/20-cluster plan | grep -E 'must be replaced|forces replacement'Catch replacements before they happen
terraform -chdir=live/prod/20-cluster state show aws_eks_cluster.thisEverything Terraform knows about the cluster
aws eks update-kubeconfig --name prod --alias prodkubeconfig after apply
aws eks describe-addon-versions --addon-name coredns --kubernetes-version 1.33 --query 'addons[].addonVersions[].addonVersion'Valid add-on versions to pin

06 · IAM as code: access entries, Pod Identity & IRSA

Review identity in code and in AWS

terraform -chdir=live/prod/30-platform state list | grep -E 'access_entry|pod_identity|iam_role'Identity resources Terraform owns
aws eks list-access-entries --cluster-name prodCompare with what exists (drift)
aws eks list-pod-identity-associations --cluster-name prod --query 'associations[].[namespace,serviceAccount]' --output tableWhich ServiceAccounts have roles
aws iam simulate-principal-policy --policy-source-arn <role-arn> --action-names s3:GetObject --resource-arns arn:aws:s3:::bucket/shop/xTest a role's permissions without calling the service

07 · Add-ons, Karpenter & the GitOps hand-off

Platform layer

terraform -chdir=live/prod/30-platform applyIAM, queues, Karpenter and Argo CD bootstrap
helm list -AHelm releases in the cluster (who installed what)
kubectl get applications -n argocdEverything Argo CD owns
kubectl get secret -n argocd -l argocd.argoproj.io/secret-type=cluster -o yamlCluster secret with Terraform-provided annotations
kubectl get nodepools,ec2nodeclassesKarpenter configuration (owned by GitOps)

08 · CI/CD for the stack

Pipeline stages

terraform fmt -check && terraform validateFormatting and syntax
tflintTerraform linter (provider-aware rules)
checkov -d . / trivy config .Security misconfiguration scanning
terraform plan -out=tfplanPlan (posted to the PR)
terraform apply tfplanApply the exact reviewed plan

GitHub OIDC to AWS

permissions: id-token: writeAllow the job to request an OIDC token
uses: aws-actions/configure-aws-credentials@v4 with role-to-assumeExchange it for temporary AWS credentials

09 · Upgrades with Terraform

Before

aws eks list-insights --cluster-name prod --filter kubernetesVersions=1.34Upgrade blockers EKS found (example target version)
kubentDeprecated/removed APIs in use (kube-no-trouble)
aws eks describe-addon-versions --kubernetes-version 1.34 --addon-name vpc-cni --query 'addons[].addonVersions[0].addonVersion'Add-on versions for the target

During

terraform -chdir=live/prod/20-cluster plan | grep -E '~ version|must be replaced'The plan should show in-place version changes only
aws eks describe-update --name prod --update-id <id>Progress of a control-plane update
kubectl get nodes -L node.kubernetes.io/instance-type -o wideNode versions rolling
kubectl get nodeclaims -o wideKarpenter replacing drifted nodes

10 · Drift, import, moved & safe destroy

Drift

terraform plan -detailed-exitcodeExit 0 = no changes, 2 = changes (drift or pending code), 1 = error
terraform plan -refresh-onlyShow only what changed outside Terraform
terraform apply -refresh-onlyAccept reality into state (no infrastructure changes)

Import and refactor

terraform plan -generate-config-out=generated.tfWrite HCL for resources declared in import blocks
terraform state listAddresses currently in state
terraform state mv aws_eks_node_group.ng module.cluster.aws_eks_node_group.systemMove an address by hand (prefer moved blocks)

Destroy

aws elbv2 describe-load-balancers --query 'LoadBalancers[?VpcId==`<vpc>`].LoadBalancerName'Load balancers still in the VPC
aws ec2 describe-volumes --filters Name=tag-key,Values=kubernetes.io/cluster/prod --query 'Volumes[].VolumeId'EBS volumes created for the cluster
aws ec2 describe-network-interfaces --filters Name=vpc-id,Values=<vpc> --query 'NetworkInterfaces[].[NetworkInterfaceId,Description]'ENIs that block subnet deletion

11 · Playbook: creating the cluster

Pre-flight

aws sts get-caller-identityRight account and role?
aws service-quotas get-service-quota --service-code ec2 --quota-code L-1216C47AOn-Demand standard vCPU quota
terraform version / kubectl version --client / helm versionTool versions match the repo's pins

Per layer (network → cluster → platform)

terraform init -backend-config=backend-prod.hclPoint at this layer's state
terraform plan -var-file=prod.tfvars -out=prod.planReview: only expected creates
terraform apply prod.planApply exactly what was reviewed

Verification gates

aws eks describe-cluster --name prod --query 'cluster.status'ACTIVE
aws eks update-kubeconfig --name prod --region eu-west-1 --alias prodkubectl access
aws eks list-addons --cluster-name prod / describe-addon …Add-ons ACTIVE
kubectl get nodes -o wide / kubectl get pods -ANodes Ready, system pods Running

12 · Playbook: day-2 operations

Everyday changes (through Git + pipeline)

NodePool / node group change in the platform or cluster layerAdd or reshape capacity
aws_eks_access_entry + aws_eks_access_policy_associationGrant a team access
aws eks describe-addon-versions --addon-name vpc-cni --kubernetes-version 1.32Compatible add-on versions

Upgrade

pluto detect-helm -o wide / pluto detect-files -d manifests/Deprecated APIs before upgrading
kubernetes_version = "1.33" → plan/apply (cluster layer)Control plane first
kubectl get nodes -L karpenter.sh/nodepool -o wideWatch nodes roll to the new version

Teardown (reverse order)

kubectl delete ingress,svc -A -l <app selector> (LB-backed)Let controllers delete ALBs/NLBs
kubectl delete nodepools --allLet Karpenter terminate its nodes
terraform destroy (platform → cluster → network)Then Terraform

13 · Challenges on this stack

Quick checks

terraform plan | grep -E 'must be replaced|forces replacement'Dangerous replacements
terraform plan -detailed-exitcode (2 = changes)Scheduled drift detection
kubectl get nodeclaims / kubectl -n kube-system logs deploy/karpenterWhy no new nodes?
kubectl describe pvc <p> / kubectl get pv -o wideVolume AZ and binding
kubectl -n kube-system logs deploy/aws-load-balancer-controller | tailWhy no ALB?

Guard rails

lifecycle { prevent_destroy = true }On the cluster and state-critical resources
volumeBindingMode: WaitForFirstConsumerEBS volumes created in the pod's AZ
VPC endpoints: ecr.api, ecr.dkr, s3 (gateway), stsCut NAT traffic

14 · Recovery playbook

State

terraform force-unlock <LOCK_ID>Release a stale lock (after confirming nothing is running)
aws s3api list-object-versions --bucket <b> --prefix eks/prod/cluster/terraform.tfstateFind earlier state versions
terraform state pull > before-repair.tfstateBackup before any repair
import { to = … id = … } then plan/applyAdopt existing resources

Network & cluster

aws ec2 describe-network-interfaces --filters Name=vpc-id,Values=<vpc> --query 'NetworkInterfaces[].[NetworkInterfaceId,Description,Status]'Leftover ENIs blocking destroy
aws eks list-access-entries --cluster-name prodWho can get in
velero restore create --from-backup <name>Restore app data/objects to a rebuilt cluster

15 · Simulator: practise for $0

Terraform without AWS

terraform fmt -check && terraform validateSyntax and references
mock_provider "aws" {} in tests/*.tftest.hclFake provider for tests (Terraform 1.7+)
terraform testRun plan-based tests with assertions

Local stand-ins

docker run -d -p 4566:4566 localstack/localstackLocalStack (check feature coverage for your version)
backend "s3" { endpoints = { s3 = "http://localhost:4566" } use_path_style = true … }State backend on LocalStack
kind create cluster --name eks-simPractise the platform layer
AWS Budgets alert + destroy after each sessionWhen you do use real AWS

16 · Capstone: the whole platform from an empty account

Readiness checks

aws eks describe-cluster --name prod --query 'cluster.resourcesVpcConfig'Endpoint access and networking
aws eks list-access-entries --cluster-name prodWho can access the cluster
aws eks list-insights --cluster-name prodUpgrade readiness
kubectl get nodepools,nodeclaimsKarpenter capacity
velero backup getBackups and their status