Part 3 — Operate · wrap-up
Cheat sheet & self-check
Every command from this section on one page.
Control plane
aws eks update-cluster-config --name prod --logging '{"clusterLogging":[{"types":["api","audit","authenticator"],"enabled":true}]}' | Send control-plane logs to CloudWatch Logs |
aws logs tail /aws/eks/prod/cluster --follow --filter-pattern authenticator | Follow control-plane logs |
kubectl get --raw /metrics | grep apiserver_request_total | head | API server metrics (Prometheus format) |
In the cluster
aws eks create-addon --cluster-name prod --addon-name amazon-cloudwatch-observability | Container Insights + Fluent Bit via managed add-on |
helm install kps prometheus-community/kube-prometheus-stack -n monitoring --create-namespace | Self-run Prometheus, Alertmanager, Grafana |
kubectl top pods -A --sort-by=memory | head | Quick resource view (metrics-server) |
Check the posture
aws eks describe-cluster --name prod --query 'cluster.[resourcesVpcConfig.endpointPublicAccess,resourcesVpcConfig.publicAccessCidrs,encryptionConfig]' | Public endpoint, allowed CIDRs, secrets encryption |
aws ec2 describe-launch-templates --query 'LaunchTemplates[].LaunchTemplateName' | Find node launch templates to check IMDS settings |
kubectl get ns -L pod-security.kubernetes.io/enforce | Pod Security level per namespace |
kubectl get networkpolicy -A | Which namespaces restrict traffic |
aws guardduty list-findings --detector-id <id> --finding-criteria '{"Criterion":{"resource.resourceType":{"Eq":["EKSCluster"]}}}' | GuardDuty findings for EKS |
Harden
kubectl label ns shop pod-security.kubernetes.io/enforce=restricted --overwrite | Enforce restricted pods |
aws eks update-addon --cluster-name prod --addon-name vpc-cni --configuration-values '{"enableNetworkPolicy":"true"}' | Enable network policy in the VPC CNI |
trivy image <acct>.dkr.ecr.<region>.amazonaws.com/shop/api:1.4.0 | Scan an image before deploying |
Before
aws eks list-insights --cluster-name prod | Upgrade insights: deprecated APIs, add-on compatibility and more |
aws eks describe-addon-versions --addon-name vpc-cni --kubernetes-version 1.33 --query 'addons[].addonVersions[0].addonVersion' | Latest add-on version for the target Kubernetes version |
kubectl get pdb -A | Budgets that could block node replacement |
Do
aws eks update-cluster-version --name prod --kubernetes-version 1.33 | Upgrade the control plane (or change the version in Terraform) |
aws eks update-addon --cluster-name prod --addon-name coredns --addon-version <v> | Upgrade an add-on |
aws eks update-nodegroup-version --cluster-name prod --nodegroup-name system | Roll a managed node group to the new version |
State recovery
aws s3api list-object-versions --bucket acme-tfstate-prod --prefix eks/prod/cluster/ | Previous versions of a state file |
aws s3api get-object --bucket … --key … --version-id <id> old.tfstate | Download an older version |
terraform state pull > backup.tfstate | Back up current state before any repair |
import { to = aws_s3_bucket.logs id = "acme-logs" } | Re-adopt an existing resource (in code, Terraform 1.5+) |
Prevention
lifecycle { prevent_destroy = true } | Refuse to plan destruction of critical resources |
deletion_protection / termination protection | Provider-level protection where available |
IAM deny on eks:DeleteCluster for CI roles (except break-glass) | Stop destructive calls at the API |
Laptop
aws sso login --profile prod-readonly | Get short-lived credentials |
aws eks update-kubeconfig --name prod --region eu-west-1 --alias prod --profile prod-readonly | Write the kubeconfig entry |
kubectl config view --minify | See the exec plugin (aws eks get-token) |
kubectl auth whoami | Who does the cluster think I am? |
Pipelines
infra: OIDC role → terraform plan (PR) / apply (main, approved) | Infrastructure changes |
app: OIDC role → docker build → push to ECR (by digest) | Images |
GitOps repo commit → Argo CD sync | Deployments |