Part 3 — Operate as code · wrap-up
Cheat sheet & self-check
25 questions across 9 lessons. Each answer links back to the lesson it came from.
Pick an answer to see if you got it, and why.
Q1. Why use OIDC federation from CI instead of storing AWS access keys as secrets?
Show answer
B. The CI provider issues a signed token per job; AWS exchanges it for temporary credentials if the trust conditions match.
From lesson 08 · CI/CD for the stackQ2. Why separate the plan role from the apply role?
Show answer
B. A malicious or buggy PR shouldn't be able to modify production just by running a plan job.
From lesson 08 · CI/CD for the stackQ3. What should apply use?
Show answer
B. Applying a saved plan prevents surprises when something changed between review and apply.
From lesson 08 · CI/CD for the stackQ4. Why split an upgrade into several pull requests (control plane, add-ons, nodes)?
Show answer
B. One giant apply mixes concerns and makes a failure hard to reason about. Small steps with checks between them are safer and easier to roll forward.
From lesson 09 · Upgrades with TerraformQ5. The control-plane version bump shows 'must be replaced' in the plan. What do you do?
Show answer
B. A Kubernetes version bump is an in-place update. Replacement means an unrelated change slipped into the plan.
From lesson 09 · Upgrades with TerraformQ6. How do Karpenter-managed nodes pick up a new AMI after the upgrade?
Show answer
B. Drift detection rolls Karpenter nodes automatically, respecting PDBs and NodePool disruption budgets.
From lesson 09 · Upgrades with TerraformQ7. You moved a node group resource into a module. How do you avoid Terraform destroying and recreating it?
Show answer
B. Addresses changed, not the resource. A moved block tells Terraform it's the same object, and it's reviewable in code.
From lesson 10 · Drift, import, moved & safe destroyQ8. Someone scaled a node group's max_size in the console. What does a scheduled drift check do?
Show answer
B. Drift detection surfaces out-of-band changes so they're decided on deliberately instead of being overwritten or forgotten.
From lesson 10 · Drift, import, moved & safe destroyQ9. terraform destroy of the network layer fails with 'DependencyViolation' on a subnet. Most likely cause?
Show answer
B. In-cluster controllers create AWS resources Terraform doesn't track. Remove them before destroying the cluster and network layers.
From lesson 10 · Drift, import, moved & safe destroyQ10. Why apply the stack layer by layer with verification gates?
Show answer
B. A broken VPC discovered after the platform layer is installed is much harder to untangle.
From lesson 11 · Playbook: creating the clusterQ11. What must you check in the plan before applying the cluster layer for the first time?
Show answer
B. Plan review is the cheapest place to catch mistakes.
From lesson 11 · Playbook: creating the clusterQ12. The platform layer fails halfway (a Helm release times out). What's the safe response?
Show answer
B. Layered states make a failed layer re-runnable without touching the others.
From lesson 11 · Playbook: creating the clusterQ13. What's the correct order for a Kubernetes minor upgrade on this stack?
Show answer
B. The control plane may be at most a limited number of versions ahead of nodes; never behind.
From lesson 12 · Playbook: day-2 operationsQ14. Why delete LoadBalancer Services/Ingresses and Karpenter NodePools before `terraform destroy`?
Show answer
B. Let the creators clean up their own resources first.
From lesson 12 · Playbook: day-2 operationsQ15. How should a team get access to the cluster?
Show answer
B. Reviewed, auditable, least-privilege access through the EKS API.
From lesson 12 · Playbook: day-2 operationsQ16. A plan shows the EKS cluster 'must be replaced'. What do you do?
Show answer
B. Replacing a cluster deletes it. Plans are where you catch this.
From lesson 13 · Challenges on this stackQ17. A pod is Pending with a volume node affinity conflict after rescheduling. Why?
Show answer
B. Use WaitForFirstConsumer, keep capacity in every AZ used by volumes, or use EFS for shared/multi-AZ needs.
From lesson 13 · Challenges on this stackQ18. Karpenter logs say no instance types satisfy the requirements. What's a likely cause?
Show answer
B. Read the NodeClaim events and Karpenter logs; widen requirements or fix selector tags.
From lesson 13 · Challenges on this stackQ19. When is `terraform force-unlock` safe?
Show answer
B. Check the lock info (who, when, operation) and your pipelines first.
From lesson 14 · Recovery playbookQ20. The state file was overwritten by a bad apply. What's the recovery path with a versioned S3 bucket?
Show answer
B. Versioning is the reason you enabled it in lesson 02. Verify with a plan afterwards.
From lesson 14 · Recovery playbookQ21. `terraform destroy` of the network layer hangs deleting a subnet. What's the usual cause?
Show answer
B. Always clean up controller-created resources before destroying lower layers.
From lesson 14 · Recovery playbookQ22. What can `terraform test` with a mock AWS provider verify?
Show answer
B. It's unit testing for Terraform code, fast and free, not an integration test.
From lesson 15 · Simulator: practise for $0Q23. Which part of this course can't be simulated faithfully offline?
Show answer
B. Use a short, budgeted real session for these, and destroy afterwards.
From lesson 15 · Simulator: practise for $0Q24. Why rehearse the platform layer on kind?
Show answer
B. Swap ALB for an ingress controller, EBS CSI for local-path, Karpenter for fixed nodes.
From lesson 15 · Simulator: practise for $0Q25. What makes a platform 'hand-over ready'?
Show answer
B. Operability is part of the deliverable. If only the builder can run it, it isn't finished.
From lesson 16 · Capstone: the whole platform from an empty account