Lesson 10 of 17 · Part 3 — Operate as code
Drift, import, moved & safe destroy
Keep code and reality in step: detect drift on a schedule, bring existing EKS resources under Terraform with import blocks, refactor addresses with moved and removed blocks without replacing anything, and tear an EKS environment down without leaving orphans behind.
Drift: find it before it finds you
Drift is any difference between code and reality: a console change, a manual hotfix, a resource deleted by hand. Run a scheduled job per layer:
$ terraform plan -detailed-exitcode -lock=false -input=false
$ echo $?
2 # changes: alert the owning team with the plan attached
When drift appears, decide: revert it (apply the code) or adopt it (change the code, or apply -refresh-only for attributes you don't manage). Never leave it; the next unrelated apply would silently undo it.
Drift is someone rearranging your kitchen while you sleep. A drift check is looking around every morning before you cook, so you don't pour tea into the sugar bowl.
Import: adopt what already exists
A cluster or node group was created by hand or by another tool, and you want Terraform to own it. Use import blocks (Terraform 1.5+), which are reviewable and repeatable:
import {
to = aws_eks_cluster.this
id = "prod"
}
import {
to = aws_eks_node_group.system
id = "prod:system" # cluster_name:node_group_name
}
import {
to = aws_eks_access_entry.admins["arn:aws:iam::111122223333:role/PlatformAdmin"]
id = "prod:arn:aws:iam::111122223333:role/PlatformAdmin"
}
$ terraform plan -generate-config-out=generated.tf # drafts HCL for anything without config
Tidy the generated code into your real modules, run plan until it shows only the imports and no changes, then apply. Import IDs differ per resource type; each resource's documentation has an "Import" section.
Refactor without replacement: moved and removed
Renaming a resource, adding for_each, or moving resources into a module changes their addresses. Without help, Terraform plans a destroy of the old address and a create of the new one, which for a cluster is a disaster.
# Moved the raw node group into the cluster module
moved {
from = aws_eks_node_group.system
to = module.cluster.aws_eks_node_group.system
}
# (alternative) if you had instead switched to for_each over a map of groups
moved {
from = aws_eks_node_group.system
to = aws_eks_node_group.this["system"]
}
To stop managing something without destroying it (for example handing it to another layer or tool), use a removed block (Terraform 1.7+):
removed {
from = helm_release.aws_load_balancer_controller # now owned by Argo CD
lifecycle { destroy = false }
}
Keep moved/removed blocks in the code for a while (until every environment has applied them), then delete them.
Protect what must never be destroyed
resource "aws_eks_cluster" "this" {
# ...
lifecycle { prevent_destroy = true }
}
Add it to the cluster, the KMS key, the state bucket and stateful data stores. A plan that would destroy them fails loudly; removing the protection becomes its own deliberate, reviewed change.
Destroying an environment safely
Controllers inside the cluster create AWS resources outside Terraform's knowledge. Tear down in this order:
- Workloads off: disable the Argo CD applications (or delete namespaces), so
LoadBalancerServices, Ingresses and PVCs are removed and their ALBs, NLBs and EBS volumes deleted by their controllers. - Karpenter nodes off: delete NodePools, wait until its nodes are gone.
- Check for orphans with the cheat-sheet commands: no load balancers, volumes or unexpected ENIs left.
- Destroy layers bottom-up:
30-platform→20-cluster(after liftingprevent_destroyin a reviewed change) →10-network. - Verify: no NAT gateways, Elastic IPs or endpoints left billing; state files archived.
Orphans cost money and block deletes
An ALB left by a forgotten Ingress keeps billing and holds ENIs in the public subnets, so the VPC destroy fails with DependencyViolation. Delete it through Kubernetes (so the controller cleans up) or by hand, then retry.
Try it: import, refactor, destroy (sandbox)
- Create a node group in the console, then import it with an import block and
-generate-config-out; get to a clean plan. - Move it into a
for_eachmap with amovedblock; confirm the plan shows no destroy. - Change
max_sizein the console and run the drift check; decide to revert it. - Deploy an app with an ALB Ingress, then try destroying the cluster and network layers without deleting it first; note the failure, clean up properly and finish the teardown.
Recap
- Run drift checks on a schedule (
plan -detailed-exitcode); revert or adopt, never ignore. - Import blocks +
-generate-config-outto adopt existing resources, ending in a no-change plan. movedfor renames and restructuring,removedto stop managing without destroying.prevent_destroyon the cluster and critical data; destroy environments workloads → Karpenter → orphans → layers bottom-up.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.