Manage EKS with Terraform›10 · Drift, import, moved & safe destroy
Learning Hub / Cloud — OpenStack, AWS & EKS / Manage EKS with Terraform

Lesson 10 of 17 · Part 3 — Operate as code

Drift, import, moved & safe destroy

Keep code and reality in step: detect drift on a schedule, bring existing EKS resources under Terraform with import blocks, refactor addresses with moved and removed blocks without replacing anything, and tear an EKS environment down without leaving orphans behind.

Advanced
Key wordsTerraformdriftdetailed-exitcoderefresh-onlyimport blockgenerate-config-outmoved blockremoved blockterraform state mvprevent_destroysafe destroyorphaned load balancers

Drift: find it before it finds you

Drift is any difference between code and reality: a console change, a manual hotfix, a resource deleted by hand. Run a scheduled job per layer:

$ terraform plan -detailed-exitcode -lock=false -input=false
$ echo $?
2      # changes: alert the owning team with the plan attached

When drift appears, decide: revert it (apply the code) or adopt it (change the code, or apply -refresh-only for attributes you don't manage). Never leave it; the next unrelated apply would silently undo it.

Drift is someone rearranging your kitchen while you sleep. A drift check is looking around every morning before you cook, so you don't pour tea into the sugar bowl.

Import: adopt what already exists

A cluster or node group was created by hand or by another tool, and you want Terraform to own it. Use import blocks (Terraform 1.5+), which are reviewable and repeatable:

import {
  to = aws_eks_cluster.this
  id = "prod"
}

import {
  to = aws_eks_node_group.system
  id = "prod:system"            # cluster_name:node_group_name
}

import {
  to = aws_eks_access_entry.admins["arn:aws:iam::111122223333:role/PlatformAdmin"]
  id = "prod:arn:aws:iam::111122223333:role/PlatformAdmin"
}
$ terraform plan -generate-config-out=generated.tf    # drafts HCL for anything without config

Tidy the generated code into your real modules, run plan until it shows only the imports and no changes, then apply. Import IDs differ per resource type; each resource's documentation has an "Import" section.

Refactor without replacement: moved and removed

Renaming a resource, adding for_each, or moving resources into a module changes their addresses. Without help, Terraform plans a destroy of the old address and a create of the new one, which for a cluster is a disaster.

# Moved the raw node group into the cluster module
moved {
  from = aws_eks_node_group.system
  to   = module.cluster.aws_eks_node_group.system
}

# (alternative) if you had instead switched to for_each over a map of groups
moved {
  from = aws_eks_node_group.system
  to   = aws_eks_node_group.this["system"]
}

To stop managing something without destroying it (for example handing it to another layer or tool), use a removed block (Terraform 1.7+):

removed {
  from = helm_release.aws_load_balancer_controller   # now owned by Argo CD
  lifecycle { destroy = false }
}

Keep moved/removed blocks in the code for a while (until every environment has applied them), then delete them.

Protect what must never be destroyed

resource "aws_eks_cluster" "this" {
  # ...
  lifecycle { prevent_destroy = true }
}

Add it to the cluster, the KMS key, the state bucket and stateful data stores. A plan that would destroy them fails loudly; removing the protection becomes its own deliberate, reviewed change.

Destroying an environment safely

Controllers inside the cluster create AWS resources outside Terraform's knowledge. Tear down in this order:

  1. Workloads off: disable the Argo CD applications (or delete namespaces), so LoadBalancer Services, Ingresses and PVCs are removed and their ALBs, NLBs and EBS volumes deleted by their controllers.
  2. Karpenter nodes off: delete NodePools, wait until its nodes are gone.
  3. Check for orphans with the cheat-sheet commands: no load balancers, volumes or unexpected ENIs left.
  4. Destroy layers bottom-up: 30-platform → 20-cluster (after lifting prevent_destroy in a reviewed change) → 10-network.
  5. Verify: no NAT gateways, Elastic IPs or endpoints left billing; state files archived.

Orphans cost money and block deletes

An ALB left by a forgotten Ingress keeps billing and holds ENIs in the public subnets, so the VPC destroy fails with DependencyViolation. Delete it through Kubernetes (so the controller cleans up) or by hand, then retry.

Try it: import, refactor, destroy (sandbox)

  1. Create a node group in the console, then import it with an import block and -generate-config-out; get to a clean plan.
  2. Move it into a for_each map with a moved block; confirm the plan shows no destroy.
  3. Change max_size in the console and run the drift check; decide to revert it.
  4. Deploy an app with an ALB Ingress, then try destroying the cluster and network layers without deleting it first; note the failure, clean up properly and finish the teardown.

Recap

  • Run drift checks on a schedule (plan -detailed-exitcode); revert or adopt, never ignore.
  • Import blocks + -generate-config-out to adopt existing resources, ending in a no-change plan.
  • moved for renames and restructuring, removed to stop managing without destroying.
  • prevent_destroy on the cluster and critical data; destroy environments workloads → Karpenter → orphans → layers bottom-up.

This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.