Manage EKS with Terraform›17 · Q&A: Terraform for EKS at scale
Learning Hub / Cloud — OpenStack, AWS & EKS / Manage EKS with Terraform

Lesson 17 of 17 · Part 4 — Q&A

Q&A: Terraform for EKS at scale

Common questions about running EKS with Terraform at scale (state design, layering, modules, pipelines, upgrades, drift, multi-account and team self-service), with detailed answers and the trade-offs involved.

Architect
Key wordsarchitectureQ&ATerraformstatelayersmodulesCI/CDOIDCdriftimportmovedupgradesmulti-accountpolicy as codeGitOps bridge

How to use these questions

Same method as the platform track: think each one through (requirements → options → trade-offs → decision → failure modes → evidence), then open the answer. They focus on what happens when Terraform runs Kubernetes platforms in production, not just on writing it.

Knowing the recipe isn't the same as running a restaurant kitchen on a Saturday night. These questions are about the Saturday night.

State and structure

Q1. How do you structure Terraform state for 12 EKS clusters across 4 accounts?

State per account × environment × layer: for each cluster, separate states for network, cluster and platform layers (plus bootstrap per account). State buckets live in each account (or a central, tightly controlled state account) with versioning, encryption, public access blocked and S3-native locking. Access to a state is limited to the pipeline role for that account; humans get read-only at most.

Why: small blast radius per apply, fast plans, independent ownership, and no single state whose corruption affects the whole fleet. Cross-layer values via SSM parameters or read-only remote state. Fleet-wide changes are rolled out by the pipeline across states in waves.

Q2. One module per cluster, or one generic module with per-cluster inputs?

One generic, versioned module set (network, cluster, platform) with per-cluster inputs. Clusters that need something special get a feature flag or an extra input, not a fork. Promotion is a version bump per environment: non-prod runs v3.5.0 for a week before prod moves from v3.4.0.

Forks per cluster feel faster at first and become an unpatchable fleet within a year. The only reasonable exception is a genuinely different platform (for example a regulated cluster with a different network model), and even then it shares the lower-level modules.

Q3. Should Kubernetes resources be managed by Terraform?

Mostly no. Terraform owns AWS resources and a minimal bootstrap (Karpenter, Argo CD, the cluster secret carrying AWS values). Argo CD owns in-cluster add-ons and workloads, because it reconciles continuously, shows health per application and doesn't need a plan/apply cycle for every chart bump.

Exceptions: objects that must exist before GitOps can run (namespaces for Argo CD, bootstrap secrets), or where the AWS and Kubernetes sides are tightly coupled and created together. Keep that list short and explicit, in a layer separate from the cluster itself to avoid the provider chicken-and-egg.

Q4. What is the provider chicken-and-egg problem, and how is it solved?

The Kubernetes and Helm providers need the cluster's endpoint and credentials, but the cluster is created in the same run. Providers are configured before resources exist, so plans fail or behave unpredictably (especially on destroy, when the cluster can disappear before the objects inside it).

Solutions: separate layers (cluster in one state, in-cluster resources in another that reads the cluster's outputs), exec-based authentication (aws eks get-token) rather than static tokens, and moving most in-cluster resources to GitOps. With layers, destroy order is explicit: platform first, cluster second.

Pipelines and change control

Q5. What should the CI/CD pipeline for the EKS Terraform repositories look like?

  • On pull request: fmt, validate, tflint, security/policy checks (Checkov or OPA/conftest on the plan JSON), then plan per affected layer and environment, posted to the PR. Guard: fail if any aws_eks_cluster, state bucket or KMS key would be destroyed or replaced.
  • Credentials: OIDC from the CI system to a per-account role; plan roles are read-only, apply roles can change only that account.
  • On merge: apply the saved plan for non-prod automatically, prod behind an approval with the plan shown again; applies run layer by layer in order; one apply per state at a time (locking).
  • After apply: smoke tests (cluster reachable, nodes Ready, add-ons ACTIVE) and a record of what changed.
  • Scheduled: drift detection per state; alerts to owners.

Q6. How do you apply a change across 12 clusters safely?

Waves with gates: sandbox → non-prod clusters → one low-risk prod cluster → remaining prod, with soak time and automatic checks (SLOs, error budgets, add-on status) between waves. The change is a module version bump per environment, so each wave is a small input change. Stop the rollout automatically on regression; roll forward with a fix rather than reverting clusters to old versions where the API doesn't allow it (control-plane versions only move forward).

Q7. A colleague applied from their laptop and now the pipeline's plan wants to undo it. What do you do, and how do you prevent it?

First understand the change: was it an emergency fix? If it was right, codify it in a PR and apply through the pipeline so state and code agree; if not, apply the pipeline's plan to revert. Then prevent it: only pipeline roles can write to state buckets and apply; humans get read-only in prod, plus a documented break-glass path with alerting. Drift detection catches console changes too.

Upgrades, refactors and drift

Q8. How do you upgrade the community EKS module across a major version?

Treat it as a refactor, not an upgrade: read the upgrade guide (renamed inputs, removed features, moved resources); bump the version in a branch for one sandbox cluster; update inputs; add moved blocks where resource addresses changed (major versions often ship guidance for these); iterate until plan shows no destroys or replacements and only expected in-place changes; then roll out environment by environment. Never combine it with a Kubernetes version upgrade or provider upgrade in the same change.

Q9. How do you bring a hand-built production cluster under Terraform?

Write the target code with the modules you'd use for new clusters, then import blocks for each existing resource (cluster, node groups, add-ons, access entries, IAM roles, security groups), using -generate-config-out to draft anything missing. Iterate until the plan shows only imports and no changes; where the existing configuration differs from your standard, decide per attribute whether to accept it (in code) or change it later in a separate, planned change. Apply the imports alone, verify, then start normalising in small steps.

Q10. Terraform keeps showing a diff on the node group's desired size and add-on configuration after every apply. Why, and how do you fix it?

Two actors own the same field. The autoscaler changes desired_size; an add-on's configuration is changed by EKS defaults or by a controller. Fix ownership: ignore_changes on fields another system manages (desired size), set add-on configuration explicitly and choose resolve_conflicts_on_update deliberately, and check for normalisation diffs (JSON ordering; use jsonencode and policy documents). Perpetual diffs train people to stop reading plans, which is the real danger.

Security and multi-account

Q11. How do you protect Terraform state for EKS?

State contains sensitive values (endpoints, sometimes secrets and certificate data). Protect it: encrypted S3 (KMS key with a restrictive key policy), versioning for recovery, public access blocked, bucket policy allowing only the pipeline roles (and break-glass), access logging, S3-native locking, and no secrets written into state where avoidable (generate secrets outside Terraform or use ephemeral values and write-only arguments where the provider supports them). Back up state buckets to another account.

Q12. How do teams self-serve IAM roles, queues and buckets for their apps without the platform team reviewing every PR?

Provide a team module (pod role with Pod Identity association, SQS, S3 prefix, secrets) that bakes in guard-rails: names prefixed by team, a mandatory permissions boundary, tags, encryption. Teams call it from their own repository or folder with their own state and pipeline role, which can only create resources with their prefix. Policy-as-code on the plan (OPA/Sentinel/Checkov) enforces the rules, and SCPs back them up. The platform team reviews the module, not every use of it.

Q13. How do you enforce standards (tags, encryption, no public endpoints) in Terraform code?

Three layers: module defaults (encryption on, endpoint private, tags from default_tags), policy-as-code on the plan in CI (block public endpoints, unencrypted volumes, missing tags, wildcard IAM), and account guard-rails (SCPs and AWS Config rules) that catch anything created outside Terraform. Report violations with the rule and the fix, so teams can self-correct.

Recovery and failure

Q14. The state file for the prod cluster layer was overwritten with an old version. How do you recover?

Stop all pipelines for that state. Restore the correct previous version from S3 versioning (compare timestamps with the last known-good apply), check the lock mechanism's metadata if you use DynamoDB locking (its digest can block you), run plan, and expect no changes if the restore was right. If resources were created after the restored version, re-import them. Then find the root cause (a local apply with an old state, a script) and remove that path.

Q15. A terraform destroy ran against the wrong environment and deleted the platform layer. What now, and what should have stopped it?

Recovery: the cluster layer still stands (layering limited the blast radius). Re-apply the platform layer from code; Argo CD then restores add-ons and workloads from Git; data in managed services is unaffected, and PV data depends on volume reclaim policies and backups (Velero/EBS snapshots). Verify with smoke tests and SLOs.

Prevention: destroy only through a dedicated pipeline job with explicit environment confirmation; prevent_destroy on critical resources; separate credentials per account so a non-prod session can't touch prod; no human write access to prod state.

Recap

  • State per account × environment × layer, protected and pipeline-only.
  • Generic versioned modules, per-cluster inputs, wave rollouts with gates.
  • Terraform owns AWS and bootstrap; GitOps owns in-cluster; layers solve the provider chicken-and-egg.
  • Refactors with moved, imports with import blocks, drift checks, policy-as-code, and guard-rails against destroys.

This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.