Manage EKS with Terraform›07 · Add-ons, Karpenter & the GitOps hand-off
Learning Hub / Cloud — OpenStack, AWS & EKS / Manage EKS with Terraform

Lesson 07 of 17 · Part 2 — Build as code

Add-ons, Karpenter & the GitOps hand-off

Decide what Terraform installs in the cluster and what GitOps owns: Terraform creates the AWS side (IAM, queues) and bootstraps Karpenter and Argo CD; Argo CD installs and upgrades everything else, reading the AWS values Terraform produced.

Advanced → Architect
Key wordsTerraformhelm_releaseKarpenterkarpenter submoduleinterruption queueAWS Load Balancer ControllerArgo CDGitOps bridgeapp of appsApplicationSetcluster secretplatform add-ons
Your laptop aws sso login assume a role aws eks update-kubeconfig exec: aws eks get-token kubectl (read-only in prod) Git + CI (OIDC to AWS) infra repo terraform plan / apply app repo build → ECR (by digest) GitOps repo new tag committed AWS account EKS cluster access entries ECR images Argo CD (in cluster) pulls + syncs humans read, pipelines write; every change goes through Git
Terraform builds AWS and bootstraps the cluster; Argo CD owns what runs inside it.

Draw the line first

Owned by Terraform Owned by GitOps (Argo CD)
IAM roles and Pod Identity associations for controllers Load Balancer Controller, ExternalDNS, cert-manager, External Secrets (Helm)
Karpenter's AWS side: interruption SQS queue, EventBridge rules, node role, access entry Karpenter NodePool and EC2NodeClass objects
EKS managed add-ons (CNI, CoreDNS, kube-proxy, Pod Identity Agent, EBS CSI) Observability stack, policy engine, platform dashboards
Bootstrap Helm releases: Karpenter and Argo CD Every application

Why not install everything with Terraform's Helm provider? Plans get slow and fragile, drift inside the cluster isn't reconciled until the next apply, and every chart upgrade becomes an infrastructure change. Argo CD continuously reconciles and shows the state per application.

Terraform is the builder who pours the foundations, runs the pipes and hands over the keys. Argo CD is the caretaker who lives on site and keeps every room exactly as the plans say, every day. You don't call the builder back to rearrange the furniture.

Karpenter: AWS side in Terraform

The EKS module's karpenter submodule creates the node IAM role, the controller's role and Pod Identity association, the interruption queue and EventBridge rules, and the access entry for Karpenter-launched nodes:

module "karpenter" {
  source  = "terraform-aws-modules/eks/aws//modules/karpenter"
  version = "~> 21.0"                 # same major as the EKS module you use

  cluster_name = var.cluster_name
  node_iam_role_additional_policies = {
    ssm = "arn:aws:iam::aws:policy/AmazonSSMManagedInstanceCore"
  }
}

Then the controller itself, pinned, onto the system node group:

resource "helm_release" "karpenter" {
  name       = "karpenter"
  namespace  = "kube-system"
  repository = "oci://public.ecr.aws/karpenter"
  chart      = "karpenter"
  version    = var.karpenter_version          # pin; upgrade in its own PR

  values = [yamlencode({
    settings = {
      clusterName       = var.cluster_name
      interruptionQueue = module.karpenter.queue_name
    }
    serviceAccount = { name = module.karpenter.service_account }
    tolerations    = [{ key = "CriticalAddonsOnly", operator = "Exists" }]
    nodeSelector   = { role = "system" }
  })]
}

Karpenter's CRDs ship with its chart; upgrading Karpenter across versions sometimes needs CRD steps, so read its upgrade guide every time.

Argo CD bootstrap and the GitOps bridge

Terraform installs Argo CD once, then registers the cluster with annotations carrying AWS values that add-ons need:

resource "helm_release" "argocd" {
  name             = "argocd"
  namespace        = "argocd"
  create_namespace = true
  repository       = "https://argoproj.github.io/argo-helm"
  chart            = "argo-cd"
  version          = var.argocd_chart_version
}

resource "kubernetes_secret" "in_cluster" {
  metadata {
    name      = "in-cluster"
    namespace = "argocd"
    labels = {
      "argocd.argoproj.io/secret-type" = "cluster"
      "environment"                    = var.environment
      "enable_aws_lb_controller"       = "true"
      "enable_external_dns"            = "true"
    }
    annotations = {
      "cluster_name"            = var.cluster_name
      "aws_region"              = var.region
      "vpc_id"                  = local.vpc_id
      "karpenter_node_role"     = module.karpenter.node_iam_role_name
      "external_dns_zone_id"    = var.route53_zone_id
    }
  }
  data = {
    name   = var.cluster_name
    server = "https://kubernetes.default.svc"
  }
  depends_on = [helm_release.argocd]
}

In the GitOps repository, an ApplicationSet with the cluster generator installs add-ons on clusters whose labels enable them, templating the annotations into Helm values:

apiVersion: argoproj.io/v1alpha1
kind: ApplicationSet
metadata: { name: aws-load-balancer-controller, namespace: argocd }
spec:
  generators:
    - clusters:
        selector: { matchLabels: { enable_aws_lb_controller: "true" } }
  template:
    metadata: { name: "lbc-{{name}}" }
    spec:
      project: platform
      source:
        repoURL: https://aws.github.io/eks-charts
        chart: aws-load-balancer-controller
        targetRevision: 1.13.0            # example: pin a version you tested
        helm:
          values: |
            clusterName: {{metadata.annotations.cluster_name}}
            region: {{metadata.annotations.aws_region}}
            vpcId: {{metadata.annotations.vpc_id}}
      destination: { server: "{{server}}", namespace: kube-system }
      syncPolicy: { automated: { prune: true, selfHeal: true } }

The controller's IAM role comes from a Pod Identity association Terraform created for kube-system/aws-load-balancer-controller, so nothing AWS-specific is typed by hand into Git.

Destroying cleanly

In-cluster controllers create AWS resources that Terraform doesn't know about: ALBs and NLBs from Ingresses and Services, EBS volumes from PVCs, EC2 instances from Karpenter. Before destroying the cluster layer:

  1. Delete workloads (or disable the Argo CD apps) so Ingresses, LoadBalancer Services and PVCs are removed.
  2. Delete Karpenter NodePools and wait for its nodes to terminate.
  3. Confirm no load balancers or volumes remain tagged for the cluster.
  4. Then destroy 30-platform, 20-cluster, 10-network in that order.

Skipping this leaves orphaned load balancers holding ENIs in your subnets, and the VPC destroy fails (lesson 10 has the checks).

Try it: bootstrap and hand over (sandbox)

  1. Apply the Karpenter submodule and Helm release; confirm the controller runs on the system nodes.
  2. Commit a NodePool and EC2NodeClass to your GitOps repo and let Argo CD apply them; scale a test Deployment and watch Karpenter launch a node.
  3. Apply the Argo CD release and the cluster secret; add the ApplicationSet above and watch the Load Balancer Controller appear with values from the annotations.
  4. Practise the clean destroy sequence and verify no ALB or EBS volume is left behind.

Going deeper: the EKS Blueprints add-ons pattern

The approach above is widely known as the GitOps bridge, used by the AWS EKS Blueprints for Terraform. It scales to fleets: one ApplicationSet per add-on, cluster labels switch add-ons on and off per environment, and annotations carry per-cluster AWS values. New clusters get the full platform by being registered, with no per-cluster Git changes.

Recap

  • Terraform: AWS resources, IAM, managed add-ons, and a minimal bootstrap (Karpenter, Argo CD).
  • Argo CD: every other add-on and all workloads, continuously reconciled.
  • The GitOps bridge: Terraform writes AWS values onto the cluster secret; ApplicationSets template them.
  • Karpenter runs on stable system nodes; destroy in-cluster AWS resources before the cluster.

This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.