Production EKS Platform — From Zero to Production›05 · The EKS cluster: control plane, compute & add-ons

Lesson 05 of 18 · Part 2 — Build the platform

The EKS cluster: control plane, compute & add-ons

Create the cluster itself and choose how pods get compute: endpoint access, Kubernetes version and support policy, managed node groups vs Fargate vs Karpenter vs Auto Mode, node AMIs, and the managed add-ons every cluster needs.

Intermediate → Advanced
Key wordsEKS clusterKubernetes versionendpoint accesspublic endpointprivate endpointmanaged node groupsself-managed nodesFargateEKS Auto ModeKarpenterAL2023Bottlerocketmanaged add-onseksctl

Decisions you make at creation

Some cluster settings are hard or impossible to change later, so decide them deliberately:

Decision Options Guidance
Kubernetes version Current supported minors Start on a recent version to maximise time in standard support
IP family IPv4 or IPv6 Fixed at creation (lesson 03)
Subnets Private subnets in 2–3 AZs The control plane places its network interfaces there
Endpoint access Public, private, or both Production: private, or public restricted to known CIDRs
Authentication mode API, API_AND_CONFIG_MAP Use API (access entries) for new clusters (lesson 06)
Secrets encryption AWS-owned or customer-managed KMS key Customer-managed key if you must control rotation and access
Control-plane logging api, audit, authenticator, controllerManager, scheduler At least audit and authenticator in production
Support policy Standard or extended support Decide whether a cluster may drift into paid extended support

Choosing cluster settings is like choosing where to pour a building's foundations. You can repaint walls and swap furniture for years, but moving the foundations means knocking the building down.

The API endpoint

The Kubernetes API server is reached through the cluster endpoint:

  • Public + private (common): kubectl works from anywhere allowed by the public CIDR list; nodes talk to the API privately inside the VPC.
  • Private only: the API is reachable only from the VPC and networks connected to it (VPN, Direct Connect, Transit Gateway), so CI runners and admins must be inside. This is the strongest option and the usual requirement in regulated environments.

Restrict the public endpoint to known CIDRs (office, VPN egress, CI) if you keep it.

Where pods run: the compute options

Option You manage Good for Watch out for
Managed node groups Instance types, scaling bounds, AMI version Steady, well-understood workloads; default choice One instance-type mix per group; scaling via Cluster Autoscaler or Karpenter
Karpenter (on nodes it launches) NodePools and constraints Mixed and spiky workloads, Spot, bin-packing You run and upgrade Karpenter itself (lesson 10)
EKS Auto Mode Node pools and policies only Teams that want minimal node operations Extra per-instance fee; less low-level control
Fargate Nothing below the pod Isolated or small workloads, per-pod billing No DaemonSets, privileged pods or GPUs; slower start
Self-managed nodes Everything (ASG, AMI, updates) Special hardware or OS needs The most operational work

A common production pattern: a small managed node group for cluster-critical components (CoreDNS, Karpenter itself, controllers), and Karpenter for everything else.

Node operating system

Use an EKS-optimised AMI: Amazon Linux 2023 or Bottlerocket (a minimal, immutable, container-only OS with fast, atomic updates). Newer Kubernetes versions no longer get Amazon Linux 2 AMIs, so don't start new node groups on AL2. Keep node disks encrypted and enforce IMDSv2 in the launch template (lesson 12).

Managed add-ons

Install the cluster's core components as EKS managed add-ons, so versions are validated against the cluster version and upgraded through the same API:

Add-on Role
vpc-cni Pod networking (lesson 04)
coredns In-cluster DNS
kube-proxy Service routing on each node
eks-pod-identity-agent AWS credentials for pods (lesson 06)
aws-ebs-csi-driver, aws-efs-csi-driver Persistent volumes (lesson 09)
metrics-server Resource metrics for HPA and kubectl top (lesson 10)
amazon-cloudwatch-observability Container Insights and logs (lesson 11)

Other platform components (AWS Load Balancer Controller, Karpenter, ExternalDNS, cert-manager, External Secrets) are usually installed with Helm, ideally through GitOps.

A quick cluster for labs

For a sandbox, eksctl builds everything in one command from a config file:

# cluster.yaml
apiVersion: eksctl.io/v1alpha5
kind: ClusterConfig
metadata:
  name: lab
  region: eu-west-1
  version: "1.33"
accessConfig:
  authenticationMode: API
managedNodeGroups:
  - name: system
    instanceType: m6i.large
    desiredCapacity: 2
    privateNetworking: true
addons:
  - name: vpc-cni
  - name: coredns
  - name: kube-proxy
  - name: eks-pod-identity-agent
$ eksctl create cluster -f cluster.yaml
$ kubectl get nodes
$ eksctl delete cluster -f cluster.yaml        # when finished: clusters and NAT are billed hourly

For production, build the same with Terraform so it's reviewed, repeatable and versioned; the "Manage EKS with Terraform" track does exactly that. (Pick a Kubernetes version that is currently supported when you run this.)

Try it: compare two compute options

  1. Create the lab cluster above.
  2. Add a Fargate profile for namespace batch and run a Job there: kubectl get pods -o wide shows each pod on its own fargate- node.
  3. Try running a DaemonSet in batch and read why it never lands on Fargate.
  4. Delete the cluster.

Going deeper: cluster topology

One cluster per environment per region is the usual starting point. Split further when you need a hard boundary: different compliance scopes, very different upgrade cadences, or blast-radius limits for critical services. Many small clusters give isolation but multiply upgrades, add-ons and monitoring; few large clusters are cheaper to run but need strong multi-tenancy (namespaces, quotas, policies, network policies). Lesson 16 turns this into a design-review question.

Recap

  • Decide version, IP family, subnets, endpoint access, auth mode, encryption and logging up front.
  • Prefer a private API endpoint (or a CIDR-restricted public one).
  • Compute: managed node groups for system pods + Karpenter for the rest is a strong default; Fargate for isolated small workloads; Auto Mode to hand node operations to AWS.
  • Use AL2023 or Bottlerocket nodes and managed add-ons for core components.

This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.