Lesson 05 of 18 · Part 2 — Build the platform
The EKS cluster: control plane, compute & add-ons
Create the cluster itself and choose how pods get compute: endpoint access, Kubernetes version and support policy, managed node groups vs Fargate vs Karpenter vs Auto Mode, node AMIs, and the managed add-ons every cluster needs.
Decisions you make at creation
Some cluster settings are hard or impossible to change later, so decide them deliberately:
| Decision | Options | Guidance |
|---|---|---|
| Kubernetes version | Current supported minors | Start on a recent version to maximise time in standard support |
| IP family | IPv4 or IPv6 | Fixed at creation (lesson 03) |
| Subnets | Private subnets in 2–3 AZs | The control plane places its network interfaces there |
| Endpoint access | Public, private, or both | Production: private, or public restricted to known CIDRs |
| Authentication mode | API, API_AND_CONFIG_MAP |
Use API (access entries) for new clusters (lesson 06) |
| Secrets encryption | AWS-owned or customer-managed KMS key | Customer-managed key if you must control rotation and access |
| Control-plane logging | api, audit, authenticator, controllerManager, scheduler | At least audit and authenticator in production |
| Support policy | Standard or extended support | Decide whether a cluster may drift into paid extended support |
Choosing cluster settings is like choosing where to pour a building's foundations. You can repaint walls and swap furniture for years, but moving the foundations means knocking the building down.
The API endpoint
The Kubernetes API server is reached through the cluster endpoint:
- Public + private (common): kubectl works from anywhere allowed by the public CIDR list; nodes talk to the API privately inside the VPC.
- Private only: the API is reachable only from the VPC and networks connected to it (VPN, Direct Connect, Transit Gateway), so CI runners and admins must be inside. This is the strongest option and the usual requirement in regulated environments.
Restrict the public endpoint to known CIDRs (office, VPN egress, CI) if you keep it.
Where pods run: the compute options
| Option | You manage | Good for | Watch out for |
|---|---|---|---|
| Managed node groups | Instance types, scaling bounds, AMI version | Steady, well-understood workloads; default choice | One instance-type mix per group; scaling via Cluster Autoscaler or Karpenter |
| Karpenter (on nodes it launches) | NodePools and constraints | Mixed and spiky workloads, Spot, bin-packing | You run and upgrade Karpenter itself (lesson 10) |
| EKS Auto Mode | Node pools and policies only | Teams that want minimal node operations | Extra per-instance fee; less low-level control |
| Fargate | Nothing below the pod | Isolated or small workloads, per-pod billing | No DaemonSets, privileged pods or GPUs; slower start |
| Self-managed nodes | Everything (ASG, AMI, updates) | Special hardware or OS needs | The most operational work |
A common production pattern: a small managed node group for cluster-critical components (CoreDNS, Karpenter itself, controllers), and Karpenter for everything else.
Node operating system
Use an EKS-optimised AMI: Amazon Linux 2023 or Bottlerocket (a minimal, immutable, container-only OS with fast, atomic updates). Newer Kubernetes versions no longer get Amazon Linux 2 AMIs, so don't start new node groups on AL2. Keep node disks encrypted and enforce IMDSv2 in the launch template (lesson 12).
Managed add-ons
Install the cluster's core components as EKS managed add-ons, so versions are validated against the cluster version and upgraded through the same API:
| Add-on | Role |
|---|---|
vpc-cni |
Pod networking (lesson 04) |
coredns |
In-cluster DNS |
kube-proxy |
Service routing on each node |
eks-pod-identity-agent |
AWS credentials for pods (lesson 06) |
aws-ebs-csi-driver, aws-efs-csi-driver |
Persistent volumes (lesson 09) |
metrics-server |
Resource metrics for HPA and kubectl top (lesson 10) |
amazon-cloudwatch-observability |
Container Insights and logs (lesson 11) |
Other platform components (AWS Load Balancer Controller, Karpenter, ExternalDNS, cert-manager, External Secrets) are usually installed with Helm, ideally through GitOps.
A quick cluster for labs
For a sandbox, eksctl builds everything in one command from a config file:
# cluster.yaml
apiVersion: eksctl.io/v1alpha5
kind: ClusterConfig
metadata:
name: lab
region: eu-west-1
version: "1.33"
accessConfig:
authenticationMode: API
managedNodeGroups:
- name: system
instanceType: m6i.large
desiredCapacity: 2
privateNetworking: true
addons:
- name: vpc-cni
- name: coredns
- name: kube-proxy
- name: eks-pod-identity-agent
$ eksctl create cluster -f cluster.yaml
$ kubectl get nodes
$ eksctl delete cluster -f cluster.yaml # when finished: clusters and NAT are billed hourly
For production, build the same with Terraform so it's reviewed, repeatable and versioned; the "Manage EKS with Terraform" track does exactly that. (Pick a Kubernetes version that is currently supported when you run this.)
Try it: compare two compute options
- Create the lab cluster above.
- Add a Fargate profile for namespace
batchand run a Job there:kubectl get pods -o wideshows each pod on its ownfargate-node. - Try running a DaemonSet in
batchand read why it never lands on Fargate. - Delete the cluster.
Going deeper: cluster topology
One cluster per environment per region is the usual starting point. Split further when you need a hard boundary: different compliance scopes, very different upgrade cadences, or blast-radius limits for critical services. Many small clusters give isolation but multiply upgrades, add-ons and monitoring; few large clusters are cheaper to run but need strong multi-tenancy (namespaces, quotas, policies, network policies). Lesson 16 turns this into a design-review question.
Recap
- Decide version, IP family, subnets, endpoint access, auth mode, encryption and logging up front.
- Prefer a private API endpoint (or a CIDR-restricted public one).
- Compute: managed node groups for system pods + Karpenter for the rest is a strong default; Fargate for isolated small workloads; Auto Mode to hand node operations to AWS.
- Use AL2023 or Bottlerocket nodes and managed add-ons for core components.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.