Lesson 03 of 18 · Part 2 — Build the platform
AWS account & VPC for EKS
Lay the ground an EKS cluster stands on: the account structure, a VPC sized for pods, public and private subnets across three AZs, NAT and VPC endpoints, and the subnet tags that let load balancers find the right subnets.
Start with the account
Put production in its own AWS account, separate from non-production, inside an AWS Organization. An account is the strongest isolation boundary AWS has: separate IAM, separate quotas, separate bill, and a mistake in dev can't touch prod. A typical minimal layout:
| Account | Holds |
|---|---|
| Management | Organizations, billing, SCPs only (no workloads) |
| Security / log archive | CloudTrail, GuardDuty and Config aggregation |
| Shared services | ECR, CI runners, DNS zones, Transit Gateway |
| Prod / non-prod workload accounts | The EKS clusters and their data |
People sign in once through IAM Identity Center and assume a role in the account they need. No IAM users with access keys.
Separate accounts are separate houses, not separate rooms. If a pipe bursts in the test house, the production house next door stays dry, because they share nothing but the street.
Size the VPC for pods, not nodes
With the Amazon VPC CNI, every pod gets an IP from the VPC. A cluster with 50 nodes running 30 pods each needs well over 1,500 addresses before you count load balancers, endpoints and headroom for rolling updates.
A common production plan:
| CIDR | Use |
|---|---|
10.20.0.0/16 |
Primary VPC range: subnets for nodes, load balancers, endpoints |
100.64.0.0/16 (secondary) |
Pod-only subnets, if routable space is scarce (lesson 04) |
Pick ranges that don't overlap with other VPCs, the corporate network or VPN ranges you might ever need to connect to. Re-addressing a VPC later means rebuilding it.
Subnets across three AZs
AZ a AZ b AZ c
Public 10.20.0.0/22 10.20.4.0/22 10.20.8.0/22 ALBs/NLBs, NAT gateways
Private 10.20.32.0/19 10.20.64.0/19 10.20.96.0/19 nodes (and pods)
(optional) 100.64.0.0/18 100.64.64.0/18 100.64.128.0/18 pods only (secondary CIDR)
- Public subnets route
0.0.0.0/0to the Internet Gateway. Only internet-facing load balancers and NAT gateways live here. - Private subnets route
0.0.0.0/0to a NAT gateway in the same AZ. Nodes live here, with no public IPs. - Use three AZs in production: losing one AZ then removes a third of capacity, not half.
Tag subnets for Kubernetes
The AWS Load Balancer Controller discovers subnets by tag:
| Subnet | Tag |
|---|---|
| Public | kubernetes.io/role/elb = 1 |
| Private | kubernetes.io/role/internal-elb = 1 |
Karpenter and some tools also select subnets and security groups by tags you choose (for example karpenter.sh/discovery = <cluster>). Decide your tag scheme once and apply it everywhere.
NAT gateways and VPC endpoints
Nodes need outbound access to pull images and reach AWS APIs. Two ways, usually combined:
- NAT gateway, one per AZ. Simple, but you pay per GB processed, and image pulls through NAT add up quickly.
- VPC endpoints keep that traffic on the AWS network:
| Endpoint | Type | Why |
|---|---|---|
| S3 | Gateway (free) | ECR image layers are served from S3 |
| ECR API, ECR DKR | Interface | Image pulls |
| STS | Interface | Pod Identity / IRSA credential exchange |
| EC2, EKS, EKS Auth | Interface | Node bootstrap, Karpenter, Pod Identity |
| CloudWatch Logs, SSM | Interface | Logging, Session Manager access to nodes |
Interface endpoints cost per hour per AZ, so add them where the traffic justifies it or where the cluster must be fully private (no NAT at all).
Cost trap
NAT data processing is one of the most common surprise lines on an EKS bill. Put the S3 gateway endpoint in on day one, and ECR endpoints as soon as image traffic is significant.
Try it: build the network (sandbox)
- In a sandbox account, create a VPC
10.20.0.0/16with the "VPC and more" console wizard: 3 AZs, 3 public and 3 private subnets, one NAT gateway per AZ, S3 gateway endpoint. - Add the two ELB tags to the subnets.
- Run the cheat-sheet commands and confirm: private route tables point at the NAT in their own AZ, public ones at the IGW.
- Delete the NAT gateways when you're done: they are billed per hour.
Going deeper: IPv6 clusters
EKS supports IPv6 clusters, where pods get IPv6 addresses from the VPC's IPv6 range and IP exhaustion effectively disappears. The trade-off is that everything the pods talk to must handle IPv6 (or go through egress-only / NAT64 paths), and it's chosen at cluster creation. Many organisations stay on IPv4 with a secondary pod CIDR; IPv6 is worth evaluating for very large or fast-growing platforms.
Recap
- One account per environment in an Organization; people use SSO roles, not access keys.
- Size the VPC for pods, avoid overlapping CIDRs, keep a secondary range for pods.
- Three AZs: public subnets for load balancers and NAT, private subnets for nodes, one NAT per AZ.
- Tag subnets for load balancers; add VPC endpoints (S3 at least) to cut NAT cost and enable private clusters.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.