Production EKS Platform — From Zero to Production›03 · AWS account & VPC for EKS

Lesson 03 of 18 · Part 2 — Build the platform

AWS account & VPC for EKS

Lay the ground an EKS cluster stands on: the account structure, a VPC sized for pods, public and private subnets across three AZs, NAT and VPC endpoints, and the subnet tags that let load balancers find the right subnets.

Intermediate
Key wordsAWS accountOrganizationslanding zoneVPCCIDRavailability zonespublic subnetprivate subnetNAT gatewayVPC endpointssubnet tagskubernetes.io/role/elb

Start with the account

Put production in its own AWS account, separate from non-production, inside an AWS Organization. An account is the strongest isolation boundary AWS has: separate IAM, separate quotas, separate bill, and a mistake in dev can't touch prod. A typical minimal layout:

Account Holds
Management Organizations, billing, SCPs only (no workloads)
Security / log archive CloudTrail, GuardDuty and Config aggregation
Shared services ECR, CI runners, DNS zones, Transit Gateway
Prod / non-prod workload accounts The EKS clusters and their data

People sign in once through IAM Identity Center and assume a role in the account they need. No IAM users with access keys.

Separate accounts are separate houses, not separate rooms. If a pipe bursts in the test house, the production house next door stays dry, because they share nothing but the street.

Size the VPC for pods, not nodes

With the Amazon VPC CNI, every pod gets an IP from the VPC. A cluster with 50 nodes running 30 pods each needs well over 1,500 addresses before you count load balancers, endpoints and headroom for rolling updates.

A common production plan:

CIDR Use
10.20.0.0/16 Primary VPC range: subnets for nodes, load balancers, endpoints
100.64.0.0/16 (secondary) Pod-only subnets, if routable space is scarce (lesson 04)

Pick ranges that don't overlap with other VPCs, the corporate network or VPN ranges you might ever need to connect to. Re-addressing a VPC later means rebuilding it.

Subnets across three AZs

                 AZ a                AZ b                AZ c
Public      10.20.0.0/22       10.20.4.0/22       10.20.8.0/22      ALBs/NLBs, NAT gateways
Private     10.20.32.0/19      10.20.64.0/19      10.20.96.0/19     nodes (and pods)
(optional)  100.64.0.0/18      100.64.64.0/18     100.64.128.0/18   pods only (secondary CIDR)
  • Public subnets route 0.0.0.0/0 to the Internet Gateway. Only internet-facing load balancers and NAT gateways live here.
  • Private subnets route 0.0.0.0/0 to a NAT gateway in the same AZ. Nodes live here, with no public IPs.
  • Use three AZs in production: losing one AZ then removes a third of capacity, not half.

Tag subnets for Kubernetes

The AWS Load Balancer Controller discovers subnets by tag:

Subnet Tag
Public kubernetes.io/role/elb = 1
Private kubernetes.io/role/internal-elb = 1

Karpenter and some tools also select subnets and security groups by tags you choose (for example karpenter.sh/discovery = <cluster>). Decide your tag scheme once and apply it everywhere.

NAT gateways and VPC endpoints

Nodes need outbound access to pull images and reach AWS APIs. Two ways, usually combined:

  • NAT gateway, one per AZ. Simple, but you pay per GB processed, and image pulls through NAT add up quickly.
  • VPC endpoints keep that traffic on the AWS network:
Endpoint Type Why
S3 Gateway (free) ECR image layers are served from S3
ECR API, ECR DKR Interface Image pulls
STS Interface Pod Identity / IRSA credential exchange
EC2, EKS, EKS Auth Interface Node bootstrap, Karpenter, Pod Identity
CloudWatch Logs, SSM Interface Logging, Session Manager access to nodes

Interface endpoints cost per hour per AZ, so add them where the traffic justifies it or where the cluster must be fully private (no NAT at all).

Cost trap

NAT data processing is one of the most common surprise lines on an EKS bill. Put the S3 gateway endpoint in on day one, and ECR endpoints as soon as image traffic is significant.

Try it: build the network (sandbox)

  1. In a sandbox account, create a VPC 10.20.0.0/16 with the "VPC and more" console wizard: 3 AZs, 3 public and 3 private subnets, one NAT gateway per AZ, S3 gateway endpoint.
  2. Add the two ELB tags to the subnets.
  3. Run the cheat-sheet commands and confirm: private route tables point at the NAT in their own AZ, public ones at the IGW.
  4. Delete the NAT gateways when you're done: they are billed per hour.

Going deeper: IPv6 clusters

EKS supports IPv6 clusters, where pods get IPv6 addresses from the VPC's IPv6 range and IP exhaustion effectively disappears. The trade-off is that everything the pods talk to must handle IPv6 (or go through egress-only / NAT64 paths), and it's chosen at cluster creation. Many organisations stay on IPv4 with a secondary pod CIDR; IPv6 is worth evaluating for very large or fast-growing platforms.

Recap

  • One account per environment in an Organization; people use SSO roles, not access keys.
  • Size the VPC for pods, avoid overlapping CIDRs, keep a secondary range for pods.
  • Three AZs: public subnets for load balancers and NAT, private subnets for nodes, one NAT per AZ.
  • Tag subnets for load balancers; add VPC endpoints (S3 at least) to cut NAT cost and enable private clusters.

This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.