Lesson 04 of 17 · Part 2 — Build as code
The VPC as code
Build the EKS network layer in Terraform: a three-AZ VPC with public and private subnets, per-AZ NAT, the load-balancer and Karpenter subnet tags, VPC endpoints, and an optional secondary CIDR with pod subnets.
The layer's job
10-network creates everything the cluster needs from the network and publishes the IDs for the next layers. It changes rarely, so its plans should almost always be empty; a non-empty network plan deserves a second reviewer.
Laying the roads before building the houses. Once houses stand on them, widening a road means knocking something down, so you draw the roads wide enough the first time.
The VPC with the community module
module "vpc" {
source = "terraform-aws-modules/vpc/aws"
version = "~> 6.0" # pin the major version you tested
name = "${var.environment}-eks"
cidr = "10.20.0.0/16"
azs = ["${var.region}a", "${var.region}b", "${var.region}c"]
public_subnets = ["10.20.0.0/22", "10.20.4.0/22", "10.20.8.0/22"]
private_subnets = ["10.20.32.0/19", "10.20.64.0/19", "10.20.96.0/19"]
enable_nat_gateway = true
single_nat_gateway = var.environment != "prod" # one NAT in dev to save cost
one_nat_gateway_per_az = var.environment == "prod"
enable_dns_hostnames = true
enable_dns_support = true
public_subnet_tags = {
"kubernetes.io/role/elb" = 1
}
private_subnet_tags = {
"kubernetes.io/role/internal-elb" = 1
"karpenter.sh/discovery" = var.cluster_name
}
}
The module turns a few inputs into dozens of resources: the VPC, subnets, route tables, IGW, NAT gateways and Elastic IPs, and their associations. Read its outputs (vpc_id, private_subnets, public_subnets, route table IDs) rather than looking IDs up by hand.
VPC endpoints
The S3 gateway endpoint is free and removes image-layer traffic from NAT. Interface endpoints cost per hour per AZ, so make them a variable per environment:
resource "aws_vpc_endpoint" "s3" {
vpc_id = module.vpc.vpc_id
service_name = "com.amazonaws.${var.region}.s3"
vpc_endpoint_type = "Gateway"
route_table_ids = module.vpc.private_route_table_ids
}
resource "aws_security_group" "endpoints" {
name = "${var.environment}-vpc-endpoints"
vpc_id = module.vpc.vpc_id
ingress {
from_port = 443
to_port = 443
protocol = "tcp"
cidr_blocks = [module.vpc.vpc_cidr_block]
}
}
resource "aws_vpc_endpoint" "interface" {
for_each = toset(var.interface_endpoints) # e.g. ["ecr.api", "ecr.dkr", "sts", "logs", "eks-auth"]
vpc_id = module.vpc.vpc_id
service_name = "com.amazonaws.${var.region}.${each.value}"
vpc_endpoint_type = "Interface"
subnet_ids = module.vpc.private_subnets
security_group_ids = [aws_security_group.endpoints.id]
private_dns_enabled = true
}
The module also ships a vpc-endpoints submodule if you prefer it.
Optional: a secondary CIDR for pods
When routable space is scarce, attach 100.64.0.0/16 and create pod-only subnets (lesson "VPC CNI & IP planning" in the platform track explains why):
resource "aws_vpc_ipv4_cidr_block_association" "pods" {
vpc_id = module.vpc.vpc_id
cidr_block = "100.64.0.0/16"
}
resource "aws_subnet" "pods" {
for_each = { a = "100.64.0.0/18", b = "100.64.64.0/18", c = "100.64.128.0/18" }
vpc_id = module.vpc.vpc_id
availability_zone = "${var.region}${each.key}"
cidr_block = each.value
tags = { Name = "${var.environment}-pods-${each.key}" }
depends_on = [aws_vpc_ipv4_cidr_block_association.pods]
}
resource "aws_route_table_association" "pods" {
for_each = aws_subnet.pods
subnet_id = each.value.id
route_table_id = module.vpc.private_route_table_ids[index(["a", "b", "c"], each.key)]
}
The VPC CNI then needs custom networking (AWS_VPC_K8S_CNI_CUSTOM_NETWORK_CFG=true) and one ENIConfig per AZ pointing at these subnets. Those are cluster-side settings, applied in the cluster and platform layers.
Publish outputs
output "vpc_id" { value = module.vpc.vpc_id }
output "private_subnets" { value = module.vpc.private_subnets }
output "pod_subnets" { value = { for k, s in aws_subnet.pods : k => s.id } }
resource "aws_ssm_parameter" "vpc_id" {
name = "/platform/${var.environment}/vpc_id"
type = "String"
value = module.vpc.vpc_id
}
Replacement traps in the network layer
Changing cidr, a subnet's CIDR, or the AZ list (including reordering it) replaces subnets, and everything in them. Read every network plan for must be replaced, keep subnet lists stable, and add capacity with a secondary CIDR rather than resizing.
Try it: build and inspect (sandbox)
- Apply the VPC module with
single_nat_gateway = truein a sandbox. - Add the S3 gateway endpoint and two interface endpoints; confirm with
aws ec2 describe-vpc-endpoints. - Change the order of
azsand runplan(don't apply): note how many resources would be replaced. - Destroy the layer when finished; NAT gateways and interface endpoints are billed hourly.
Recap
- One network layer with the VPC module: 3 AZs, public and private subnets, per-AZ NAT in prod.
- Tag subnets for load balancers and Karpenter discovery.
- S3 gateway endpoint always; interface endpoints per environment need.
- Optional secondary CIDR with pod subnets for custom networking.
- CIDR and AZ changes replace subnets: plan space once, read plans carefully.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.