Lesson 08 of 18 · Part 2 — Build the platform
ECR & application deployment
Get an application from a laptop into production on EKS: private ECR repositories with scanning, immutable tags and lifecycle rules, how nodes pull images, and a Deployment built to roll out safely across AZs.
Images live in ECR
Amazon ECR is the private registry next to your clusters. Set every repository up the same way:
| Setting | Why |
|---|---|
| Tag immutability | 1.4.0 always means the same image; deploy by tag + digest for extra certainty |
| Scan on push (basic) or enhanced scanning (Amazon Inspector, continuous) | Know about CVEs before and after release |
| Lifecycle policy | Keep the last N releases, expire untagged images: storage stays bounded |
| Encryption | AWS-managed or KMS key |
| Repository policy | Cross-account pulls (e.g. a shared-services ECR used by prod and non-prod accounts) |
| Replication | Copy images to other regions/accounts for DR and latency |
| Pull-through cache | Mirror upstream registries (Docker Hub, Quay, GitHub) into ECR: no rate limits, one place to scan |
A lifecycle policy that keeps the last 30 release tags and cleans untagged images after a week:
{
"rules": [
{ "rulePriority": 1, "description": "expire untagged after 7 days",
"selection": { "tagStatus": "untagged", "countType": "sinceImagePushed", "countUnit": "days", "countNumber": 7 },
"action": { "type": "expire" } },
{ "rulePriority": 2, "description": "keep last 30 releases",
"selection": { "tagStatus": "tagged", "tagPrefixList": ["1.", "2."], "countType": "imageCountMoreThan", "countNumber": 30 },
"action": { "type": "expire" } }
]
}
ECR is the warehouse, and every box has a label that can never be changed once it's on the shelf. The delivery van (the node) has a key card for the warehouse built in, so nobody has to hand the driver a password each morning.
How nodes pull
The kubelet pulls images using the node IAM role, which carries an ECR pull policy. For images in the same account there are no pull secrets to manage. For a shared-services account, add a repository policy allowing the workload accounts' node roles (or the accounts) to pull.
Pull traffic goes to ECR's API and to S3 (where layers live), so the S3 gateway endpoint and ECR interface endpoints from lesson 03 matter here, both for cost and for private clusters.
A Deployment built for production
apiVersion: apps/v1
kind: Deployment
metadata: { name: api, namespace: shop }
spec:
replicas: 3
strategy:
rollingUpdate: { maxUnavailable: 0, maxSurge: 1 } # never drop below 3 serving pods
selector: { matchLabels: { app: api } }
template:
metadata: { labels: { app: api } }
spec:
serviceAccountName: api # its own AWS role (lesson 06)
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: DoNotSchedule
labelSelector: { matchLabels: { app: api } }
containers:
- name: api
image: 111122223333.dkr.ecr.eu-west-1.amazonaws.com/shop/api:1.4.0
ports: [ { containerPort: 8080 } ]
resources:
requests: { cpu: 250m, memory: 256Mi }
limits: { memory: 512Mi }
readinessProbe: { httpGet: { path: /ready, port: 8080 }, periodSeconds: 5 }
livenessProbe: { httpGet: { path: /healthz, port: 8080 }, initialDelaySeconds: 20 }
lifecycle:
preStop: { exec: { command: ["sleep", "15"] } } # let the load balancer deregister first
securityContext:
runAsNonRoot: true
readOnlyRootFilesystem: true
allowPrivilegeEscalation: false
---
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata: { name: api, namespace: shop }
spec:
minAvailable: 2
selector: { matchLabels: { app: api } }
What each part buys you:
| Part | Protects against |
|---|---|
maxUnavailable: 0 + readiness probe |
Traffic loss during rollouts |
preStop sleep |
Requests hitting a pod the ALB hasn't deregistered yet |
| Topology spread across zones | One AZ failure taking out every replica |
| PodDisruptionBudget | Node drains and consolidation removing too many replicas at once |
| Requests and memory limit | Noisy neighbours; scheduling on nodes that can't fit it |
| Security context | Container escape and tampering (lesson 12) |
Expose it with a Service and an Ingress (lesson 07); scale it with an HPA (lesson 10).
Getting it there: pipelines, not laptops
In production, humans don't kubectl apply. The usual flow:
- CI builds, tests and scans the image, and pushes
shop/api:1.4.1to ECR. - CI updates the image tag in the deployment repository (Helm values or Kustomize).
- Argo CD (or Flux) in the cluster sees the change and syncs it; the rollout is visible and reversible in Git.
Lesson 15 covers how people and pipelines reach the cluster; the "Argo CD — Level by Level" track covers the delivery side in depth.
Try it: ship and roll back (sandbox)
- Create an ECR repository with immutable tags and scan on push; push two versions of any small web image (
1.0.0,1.0.1). - Deploy
1.0.0with the manifest above (adjust image and probes); check the spread across zones. kubectl set imageto1.0.1and watchkubectl rollout status; thenkubectl rollout undo.- Try pushing
1.0.0again with different content: ECR refuses it. - Drain one node (
kubectl drain <node> --ignore-daemonsets --delete-emptydir-data) and watch the PDB keep two replicas serving.
Recap
- ECR repositories with immutable tags, scanning, lifecycle rules; pull-through cache for upstream images; replication for DR.
- Nodes pull with the node role; cross-account pulls need a repository policy; endpoints keep pulls private and cheap.
- Production Deployments: readiness, preStop, maxUnavailable 0, zone spread, PDB, requests and limits, a restrictive security context.
- Changes flow through CI → ECR → Git → Argo CD, not from laptops.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.