Production EKS Platform — From Zero to Production›08 · ECR & application deployment

Lesson 08 of 18 · Part 2 — Build the platform

ECR & application deployment

Get an application from a laptop into production on EKS: private ECR repositories with scanning, immutable tags and lifecycle rules, how nodes pull images, and a Deployment built to roll out safely across AZs.

Intermediate
Key wordsAmazon ECRprivate repositoryimage scanningAmazon Inspectortag immutabilitylifecycle policypull-through cachereplicationnode roleDeploymentrolling updatereadiness probePodDisruptionBudgettopology spread

Images live in ECR

Amazon ECR is the private registry next to your clusters. Set every repository up the same way:

Setting Why
Tag immutability 1.4.0 always means the same image; deploy by tag + digest for extra certainty
Scan on push (basic) or enhanced scanning (Amazon Inspector, continuous) Know about CVEs before and after release
Lifecycle policy Keep the last N releases, expire untagged images: storage stays bounded
Encryption AWS-managed or KMS key
Repository policy Cross-account pulls (e.g. a shared-services ECR used by prod and non-prod accounts)
Replication Copy images to other regions/accounts for DR and latency
Pull-through cache Mirror upstream registries (Docker Hub, Quay, GitHub) into ECR: no rate limits, one place to scan

A lifecycle policy that keeps the last 30 release tags and cleans untagged images after a week:

{
  "rules": [
    { "rulePriority": 1, "description": "expire untagged after 7 days",
      "selection": { "tagStatus": "untagged", "countType": "sinceImagePushed", "countUnit": "days", "countNumber": 7 },
      "action": { "type": "expire" } },
    { "rulePriority": 2, "description": "keep last 30 releases",
      "selection": { "tagStatus": "tagged", "tagPrefixList": ["1.", "2."], "countType": "imageCountMoreThan", "countNumber": 30 },
      "action": { "type": "expire" } }
  ]
}

ECR is the warehouse, and every box has a label that can never be changed once it's on the shelf. The delivery van (the node) has a key card for the warehouse built in, so nobody has to hand the driver a password each morning.

How nodes pull

The kubelet pulls images using the node IAM role, which carries an ECR pull policy. For images in the same account there are no pull secrets to manage. For a shared-services account, add a repository policy allowing the workload accounts' node roles (or the accounts) to pull.

Pull traffic goes to ECR's API and to S3 (where layers live), so the S3 gateway endpoint and ECR interface endpoints from lesson 03 matter here, both for cost and for private clusters.

A Deployment built for production

apiVersion: apps/v1
kind: Deployment
metadata: { name: api, namespace: shop }
spec:
  replicas: 3
  strategy:
    rollingUpdate: { maxUnavailable: 0, maxSurge: 1 }   # never drop below 3 serving pods
  selector: { matchLabels: { app: api } }
  template:
    metadata: { labels: { app: api } }
    spec:
      serviceAccountName: api                            # its own AWS role (lesson 06)
      topologySpreadConstraints:
        - maxSkew: 1
          topologyKey: topology.kubernetes.io/zone
          whenUnsatisfiable: DoNotSchedule
          labelSelector: { matchLabels: { app: api } }
      containers:
        - name: api
          image: 111122223333.dkr.ecr.eu-west-1.amazonaws.com/shop/api:1.4.0
          ports: [ { containerPort: 8080 } ]
          resources:
            requests: { cpu: 250m, memory: 256Mi }
            limits: { memory: 512Mi }
          readinessProbe: { httpGet: { path: /ready, port: 8080 }, periodSeconds: 5 }
          livenessProbe: { httpGet: { path: /healthz, port: 8080 }, initialDelaySeconds: 20 }
          lifecycle:
            preStop: { exec: { command: ["sleep", "15"] } }   # let the load balancer deregister first
          securityContext:
            runAsNonRoot: true
            readOnlyRootFilesystem: true
            allowPrivilegeEscalation: false
---
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata: { name: api, namespace: shop }
spec:
  minAvailable: 2
  selector: { matchLabels: { app: api } }

What each part buys you:

Part Protects against
maxUnavailable: 0 + readiness probe Traffic loss during rollouts
preStop sleep Requests hitting a pod the ALB hasn't deregistered yet
Topology spread across zones One AZ failure taking out every replica
PodDisruptionBudget Node drains and consolidation removing too many replicas at once
Requests and memory limit Noisy neighbours; scheduling on nodes that can't fit it
Security context Container escape and tampering (lesson 12)

Expose it with a Service and an Ingress (lesson 07); scale it with an HPA (lesson 10).

Getting it there: pipelines, not laptops

In production, humans don't kubectl apply. The usual flow:

  1. CI builds, tests and scans the image, and pushes shop/api:1.4.1 to ECR.
  2. CI updates the image tag in the deployment repository (Helm values or Kustomize).
  3. Argo CD (or Flux) in the cluster sees the change and syncs it; the rollout is visible and reversible in Git.

Lesson 15 covers how people and pipelines reach the cluster; the "Argo CD — Level by Level" track covers the delivery side in depth.

Try it: ship and roll back (sandbox)

  1. Create an ECR repository with immutable tags and scan on push; push two versions of any small web image (1.0.0, 1.0.1).
  2. Deploy 1.0.0 with the manifest above (adjust image and probes); check the spread across zones.
  3. kubectl set image to 1.0.1 and watch kubectl rollout status; then kubectl rollout undo.
  4. Try pushing 1.0.0 again with different content: ECR refuses it.
  5. Drain one node (kubectl drain <node> --ignore-daemonsets --delete-emptydir-data) and watch the PDB keep two replicas serving.

Recap

  • ECR repositories with immutable tags, scanning, lifecycle rules; pull-through cache for upstream images; replication for DR.
  • Nodes pull with the node role; cross-account pulls need a repository policy; endpoints keep pulls private and cheap.
  • Production Deployments: readiness, preStop, maxUnavailable 0, zone spread, PDB, requests and limits, a restrictive security context.
  • Changes flow through CI → ECR → Git → Argo CD, not from laptops.

This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.