Lesson 12 of 18 · Part 3 — Operate
Security: KMS, GuardDuty, IMDSv2, Pod Security & network policy
Layered security for an EKS platform: the shared-responsibility line, protecting the API and secrets, hardening nodes, restricting pods and their network, supply-chain checks, secrets from Secrets Manager, and detection with GuardDuty and audit logs.
Who secures what
| AWS | You |
|---|---|
| Control plane: API servers, etcd, their patching and availability | API endpoint exposure, authentication mode, who has access |
| The underlying hardware and hypervisor | Node OS updates (managed node groups: you roll AMIs), node configuration |
| Fargate and Auto Mode node patching | Pod security, network policy, images, secrets, IAM for pods |
Security on EKS is defence in depth: no single control is expected to stop everything.
A bank doesn't rely only on the front door lock. There's a guard, cameras, a vault with its own code, and a rule that no single employee can open it alone. If one layer fails, the next one still holds.
1. Protect the API
- Private endpoint, or a public one restricted to known CIDRs (lesson 05).
- Access entries with least privilege and SSO roles; no IAM users; a tested break-glass role (lesson 06).
- Audit and authenticator logs enabled and retained (lesson 11).
2. Protect data and secrets
- Kubernetes Secrets are stored encrypted by EKS; use envelope encryption with a customer-managed KMS key when you must control key policy, rotation and access, and audit key use in CloudTrail.
- Keep application secrets in AWS Secrets Manager or Parameter Store. Either sync them into Kubernetes Secrets with the External Secrets Operator, or let the app read them directly with its own pod role.
- Encrypt EBS and EFS volumes (default encryption per account is a one-time setting worth enabling).
3. Harden the nodes
- EKS-optimised AMIs: AL2023 or Bottlerocket (minimal, immutable, fewer packages to patch); roll to new AMIs on a schedule.
- IMDSv2 required, hop limit 1 so pods can't use the node role (lesson 06).
- No SSH: use SSM Session Manager when you must get on a node, with logging.
- Nodes in private subnets with tight security groups.
- Check against the CIS Amazon EKS Benchmark (kube-bench can run the node checks).
4. Restrict pods
Enforce Pod Security Standards with labels on each namespace:
metadata:
labels:
pod-security.kubernetes.io/enforce: restricted # reject what violates it
pod-security.kubernetes.io/warn: restricted # warn clients too
Add an admission policy engine (Kyverno or Gatekeeper) for rules Pod Security doesn't cover: approved registries only (your ECR), no latest tags, required labels, required resource limits, signed images.
5. Restrict the network
Pods can talk to every other pod by default. Turn on network policy support in the VPC CNI (or run Cilium/Calico for it), then default-deny per namespace and allow what's needed:
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata: { name: default-deny, namespace: shop }
spec:
podSelector: {}
policyTypes: [Ingress, Egress]
---
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata: { name: api-from-ingress, namespace: shop }
spec:
podSelector: { matchLabels: { app: api } }
ingress:
- ports: [ { port: 8080 } ] # traffic from the ALB (IP targets) reaches pods directly
egress:
- to: [ { namespaceSelector: { matchLabels: { kubernetes.io/metadata.name: kube-system } }, podSelector: { matchLabels: { k8s-app: kube-dns } } } ]
ports: [ { port: 53, protocol: UDP } ]
Remember to allow DNS, or everything breaks at once. For traffic to AWS resources like RDS, security groups for pods let a database's security group allow only specific pods rather than every node.
6. Secure the supply chain
Scan images in CI and in ECR (enhanced scanning with Amazon Inspector re-scans as new CVEs appear), sign images (cosign, or AWS Signer with Notation) and verify signatures at admission. See the "GitHub Actions — Level by Level" track.
7. Detect and respond
- GuardDuty EKS protection (audit-log analysis) and runtime monitoring (agent on nodes).
- CloudTrail for AWS API changes; Security Hub to aggregate findings.
- Alerts on: new cluster-admin access entries, break-glass use, changes to the endpoint's public CIDRs, disabled logging.
Try it: find the gaps (sandbox cluster)
- Run the posture commands in the cheat sheet and list every "no" you find (public endpoint open to
0.0.0.0/0, no encryption config, namespaces without Pod Security labels, no network policies). - Label a namespace
restrictedand try a privileged pod; read the rejection. - Enable network policy in the VPC CNI, apply the two policies above and prove a pod in another namespace can no longer reach
api, while DNS still works. - Run
trivy imageon one of your images and fix one HIGH finding by updating the base image.
Going deeper: prove it continuously
Point-in-time hardening decays. Run policy checks in CI (Kyverno CLI or conftest against manifests), keep kube-bench and an image-scan report in each release, and track posture as a small set of measurable controls: percentage of namespaces enforcing restricted, images older than 30 days, critical CVEs in running images, and time to rotate a compromised credential.
Recap
- AWS secures the control plane; you secure access, nodes, pods, network, images and secrets.
- Private or restricted API, SSO roles, audit logs; KMS for secrets; Secrets Manager as the source of truth.
- Bottlerocket/AL2023, IMDSv2 hop limit 1, no SSH; Pod Security
restricted, admission policies, default-deny network policies. - Scan and sign images; detect with GuardDuty, CloudTrail and alerts on risky changes.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.