Lesson 06 of 18 · Part 2 — Build the platform
IAM & governance: access entries, Pod Identity, IRSA, policies
Everything about who can do what on an EKS platform: IAM roles and policies, the cluster and node roles, access entries and RBAC for people and pipelines, Pod Identity and IRSA for workloads, and the governance guard-rails (quotas, policies, tagging, audit, break-glass) that keep it safe as teams grow.
Three questions, three mechanisms
Almost every EKS security discussion comes down to one of three questions. Keep them separate and the design stays clear:
| Question | Mechanism | Lives in |
|---|---|---|
| Which AWS identities may use the Kubernetes API, and with which Kubernetes permissions? | Access entries + access policies or Kubernetes RBAC | EKS API + cluster RBAC |
| Which AWS permissions may a pod use? | EKS Pod Identity (or IRSA) | IAM + EKS |
| What are teams allowed to consume and deploy at all? | Governance: quotas, admission policies, Pod Security, network policy, tagging, audit | Organizations, IAM, Kubernetes admission |
A hospital. The staff badge says who may enter which ward (access entries). The medicine cabinet key each nurse carries opens only the cabinets for their patients (pod roles). The hospital rules say how many beds each ward gets, which medicines are allowed and that every cabinet opening is logged (governance). A badge that opens every door and every cabinet is convenient right up until it's stolen.
IAM building blocks, in one table
| Concept | What it does | EKS example |
|---|---|---|
| Role | An identity with no long-term credentials; assumed for temporary credentials | Cluster role, node role, pod roles, CI role, admin SSO roles |
| Trust policy | Who may assume the role | pods.eks.amazonaws.com (Pod Identity), the cluster's OIDC provider (IRSA), ec2.amazonaws.com (nodes) |
| Permissions policy | What the role may do | s3:GetObject on one prefix |
| AWS-managed policy | Maintained by AWS | AmazonEKSClusterPolicy, AmazonEKSWorkerNodePolicy, AmazonEKS_CNI_Policy |
| Permissions boundary | Maximum permissions a role can ever have | Let teams create pod roles, but never beyond "their" S3 and SQS |
| SCP (Organizations) | Maximum permissions for a whole account | Deny leaving the Organization, deny unapproved regions |
| Condition keys / tags | Fine-grained rules | Only resources tagged with the pod's namespace (ABAC) |
Policy evaluation, simplified: an explicit deny anywhere wins; otherwise the action must be allowed by the identity policy and not blocked by the boundary or SCP.
The platform's own roles
| Role | Trusted by | Typical policies | Keep in mind |
|---|---|---|---|
| Cluster role | eks.amazonaws.com |
AmazonEKSClusterPolicy (Auto Mode adds more) |
Used by EKS itself, not by people |
| Node role | ec2.amazonaws.com |
AmazonEKSWorkerNodePolicy, ECR pull (AmazonEC2ContainerRegistryPullOnly or ...ReadOnly), AmazonSSMManagedInstanceCore |
Every pod on the node can potentially reach it, so keep it minimal |
| VPC CNI role | Pod Identity / IRSA | AmazonEKS_CNI_Policy |
Give the CNI its own role rather than putting its policy on the node role |
| Controller roles | Pod Identity / IRSA | Load Balancer Controller, EBS CSI, Karpenter, ExternalDNS policies | One role per controller, scoped by tags where possible |
People and pipelines: access entries
Set the cluster's authentication mode to API for new clusters. API_AND_CONFIG_MAP exists for migrating from the legacy aws-auth ConfigMap; the switch to API is one-way.
An access entry says "this IAM principal may use this cluster". What it may then do comes from either:
- Access policies, AWS-managed permission sets, scoped to the cluster or to namespaces:
| Access policy | Roughly equivalent to |
|---|---|
AmazonEKSClusterAdminPolicy |
cluster-admin |
AmazonEKSAdminPolicy |
admin (namespace-scopable) |
AmazonEKSEditPolicy |
edit (namespace-scopable) |
AmazonEKSViewPolicy |
view (namespace-scopable) |
- or Kubernetes groups on the entry, bound by your own
RoleBindings, for anything the managed policies don't express.
$ aws eks create-access-entry --cluster-name prod \
--principal-arn arn:aws:iam::111122223333:role/aws-reserved/sso.amazonaws.com/AWSReservedSSO_ShopDev_abc123
$ aws eks associate-access-policy --cluster-name prod \
--principal-arn arn:aws:iam::111122223333:role/aws-reserved/sso.amazonaws.com/AWSReservedSSO_ShopDev_abc123 \
--policy-arn arn:aws:eks::aws:cluster-access-policy/AmazonEKSEditPolicy \
--access-scope type=namespace,namespaces=shop
Rules that keep this clean:
- Map roles, never IAM users. People arrive through IAM Identity Center permission sets; CI through its own OIDC-federated role.
- Create the cluster with
bootstrap_cluster_creator_admin_permissionsoff (in Terraform) and grant admin explicitly, so admin access doesn't silently belong to whichever identity happened to runcreate. - Keep one break-glass role with cluster-admin, protected by MFA, alarmed on use, and tested twice a year.
Workloads: Pod Identity
EKS Pod Identity gives a pod the credentials of an IAM role, chosen per namespace + ServiceAccount:
- Install the
eks-pod-identity-agentmanaged add-on (a DaemonSet on each node). - Create a role whose trust policy allows the Pod Identity service principal:
{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Principal": { "Service": "pods.eks.amazonaws.com" },
"Action": ["sts:AssumeRole", "sts:TagSession"]
}]
}
- Attach a least-privilege permissions policy:
{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Action": ["s3:GetObject"],
"Resource": "arn:aws:s3:::shop-assets-prod/orders/*"
}]
}
- Associate the role with the ServiceAccount:
$ aws eks create-pod-identity-association --cluster-name prod \
--namespace shop --service-account orders \
--role-arn arn:aws:iam::111122223333:role/shop-orders
Pods with serviceAccountName: orders now get that role's temporary credentials; AWS SDKs find them automatically. No annotations, no per-cluster OIDC setup.
One policy for many teams: session tags (ABAC)
Pod Identity adds session tags such as kubernetes-namespace, kubernetes-service-account and eks-cluster-name. Policies can use them, so one role and policy can serve every namespace safely:
{
"Effect": "Allow",
"Action": ["s3:GetObject", "s3:PutObject"],
"Resource": "arn:aws:s3:::team-data-prod/${aws:PrincipalTag/kubernetes-namespace}/*"
}
A pod in namespace shop can only touch team-data-prod/shop/*.
Workloads: IRSA
IAM Roles for Service Accounts is the older mechanism, still everywhere:
- Register the cluster's OIDC issuer as an IAM identity provider.
- The role's trust policy allows
sts:AssumeRoleWithWebIdentityfor exactly one ServiceAccount:
{
"Effect": "Allow",
"Principal": { "Federated": "arn:aws:iam::111122223333:oidc-provider/oidc.eks.eu-west-1.amazonaws.com/id/EXAMPLE123" },
"Action": "sts:AssumeRoleWithWebIdentity",
"Condition": {
"StringEquals": {
"oidc.eks.eu-west-1.amazonaws.com/id/EXAMPLE123:sub": "system:serviceaccount:shop:orders",
"oidc.eks.eu-west-1.amazonaws.com/id/EXAMPLE123:aud": "sts.amazonaws.com"
}
}
}
- Annotate the ServiceAccount:
eks.amazonaws.com/role-arn: arn:aws:iam::111122223333:role/shop-orders.
| Pod Identity | IRSA | |
|---|---|---|
| Trust policy | One fixed service principal; reusable across clusters | Per-cluster OIDC provider in every trust policy |
| Setup per cluster | Agent add-on | OIDC provider registration |
| Linking | Association via EKS API | Annotation on the ServiceAccount |
| ABAC session tags | Built in | Not built in |
| Where it works | EKS pods on EC2 nodes (not Fargate at the time of writing) | Also Fargate pods and self-managed Kubernetes clusters |
| Use for | New EKS workloads | Existing setups, non-EKS clusters, edge cases |
Always check which role a pod actually gets (aws sts get-caller-identity from inside it). Mistakes here are silent.
Close the back door: the node role
If a pod can reach the instance metadata service, it can use the node role, which is shared by every pod on that node. Enforce IMDSv2 with a hop limit of 1 in the node launch template (the default on EKS Auto Mode and many current templates, but verify), keep the node role minimal, and move the CNI's permissions off it.
Governance: guard-rails that scale with teams
Identity decides who; governance bounds what and how much.
| Guard-rail | Tool | What it prevents |
|---|---|---|
| Account boundaries | Separate accounts, SCPs | Dev mistakes reaching prod; unapproved regions and services |
| Namespace per team/app | Namespaces + access scoped to them | Teams touching each other's workloads |
| Quotas | ResourceQuota, LimitRange |
One team consuming the cluster; pods without requests/limits |
| Admission policies | Kyverno or OPA Gatekeeper | Unapproved registries, latest tags, missing labels, privileged pods |
| Pod Security | Pod Security Admission (restricted / baseline) |
Privileged containers, host mounts, root users |
| Network policy | VPC CNI network policy (or Cilium/Calico) | Any pod talking to any pod |
| Tagging & cost | Required tags, cost allocation, split cost data per namespace | Unknown spend; no chargeback |
| Service quotas | Service Quotas + alarms | Scaling blocked by vCPU, IP, ELB or EKS limits at the worst moment |
| Audit | CloudTrail, EKS audit logs, GuardDuty | Unexplained changes; undetected misuse |
| Access reviews | list-access-entries, IAM Access Analyzer |
Access that outlived its reason |
Example: a namespace that can't be abused by accident:
apiVersion: v1
kind: Namespace
metadata:
name: shop
labels:
pod-security.kubernetes.io/enforce: restricted
team: shop
cost-centre: cc-1042
---
apiVersion: v1
kind: ResourceQuota
metadata: { name: shop-quota, namespace: shop }
spec:
hard:
requests.cpu: "40"
requests.memory: 80Gi
limits.memory: 120Gi
persistentvolumeclaims: "20"
services.loadbalancer: "2"
---
apiVersion: v1
kind: LimitRange
metadata: { name: shop-defaults, namespace: shop }
spec:
limits:
- type: Container
default: { cpu: 500m, memory: 512Mi }
defaultRequest: { cpu: 100m, memory: 128Mi }
Try it: least privilege end to end (sandbox)
- Create a cluster with authentication mode
API. Grant a second role view on namespaceshoponly; assume it and confirmkubectl get pods -n shopworks and-n kube-systemis forbidden. - Install the Pod Identity Agent add-on. Create a role that can read only
s3://<bucket>/shop/*and associate it withshop/orders. - Run the
amazon/aws-clipod from the cheat sheet withserviceAccountName: orders: check the role, read an allowed and a denied object. - Run the same pod with the
defaultServiceAccount: it must get no credentials and must not reach169.254.169.254. - Apply the namespace, quota and limit range above, then try to create a pod without limits and a privileged pod. Note what's defaulted and what's rejected.
Going deeper: delegating safely to many teams
At scale the platform team can't hand-craft every pod role. A proven pattern: teams create their own roles through a pipeline, but every role must carry a platform-owned permissions boundary and a naming prefix, and a Kyverno policy only allows ServiceAccounts in a namespace to be associated with roles carrying that namespace's tag. Combine with ABAC session tags so a single policy template serves every namespace, and review drift with IAM Access Analyzer's unused-access findings.
Recap
- Keep three things separate: cluster access (access entries + RBAC), pod permissions (Pod Identity / IRSA), governance (quotas, admission, Pod Security, network policy, tags, audit).
- People via SSO roles, CI via OIDC roles, never IAM users or shared kubeconfigs; keep a tested break-glass role.
- Pod Identity for new workloads (fixed service principal, associations, ABAC tags); IRSA where needed; always verify the effective role.
- Minimal node role, IMDSv2 hop limit 1, CNI and controllers on their own roles.
- Permissions boundaries and SCPs make delegation safe as the platform grows.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.