Production EKS Platform — From Zero to Production›06 · IAM & governance: access entries, Pod Identity, IRSA, policies

Lesson 06 of 18 · Part 2 — Build the platform

IAM & governance: access entries, Pod Identity, IRSA, policies

Everything about who can do what on an EKS platform: IAM roles and policies, the cluster and node roles, access entries and RBAC for people and pipelines, Pod Identity and IRSA for workloads, and the governance guard-rails (quotas, policies, tagging, audit, break-glass) that keep it safe as teams grow.

Intermediate → Architect
Key wordsIAM rolesIAM policiestrust policypermission boundarySCPcluster rolenode roleaccess entriesaccess policiesauthentication modeaws-authKubernetes RBACEKS Pod IdentityIRSAOIDCsession tagsABACIMDSv2ResourceQuotaKyvernogovernancebreak-glass
People and pipelines -> Kubernetes API SSO role / CI role IAM Identity Center, OIDC from CI access entry + access policy or Kubernetes groups RBAC / namespace scope Pods -> AWS APIs ServiceAccount shop/orders + association IAM role Pod Identity (or IRSA) least-privilege policy Governance around both SCPs & boundaries max permissions Quotas & limits per namespace Admission policies Pod Security, Kyverno Audit CloudTrail, EKS logs node role stays minimal; IMDSv2 hop limit 1 keeps pods off it
People and pipelines reach the Kubernetes API through access entries; pods reach AWS through their own roles; governance bounds both.

Three questions, three mechanisms

Almost every EKS security discussion comes down to one of three questions. Keep them separate and the design stays clear:

Question Mechanism Lives in
Which AWS identities may use the Kubernetes API, and with which Kubernetes permissions? Access entries + access policies or Kubernetes RBAC EKS API + cluster RBAC
Which AWS permissions may a pod use? EKS Pod Identity (or IRSA) IAM + EKS
What are teams allowed to consume and deploy at all? Governance: quotas, admission policies, Pod Security, network policy, tagging, audit Organizations, IAM, Kubernetes admission

A hospital. The staff badge says who may enter which ward (access entries). The medicine cabinet key each nurse carries opens only the cabinets for their patients (pod roles). The hospital rules say how many beds each ward gets, which medicines are allowed and that every cabinet opening is logged (governance). A badge that opens every door and every cabinet is convenient right up until it's stolen.

IAM building blocks, in one table

Concept What it does EKS example
Role An identity with no long-term credentials; assumed for temporary credentials Cluster role, node role, pod roles, CI role, admin SSO roles
Trust policy Who may assume the role pods.eks.amazonaws.com (Pod Identity), the cluster's OIDC provider (IRSA), ec2.amazonaws.com (nodes)
Permissions policy What the role may do s3:GetObject on one prefix
AWS-managed policy Maintained by AWS AmazonEKSClusterPolicy, AmazonEKSWorkerNodePolicy, AmazonEKS_CNI_Policy
Permissions boundary Maximum permissions a role can ever have Let teams create pod roles, but never beyond "their" S3 and SQS
SCP (Organizations) Maximum permissions for a whole account Deny leaving the Organization, deny unapproved regions
Condition keys / tags Fine-grained rules Only resources tagged with the pod's namespace (ABAC)

Policy evaluation, simplified: an explicit deny anywhere wins; otherwise the action must be allowed by the identity policy and not blocked by the boundary or SCP.

The platform's own roles

Role Trusted by Typical policies Keep in mind
Cluster role eks.amazonaws.com AmazonEKSClusterPolicy (Auto Mode adds more) Used by EKS itself, not by people
Node role ec2.amazonaws.com AmazonEKSWorkerNodePolicy, ECR pull (AmazonEC2ContainerRegistryPullOnly or ...ReadOnly), AmazonSSMManagedInstanceCore Every pod on the node can potentially reach it, so keep it minimal
VPC CNI role Pod Identity / IRSA AmazonEKS_CNI_Policy Give the CNI its own role rather than putting its policy on the node role
Controller roles Pod Identity / IRSA Load Balancer Controller, EBS CSI, Karpenter, ExternalDNS policies One role per controller, scoped by tags where possible

People and pipelines: access entries

Set the cluster's authentication mode to API for new clusters. API_AND_CONFIG_MAP exists for migrating from the legacy aws-auth ConfigMap; the switch to API is one-way.

An access entry says "this IAM principal may use this cluster". What it may then do comes from either:

  • Access policies, AWS-managed permission sets, scoped to the cluster or to namespaces:
Access policy Roughly equivalent to
AmazonEKSClusterAdminPolicy cluster-admin
AmazonEKSAdminPolicy admin (namespace-scopable)
AmazonEKSEditPolicy edit (namespace-scopable)
AmazonEKSViewPolicy view (namespace-scopable)
  • or Kubernetes groups on the entry, bound by your own RoleBindings, for anything the managed policies don't express.
$ aws eks create-access-entry --cluster-name prod \
    --principal-arn arn:aws:iam::111122223333:role/aws-reserved/sso.amazonaws.com/AWSReservedSSO_ShopDev_abc123
$ aws eks associate-access-policy --cluster-name prod \
    --principal-arn arn:aws:iam::111122223333:role/aws-reserved/sso.amazonaws.com/AWSReservedSSO_ShopDev_abc123 \
    --policy-arn arn:aws:eks::aws:cluster-access-policy/AmazonEKSEditPolicy \
    --access-scope type=namespace,namespaces=shop

Rules that keep this clean:

  • Map roles, never IAM users. People arrive through IAM Identity Center permission sets; CI through its own OIDC-federated role.
  • Create the cluster with bootstrap_cluster_creator_admin_permissions off (in Terraform) and grant admin explicitly, so admin access doesn't silently belong to whichever identity happened to run create.
  • Keep one break-glass role with cluster-admin, protected by MFA, alarmed on use, and tested twice a year.

Workloads: Pod Identity

EKS Pod Identity gives a pod the credentials of an IAM role, chosen per namespace + ServiceAccount:

  1. Install the eks-pod-identity-agent managed add-on (a DaemonSet on each node).
  2. Create a role whose trust policy allows the Pod Identity service principal:
{
  "Version": "2012-10-17",
  "Statement": [{
    "Effect": "Allow",
    "Principal": { "Service": "pods.eks.amazonaws.com" },
    "Action": ["sts:AssumeRole", "sts:TagSession"]
  }]
}
  1. Attach a least-privilege permissions policy:
{
  "Version": "2012-10-17",
  "Statement": [{
    "Effect": "Allow",
    "Action": ["s3:GetObject"],
    "Resource": "arn:aws:s3:::shop-assets-prod/orders/*"
  }]
}
  1. Associate the role with the ServiceAccount:
$ aws eks create-pod-identity-association --cluster-name prod \
    --namespace shop --service-account orders \
    --role-arn arn:aws:iam::111122223333:role/shop-orders

Pods with serviceAccountName: orders now get that role's temporary credentials; AWS SDKs find them automatically. No annotations, no per-cluster OIDC setup.

One policy for many teams: session tags (ABAC)

Pod Identity adds session tags such as kubernetes-namespace, kubernetes-service-account and eks-cluster-name. Policies can use them, so one role and policy can serve every namespace safely:

{
  "Effect": "Allow",
  "Action": ["s3:GetObject", "s3:PutObject"],
  "Resource": "arn:aws:s3:::team-data-prod/${aws:PrincipalTag/kubernetes-namespace}/*"
}

A pod in namespace shop can only touch team-data-prod/shop/*.

Workloads: IRSA

IAM Roles for Service Accounts is the older mechanism, still everywhere:

  1. Register the cluster's OIDC issuer as an IAM identity provider.
  2. The role's trust policy allows sts:AssumeRoleWithWebIdentity for exactly one ServiceAccount:
{
  "Effect": "Allow",
  "Principal": { "Federated": "arn:aws:iam::111122223333:oidc-provider/oidc.eks.eu-west-1.amazonaws.com/id/EXAMPLE123" },
  "Action": "sts:AssumeRoleWithWebIdentity",
  "Condition": {
    "StringEquals": {
      "oidc.eks.eu-west-1.amazonaws.com/id/EXAMPLE123:sub": "system:serviceaccount:shop:orders",
      "oidc.eks.eu-west-1.amazonaws.com/id/EXAMPLE123:aud": "sts.amazonaws.com"
    }
  }
}
  1. Annotate the ServiceAccount: eks.amazonaws.com/role-arn: arn:aws:iam::111122223333:role/shop-orders.
Pod Identity IRSA
Trust policy One fixed service principal; reusable across clusters Per-cluster OIDC provider in every trust policy
Setup per cluster Agent add-on OIDC provider registration
Linking Association via EKS API Annotation on the ServiceAccount
ABAC session tags Built in Not built in
Where it works EKS pods on EC2 nodes (not Fargate at the time of writing) Also Fargate pods and self-managed Kubernetes clusters
Use for New EKS workloads Existing setups, non-EKS clusters, edge cases

Always check which role a pod actually gets (aws sts get-caller-identity from inside it). Mistakes here are silent.

Close the back door: the node role

If a pod can reach the instance metadata service, it can use the node role, which is shared by every pod on that node. Enforce IMDSv2 with a hop limit of 1 in the node launch template (the default on EKS Auto Mode and many current templates, but verify), keep the node role minimal, and move the CNI's permissions off it.

Governance: guard-rails that scale with teams

Identity decides who; governance bounds what and how much.

Guard-rail Tool What it prevents
Account boundaries Separate accounts, SCPs Dev mistakes reaching prod; unapproved regions and services
Namespace per team/app Namespaces + access scoped to them Teams touching each other's workloads
Quotas ResourceQuota, LimitRange One team consuming the cluster; pods without requests/limits
Admission policies Kyverno or OPA Gatekeeper Unapproved registries, latest tags, missing labels, privileged pods
Pod Security Pod Security Admission (restricted / baseline) Privileged containers, host mounts, root users
Network policy VPC CNI network policy (or Cilium/Calico) Any pod talking to any pod
Tagging & cost Required tags, cost allocation, split cost data per namespace Unknown spend; no chargeback
Service quotas Service Quotas + alarms Scaling blocked by vCPU, IP, ELB or EKS limits at the worst moment
Audit CloudTrail, EKS audit logs, GuardDuty Unexplained changes; undetected misuse
Access reviews list-access-entries, IAM Access Analyzer Access that outlived its reason

Example: a namespace that can't be abused by accident:

apiVersion: v1
kind: Namespace
metadata:
  name: shop
  labels:
    pod-security.kubernetes.io/enforce: restricted
    team: shop
    cost-centre: cc-1042
---
apiVersion: v1
kind: ResourceQuota
metadata: { name: shop-quota, namespace: shop }
spec:
  hard:
    requests.cpu: "40"
    requests.memory: 80Gi
    limits.memory: 120Gi
    persistentvolumeclaims: "20"
    services.loadbalancer: "2"
---
apiVersion: v1
kind: LimitRange
metadata: { name: shop-defaults, namespace: shop }
spec:
  limits:
    - type: Container
      default: { cpu: 500m, memory: 512Mi }
      defaultRequest: { cpu: 100m, memory: 128Mi }

Try it: least privilege end to end (sandbox)

  1. Create a cluster with authentication mode API. Grant a second role view on namespace shop only; assume it and confirm kubectl get pods -n shop works and -n kube-system is forbidden.
  2. Install the Pod Identity Agent add-on. Create a role that can read only s3://<bucket>/shop/* and associate it with shop/orders.
  3. Run the amazon/aws-cli pod from the cheat sheet with serviceAccountName: orders: check the role, read an allowed and a denied object.
  4. Run the same pod with the default ServiceAccount: it must get no credentials and must not reach 169.254.169.254.
  5. Apply the namespace, quota and limit range above, then try to create a pod without limits and a privileged pod. Note what's defaulted and what's rejected.

Going deeper: delegating safely to many teams

At scale the platform team can't hand-craft every pod role. A proven pattern: teams create their own roles through a pipeline, but every role must carry a platform-owned permissions boundary and a naming prefix, and a Kyverno policy only allows ServiceAccounts in a namespace to be associated with roles carrying that namespace's tag. Combine with ABAC session tags so a single policy template serves every namespace, and review drift with IAM Access Analyzer's unused-access findings.

Recap

  • Keep three things separate: cluster access (access entries + RBAC), pod permissions (Pod Identity / IRSA), governance (quotas, admission, Pod Security, network policy, tags, audit).
  • People via SSO roles, CI via OIDC roles, never IAM users or shared kubeconfigs; keep a tested break-glass role.
  • Pod Identity for new workloads (fixed service principal, associations, ABAC tags); IRSA where needed; always verify the effective role.
  • Minimal node role, IMDSv2 hop limit 1, CNI and controllers on their own roles.
  • Permissions boundaries and SCPs make delegation safe as the platform grows.

This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.