Cheat Sheets / Kubernetes & Platform

Kubernetes Administration cheat sheet

283 commands from every lesson of Kubernetes Administration — Level by Level, on one page.

01 · Set up your practice lab

Cluster lifecycle (kind)

kind create cluster --name lab --config kind-lab.yamlCreate the 3-node practice cluster
kind get clustersList your kind clusters
kind delete cluster --name labDelete it (start fresh any time)

First kubectl commands

kubectl versionClient and server versions
kubectl cluster-infoWhere the API server lives
kubectl get nodes -o wideNodes, their IPs, OS and runtime
kubectl config get-contextsWhich clusters kubectl knows about
kubectl config use-context kind-labPoint kubectl at the lab cluster

Make life easier

alias k=kubectlShorter commands
source <(kubectl completion bash)Tab completion in bash
kubectl explain pod.specBuilt-in docs for any field

02 · Architecture & the control plane

See the control plane

kubectl get pods -n kube-system -o wideControl-plane and system pods, and which node they run on
kubectl get --raw='/readyz?verbose'API server health checks, one line each
kubectl get events -A --sort-by=.lastTimestampRecent cluster events, newest last
kubectl api-resourcesEvery object type the API server knows

Look inside a node (kind)

docker exec -it lab-control-plane bashOpen a shell 'on' the node
ls /etc/kubernetes/manifestsStatic pod manifests for the control plane
crictl psContainers the kubelet is running
journalctl -u kubelet -fFollow kubelet logs

03 · Pods, Deployments & workload types

Pods

kubectl apply -f pod.yamlCreate or update from a file
kubectl get pods -o widePods with IP and node
kubectl describe pod <name>Details and events: first stop when debugging
kubectl logs <pod> [-c container] [--previous]Logs; --previous shows the crashed run
kubectl exec -it <pod> -- shShell inside the container
kubectl run tmp --rm -it --image=busybox:1.36 -- shThrowaway debug pod

Deployments

kubectl create deployment web --image=nginx:1.27Quick Deployment
kubectl scale deployment web --replicas=5Scale out or in
kubectl set image deployment/web nginx=nginx:1.27-alpineStart a rolling update
kubectl rollout status deployment/webWait for the rollout to finish
kubectl rollout history deployment/webPast revisions
kubectl rollout undo deployment/webRoll back to the previous revision
kubectl rollout restart deployment/webRestart all pods, one by one

Generate YAML instead of typing it

kubectl create deployment web --image=nginx:1.27 --dry-run=client -o yaml > web.yamlA starter manifest

04 · Services, DNS & basic networking

Services

kubectl expose deployment web --port=80Create a ClusterIP Service for a Deployment
kubectl get svc,endpointslices -l app=webService and the pod IPs behind it
kubectl describe svc webSelector, ports and endpoints in one view
kubectl port-forward svc/web 8080:80Reach a Service from your laptop

DNS & connectivity tests

kubectl run tmp --rm -it --image=busybox:1.36 --restart=Never -- shDebug shell inside the cluster
nslookup webResolve a Service name (inside the debug pod)
wget -qO- http://webCall the Service (inside the debug pod)
kubectl get pods -n kube-system -l k8s-app=kube-dnsAre the CoreDNS pods healthy?

DNS names

webSame namespace
web.shopService 'web' in namespace 'shop'
web.shop.svc.cluster.localFully qualified name

05 · Ingress: getting traffic in

Install & inspect

kubectl apply -f https://kind.sigs.k8s.io/examples/ingress/deploy-ingress-nginx.yamlingress-nginx for kind (cluster needs port mappings + ingress-ready label)
helm upgrade --install ingress-nginx ingress-nginx --repo https://kubernetes.github.io/ingress-nginx -n ingress-nginx --create-namespaceingress-nginx with Helm (other clusters)
kubectl get svc -n ingress-nginx ingress-nginx-controllerThe controller's Service: EXTERNAL-IP or NodePorts
kubectl get ingressclassWhich controllers exist (e.g. nginx)

Rules & debugging

kubectl get ingress -A / kubectl describe ingress shopRules, backends, address
kubectl get endpointslices -l kubernetes.io/service-name=shopPod IPs the controller will use
kubectl logs -n ingress-nginx deploy/ingress-nginx-controllerAccess log: status, upstream, latency
curl -H 'Host: shop.example.com' http://<ingress-ip>/Test a host rule without DNS

06 · Configuration & Secrets

ConfigMaps

kubectl create configmap app-config --from-literal=APP_COLOR=blueFrom key=value pairs
kubectl create configmap nginx-conf --from-file=default.confFrom a file (key = file name)
kubectl get configmap app-config -o yamlSee the contents

Secrets

kubectl create secret generic db-cred --from-literal=username=app --from-literal=password='S3cr3t!'Generic secret
kubectl create secret tls web-tls --cert=tls.crt --key=tls.keyTLS certificate + key
kubectl get secret db-cred -o jsonpath='{.data.password}' | base64 -dDecode one value (anyone with read access can)

Apply changes

kubectl rollout restart deployment/appPick up changed env vars
kubectl exec deploy/app -- env | grep APP_Check what a pod actually sees

07 · Storage basics

Inspect storage

kubectl get storageclassAvailable storage types (default marked)
kubectl get pvc,pvClaims and the volumes bound to them
kubectl describe pvc <name>Why a claim is Pending (events)

Access modes

ReadWriteOnce (RWO)Read-write by pods on ONE node
ReadOnlyMany (ROX)Read-only by many nodes
ReadWriteMany (RWX)Read-write by many nodes (NFS, CephFS, EFS)
ReadWriteOncePod (RWOP)Read-write by exactly one pod

Volume types

emptyDirScratch space; lives and dies with the pod
hostPathA folder on the node; avoid for apps
persistentVolumeClaimDurable storage that outlives pods
configMap / secretConfig files (lesson 05)

08 · Checkpoint: deploy a 3-tier app

Namespaces

kubectl create namespace shopCreate a namespace
kubectl config set-context --current --namespace=shopMake it your default for this context
kubectl get all -n shopDeployments, pods, Services in one list
kubectl delete namespace shopDelete it and everything inside

Everyday debugging

kubectl get pods -n shop -wWatch pods change state
kubectl -n shop describe pod <name>Events: scheduling, pulls, probes, mounts
kubectl -n shop logs deploy/webLogs from one pod of a Deployment
kubectl -n shop get endpointslicesWhich pods each Service really points to
kubectl -n shop exec deploy/postgres -- psql -U shop -d shop -c 'select 1'Run a command in a pod

09 · Build a cluster with kubeadm

Node preparation

sudo swapoff -aDisable swap now (also remove it from /etc/fstab)
sudo modprobe overlay && sudo modprobe br_netfilterKernel modules containers and bridged traffic need
containerd config default | sudo tee /etc/containerd/config.tomlWrite a default containerd config (then set SystemdCgroup = true)
sudo apt-mark hold kubelet kubeadm kubectlStop unattended upgrades changing versions

Build the cluster

sudo kubeadm init --control-plane-endpoint <LB-or-DNS>:6443 --pod-network-cidr 10.244.0.0/16 --upload-certsFirst control-plane node
kubeadm token create --print-join-commandNew worker join command (tokens expire after 24h)
sudo kubeadm join <endpoint>:6443 --token … --discovery-token-ca-cert-hash sha256:…Join a worker
sudo kubeadm resetUndo init/join on a node (then clean CNI config and iptables)

Certificates

sudo kubeadm certs check-expirationWhen each certificate expires
sudo kubeadm certs renew allRenew all kubeadm-managed certificates
ls /etc/kubernetes/pkiWhere the cluster PKI lives

10 · kubeadm at scale: immutable images

Image & boot

Base image: OS + containerd + kubelet/kubeadm/kubectl (pinned) + kernel settingsEverything that's the same on every node
Role at boot: cloud-init user-data (init / join / join --control-plane)One image, many roles
DHCP reservation (MAC → IP) or static network-configPredictable node IPs
cloud-init clean && truncate -s0 /etc/machine-idMake a template clonable (no duplicate identities)

kubeadm automation

kubeadm token generateCreate a bootstrap token value in advance
kubeadm init --config init.yaml --upload-certsFirst control plane from a config file
kubeadm init phase upload-certs --upload-certsRe-upload certs for control-plane joins (the key expires after 2 h)
kubeadm join --config join.yamlJoin as worker or control plane from a config file
openssl x509 -pubkey -in /etc/kubernetes/pki/ca.crt | openssl rsa -pubin -outform der 2>/dev/null | openssl dgst -sha256 -hexCompute the CA cert hash for discovery

11 · Bare-metal load balancing with MetalLB

Install & configure

helm repo add metallb https://metallb.github.io/metallb && helm install metallb metallb/metallb -n metallb-system --create-namespaceInstall MetalLB
IPAddressPool (spec.addresses: [10.0.0.240-10.0.0.250])Which IPs MetalLB may hand out
L2Advertisement (spec.ipAddressPools: [...])Announce those IPs with ARP/NDP
BGPPeer + BGPAdvertisementAnnounce them to routers with BGP

Check

kubectl get svc -A | grep LoadBalancerEXTERNAL-IP assigned?
kubectl get ipaddresspools,l2advertisements,bgppeers -n metallb-systemMetalLB config
kubectl logs -n metallb-system -l app.kubernetes.io/component=speakerWhich node announces what
arping -I eth0 10.0.0.240 (from another host on the segment)Who answers ARP for the VIP

12 · Upgrades & version skew

Before you start

kubectl versionCurrent client and server versions
kubectl get nodesEvery node's kubelet version
sudo kubeadm upgrade planWhat you can upgrade to, component by component
kubectl get pdb -APodDisruptionBudgets that could block drains

Control plane (first node)

sudo apt-mark unhold kubeadm && sudo apt-get install -y kubeadm='1.32.x-*' && sudo apt-mark hold kubeadmUpgrade the kubeadm binary first
sudo kubeadm upgrade apply v1.32.xUpgrade control-plane components (first control-plane node only)
sudo kubeadm upgrade nodeOther control-plane nodes, and every worker

Each node's kubelet

kubectl drain <node> --ignore-daemonsets --delete-emptydir-dataMove workloads off safely
sudo apt-get install -y kubelet='1.32.x-*' kubectl='1.32.x-*'Upgrade kubelet and kubectl (unhold/hold around it)
sudo systemctl daemon-reload && sudo systemctl restart kubeletRestart with the new version
kubectl uncordon <node>Allow scheduling again

13 · etcd backup & restore

Take and check a snapshot

ETCDCTL_API=3 etcdctl --endpoints=https://127.0.0.1:2379 --cacert=/etc/kubernetes/pki/etcd/ca.crt --cert=/etc/kubernetes/pki/etcd/server.crt --key=/etc/kubernetes/pki/etcd/server.key snapshot save /backup/etcd.dbSave a snapshot (kubeadm certificate paths)
etcdutl snapshot status /backup/etcd.db -w tableHash, revision, key count, size
etcdctl … endpoint status -w tableLeader, DB size, raft term per member
etcdctl … member list -w tableMembers of the etcd cluster

Restore (single control plane, kubeadm)

sudo mv /etc/kubernetes/manifests/kube-apiserver.yaml /root/Stop the API server (no writes during restore)
sudo etcdutl snapshot restore /backup/etcd.db --data-dir /var/lib/etcd-restoredRestore into a NEW data directory
sudo vi /etc/kubernetes/manifests/etcd.yamlPoint the etcd-data hostPath at /var/lib/etcd-restored
sudo mv /root/kube-apiserver.yaml /etc/kubernetes/manifests/Start the API server again

14 · Scheduling in depth

Labels & taints

kubectl label node w1 disktype=ssdAdd a node label
kubectl get nodes -L disktype,topology.kubernetes.io/zoneShow labels as columns
kubectl taint node w1 dedicated=gpu:NoScheduleRepel pods without a matching toleration
kubectl taint node w1 dedicated=gpu:NoSchedule-Remove that taint (trailing minus)
kubectl describe node w1 | grep -A3 TaintsSee a node's taints

Why is it Pending?

kubectl describe pod <pod>The FailedScheduling event lists each reason
kubectl get events --field-selector reason=FailedSchedulingAll scheduling failures
kubectl describe node <node> | grep -A8 'Allocated resources'How much CPU/memory is already requested

Priority

kubectl get priorityclassAvailable priorities (system-cluster-critical, …)

15 · RBAC & service accounts

Check permissions

kubectl auth can-i create deployments -n shopCan I do this?
kubectl auth can-i --list -n shopEverything I can do in a namespace
kubectl auth can-i get secrets -n shop --as=system:serviceaccount:shop:appTest as a service account
kubectl auth whoamiWho does the API server think I am?

Create RBAC quickly

kubectl create role pod-reader --verb=get,list,watch --resource=pods -n shopA namespaced Role
kubectl create rolebinding read-pods --role=pod-reader --user=asha -n shopBind it to a user
kubectl create clusterrolebinding ops-view --clusterrole=view --group=opsBuilt-in read-only role, cluster-wide, for a group
kubectl create serviceaccount app -n shopAn identity for a workload
kubectl create token app -n shop --duration=1hA short-lived token for that service account

Inspect

kubectl get roles,rolebindings -n shopNamespaced RBAC objects
kubectl describe clusterrole editWhat a built-in role allows

16 · Troubleshooting: nodes, pods, networking

Scope it

kubectl get nodesIs it one node or all of them?
kubectl get pods -A -o wide | grep -v RunningEverything not Running, and where
kubectl get events -A --sort-by=.lastTimestamp | tail -30What happened most recently

Pods

kubectl describe pod <pod>Events, state, last state, exit code
kubectl logs <pod> --previousLogs from the crashed container
kubectl debug -it <pod> --image=busybox:1.36 --target=<container>Ephemeral debug container sharing the pod's namespaces
kubectl get pod <pod> -o jsonpath='{.status.containerStatuses[*].lastState}'Why the last run ended

Nodes

kubectl describe node <node>Conditions: Ready, MemoryPressure, DiskPressure, PIDPressure
kubectl debug node/<node> -it --image=busybox:1.36Shell on the node (host filesystem at /host)
systemctl status kubelet containerdAre the node agents running? (on the node)
journalctl -u kubelet --since '15 min ago'Why the kubelet is unhappy (on the node)
crictl ps -aContainers the runtime knows about, even if the API is down

17 · Deploy an app end to end

Build it

kubectl create namespace shop-webA home for the app
kubectl apply -f deploy.yaml -f svc.yaml -f ingress.yaml -n shop-webDeployment → Service → Ingress
kubectl rollout status deploy/podinfo -n shop-webWait for a rollout
kubectl set image deploy/podinfo podinfo=ghcr.io/stefanprodan/podinfo:6.7.1 -n shop-webRoll out a new version
kubectl rollout undo deploy/podinfo -n shop-webRoll back

Debug hop by hop

kubectl get pods -n shop-web -o wide1. Pods Running and READY?
kubectl get endpointslices -n shop-web2. Service has endpoints?
kubectl run t --rm -it --image=busybox -- wget -qO- http://podinfo.shop-web3. Service works inside the cluster?
kubectl describe ingress podinfo -n shop-web4. Ingress rule, class, address?
curl -v http://podinfo.127.0.0.1.nip.io/5. From outside?

18 · Namespaces & an app with a database

Namespaces & DNS

<service>.<namespace>.svc.cluster.localFull DNS name of a Service
postgres.shop-data (from another namespace)Short form that works across namespaces
kubectl get all -n shop-dataWhat's in a namespace
Secrets and ConfigMaps are namespacedCreate them where the pods run

Database pieces

StatefulSet + volumeClaimTemplatesStable name (postgres-0) and its own PVC
Service with clusterIP: NoneHeadless: DNS for each StatefulSet pod
initContainer: pg_isready -h postgres.shop-dataWait for the DB before starting the app
NetworkPolicy (namespaceSelector + podSelector)Only the app may reach port 5432

19 · HTTPS & TLS termination

Certificates

openssl req -x509 -nodes -newkey rsa:2048 -days 30 -keyout tls.key -out tls.crt -subj '/CN=shop.127.0.0.1.nip.io' -addext 'subjectAltName=DNS:shop.127.0.0.1.nip.io'Self-signed cert for a lab
kubectl create secret tls shop-tls --cert=tls.crt --key=tls.key -n shop-webStore it as a TLS Secret
kubectl apply -f https://github.com/cert-manager/cert-manager/releases/download/v1.16.2/cert-manager.yamlInstall cert-manager (pick a current release)
kubectl get certificate,certificaterequest -Acert-manager status

ingress-nginx annotations

cert-manager.io/cluster-issuer: lab-caAsk cert-manager for this Ingress's certificate
nginx.ingress.kubernetes.io/ssl-redirect: "true"HTTP → HTTPS (default when TLS is set for the host)
nginx.ingress.kubernetes.io/ssl-passthrough: "true"Passthrough (controller needs --enable-ssl-passthrough)
nginx.ingress.kubernetes.io/backend-protocol: "HTTPS"Re-encrypt to the pod

20 · Ingress & egress gateways

Ingress gateway (Gateway API)

kubectl apply -f https://github.com/kubernetes-sigs/gateway-api/releases/download/v1.2.1/standard-install.yamlGateway API CRDs (pick a current release)
GatewayClass → Gateway (listeners) → HTTPRoute (rules)Platform owns Gateways, teams own Routes
kubectl get gatewayclass,gateway,httproute -AStatus of each layer

Egress control

NetworkPolicy podSelector: {} + policyTypes: [Egress]Default-deny egress for a namespace
Allow kube-dns on UDP/TCP 53Don't forget DNS
CiliumNetworkPolicy toFQDNsAllow by DNS name (Cilium)
CiliumEgressGatewayPolicy / Istio egress gatewayFixed exit point and source IP

21 · Release strategies: rolling, blue-green, canary

Rolling & blue-green

strategy.rollingUpdate: { maxSurge: 1, maxUnavailable: 0 }Rolling: capacity never drops
strategy.type: RecreateStop all old pods, then start new (downtime)
kubectl patch svc shop -p '{"spec":{"selector":{"app":"shop","version":"green"}}}'Blue-green: switch the Service

Canary

nginx.ingress.kubernetes.io/canary: "true" + canary-weight: "10"ingress-nginx canary Ingress
HTTPRoute backendRefs weights 90 / 10Gateway API canary
Argo Rollouts: steps [ setWeight: 20, pause, analysis ]Automated progressive delivery
kubectl argo rollouts promote|abort <name>Move forward or roll back

22 · HA control plane & load balancing

Join more control-plane nodes

sudo kubeadm init phase upload-certs --upload-certsRe-upload control-plane certs; prints a new certificate key (valid 2 hours)
kubeadm token create --print-join-commandBase join command (append --control-plane --certificate-key <key>)
sudo kubeadm join <VIP>:6443 --token … --discovery-token-ca-cert-hash … --control-plane --certificate-key …Join as a control-plane node

Check HA health

kubectl get nodes -l node-role.kubernetes.io/control-planeAll control-plane nodes
kubectl get lease -n kube-system kube-scheduler kube-controller-managerWhich node currently holds each leader lease
etcdctl … endpoint status --cluster -w tableetcd members, leader and DB size
curl -k https://<VIP>:6443/readyzIs the API reachable through the load balancer?

23 · Resource management & QoS

See usage vs requests

kubectl top nodesReal CPU/memory usage per node (needs metrics-server)
kubectl top pods -A --sort-by=memoryHungriest pods
kubectl describe node <node> | grep -A10 'Allocated resources'Sum of requests and limits on a node
kubectl get pod <pod> -o jsonpath='{.status.qosClass}'A pod's QoS class

Namespace guard-rails

kubectl get limitrange,resourcequota -n <ns>Defaults and totals for a namespace
kubectl describe resourcequota -n <ns>Used vs hard limits

Units

cpu: 500mHalf a CPU core (1000m = 1 core)
memory: 256Mi256 mebibytes (Mi/Gi are powers of 2; M/G are powers of 10)

24 · Autoscaling

HPA

kubectl autoscale deployment web --cpu-percent=60 --min=2 --max=10Quick CPU-based HPA
kubectl get hpa -wWatch current vs target and replica count
kubectl describe hpa webConditions and scaling events (why it did or didn't scale)

Which autoscaler?

HPAMore or fewer pods, from CPU/memory/custom metrics
VPABigger or smaller requests per pod, from observed usage
Cluster AutoscalerMore or fewer nodes in existing node groups when pods are Pending
KarpenterLaunches right-sized nodes directly for Pending pods; consolidates
KEDAScales on events (queue length, Kafka lag, cron), including to zero

25 · Cluster lifecycle at scale

Fleet inventory

kubectl config get-contexts -o nameEvery cluster in your kubeconfig
kubectl --context <ctx> version -o json | jq -r .serverVersion.gitVersionOne cluster's server version
kubectl --context <ctx> get nodes -o custom-columns=NAME:.metadata.name,KUBELET:.status.nodeInfo.kubeletVersionKubelet versions per node

Pre-flight checks (per cluster)

kubectl get nodes | grep -v ' Ready'Anything not Ready?
kubectl get pods -A | grep -vE 'Running|Completed'Unhealthy pods before you start
kubectl get pdb -ABudgets that could block drains
kubectl get --raw /metrics | grep apiserver_requested_deprecated_apisDeprecated API usage
sudo kubeadm certs check-expirationCertificate expiry (kubeadm clusters)

26 · Performance tuning

Health and latency metrics

etcd_disk_wal_fsync_duration_secondsetcd write-ahead-log fsync latency (keep p99 low, ~10ms)
etcd_disk_backend_commit_duration_secondsetcd commit latency
apiserver_request_duration_secondsAPI latency by verb and resource
apiserver_flowcontrol_rejected_requests_totalRequests rejected by API Priority and Fairness

Inspect

kubectl get flowschemas,prioritylevelconfigurationsHow the API server prioritises clients
kubectl get --raw /metrics | grep apiserver_request_totalRequest counts by verb/resource/code
sudo cat /var/lib/kubelet/config.yamlThe kubelet's configuration (kubeadm nodes)

27 · Disaster recovery

Velero

velero backup create shop-$(date +%F) --include-namespaces shopBack up one namespace (objects + volumes as configured)
velero backup describe <backup> --detailsWhat was captured, warnings and errors
velero restore create --from-backup <backup>Restore it
velero schedule create daily-shop --schedule='0 2 * * *' --include-namespaces shop --ttl 720hNightly backup kept for 30 days
velero backup-location getIs the object storage target available?

Words to agree on

RPORecovery Point Objective: how much data you can afford to lose (time)
RTORecovery Time Objective: how long until service is back

28 · Multi-cluster & fleet management

Cluster API (clusterctl)

clusterctl init --infrastructure <provider>Turn a cluster into a management cluster
clusterctl generate cluster <name> --kubernetes-version <v> > cluster.yamlGenerate a workload cluster definition
kubectl apply -f cluster.yamlCreate it (declaratively)
clusterctl describe cluster <name>Tree view of the cluster's machines and conditions
clusterctl get kubeconfig <name> > <name>.kubeconfigGet access to the new cluster
kubectl get clusters,machinedeployments,machinesFleet objects in the management cluster

Fleet GitOps

argocd cluster add <context>Register a cluster with Argo CD
kubectl get applicationsets -n argocdTemplates that fan out apps to clusters

29 · Multi-tenancy models

Tenant kit (per namespace)

kubectl create namespace team-aThe tenant boundary
kubectl label ns team-a pod-security.kubernetes.io/enforce=restrictedEnforce the 'restricted' Pod Security Standard
kubectl create rolebinding team-a-edit --clusterrole=edit --group=team-a -n team-aThe team may manage its own namespace
kubectl apply -f quota.yaml -f limitrange.yaml -f netpol-default-deny.yamlBudget, defaults, and network isolation

Verify isolation

kubectl auth can-i list secrets -n team-b --as-group=team-a --as=someoneCan team A read team B's secrets? (should be no)
kubectl get networkpolicy -AWhich namespaces are isolated

30 · Internal Developer Platform design

Platform building blocks

Golden pathAn opinionated, supported way to build and ship a service end to end
Platform APIA simple, declarative interface (CRD, Terraform module, template) hiding infrastructure detail
Developer portalA catalog of services, owners, docs and self-service actions (e.g. Backstage)
Guard-railsPolicies that prevent unsafe choices automatically, instead of manual approvals

Measure it

Deployment frequency / lead timeDORA: how fast changes reach production
Change failure rate / time to restoreDORA: how safely
Time to first deploy (new service)How long a golden path takes, end to end
Developer satisfactionSurvey the platform's users regularly

31 · SLOs for the platform itself

Definitions

SLIA measured ratio of good events: good ÷ total
SLOThe target for that SLI over a window, e.g. 99.9% over 30 days
Error budgetWhat the SLO allows to fail: 0.1% of 30 days ≈ 43 minutes
Burn rateHow fast the budget is being used (1 = exactly on budget)

Example PromQL

sum(rate(apiserver_request_total{code=~"5.."}[5m])) / sum(rate(apiserver_request_total[5m]))API server error ratio
histogram_quantile(0.99, sum(rate(apiserver_request_duration_seconds_bucket{verb!~"WATCH|CONNECT"}[5m])) by (le))p99 API latency (excluding long-lived watches)

32 · Capstone: platform design review

Design document outline

1. Context & requirementsWho, what, scale, constraints, SLOs, RPO/RTO
2. Cluster topologyHow many clusters, where, why (with an ADR)
3. Tenancy & accessModel, tenant kit, identity, RBAC
4. LifecycleProvisioning, upgrades, waves, add-ons
5. Resilience & DRFailure domains, backups, drills, runbooks
6. Observability & SLOsSLIs, SLOs, probes, alerting
7. Cost & risksEstimate, top risks, open questions

Questions reviewers always ask

What happens when X fails?For every component and dependency
How do you upgrade it?Without downtime, repeatedly, for years
How do you know it's working?SLIs, probes, alerts
What does it cost, and what could we drop?Trade-offs made explicit