Kubernetes Administration cheat sheet
283 commands from every lesson of Kubernetes Administration — Level by Level, on one page.
Cluster lifecycle (kind)
kind create cluster --name lab --config kind-lab.yaml | Create the 3-node practice cluster |
kind get clusters | List your kind clusters |
kind delete cluster --name lab | Delete it (start fresh any time) |
First kubectl commands
kubectl version | Client and server versions |
kubectl cluster-info | Where the API server lives |
kubectl get nodes -o wide | Nodes, their IPs, OS and runtime |
kubectl config get-contexts | Which clusters kubectl knows about |
kubectl config use-context kind-lab | Point kubectl at the lab cluster |
Make life easier
alias k=kubectl | Shorter commands |
source <(kubectl completion bash) | Tab completion in bash |
kubectl explain pod.spec | Built-in docs for any field |
See the control plane
kubectl get pods -n kube-system -o wide | Control-plane and system pods, and which node they run on |
kubectl get --raw='/readyz?verbose' | API server health checks, one line each |
kubectl get events -A --sort-by=.lastTimestamp | Recent cluster events, newest last |
kubectl api-resources | Every object type the API server knows |
Look inside a node (kind)
docker exec -it lab-control-plane bash | Open a shell 'on' the node |
ls /etc/kubernetes/manifests | Static pod manifests for the control plane |
crictl ps | Containers the kubelet is running |
journalctl -u kubelet -f | Follow kubelet logs |
Pods
kubectl apply -f pod.yaml | Create or update from a file |
kubectl get pods -o wide | Pods with IP and node |
kubectl describe pod <name> | Details and events: first stop when debugging |
kubectl logs <pod> [-c container] [--previous] | Logs; --previous shows the crashed run |
kubectl exec -it <pod> -- sh | Shell inside the container |
kubectl run tmp --rm -it --image=busybox:1.36 -- sh | Throwaway debug pod |
Deployments
kubectl create deployment web --image=nginx:1.27 | Quick Deployment |
kubectl scale deployment web --replicas=5 | Scale out or in |
kubectl set image deployment/web nginx=nginx:1.27-alpine | Start a rolling update |
kubectl rollout status deployment/web | Wait for the rollout to finish |
kubectl rollout history deployment/web | Past revisions |
kubectl rollout undo deployment/web | Roll back to the previous revision |
kubectl rollout restart deployment/web | Restart all pods, one by one |
Generate YAML instead of typing it
kubectl create deployment web --image=nginx:1.27 --dry-run=client -o yaml > web.yaml | A starter manifest |
Services
kubectl expose deployment web --port=80 | Create a ClusterIP Service for a Deployment |
kubectl get svc,endpointslices -l app=web | Service and the pod IPs behind it |
kubectl describe svc web | Selector, ports and endpoints in one view |
kubectl port-forward svc/web 8080:80 | Reach a Service from your laptop |
DNS & connectivity tests
kubectl run tmp --rm -it --image=busybox:1.36 --restart=Never -- sh | Debug shell inside the cluster |
nslookup web | Resolve a Service name (inside the debug pod) |
wget -qO- http://web | Call the Service (inside the debug pod) |
kubectl get pods -n kube-system -l k8s-app=kube-dns | Are the CoreDNS pods healthy? |
DNS names
web | Same namespace |
web.shop | Service 'web' in namespace 'shop' |
web.shop.svc.cluster.local | Fully qualified name |
Install & inspect
kubectl apply -f https://kind.sigs.k8s.io/examples/ingress/deploy-ingress-nginx.yaml | ingress-nginx for kind (cluster needs port mappings + ingress-ready label) |
helm upgrade --install ingress-nginx ingress-nginx --repo https://kubernetes.github.io/ingress-nginx -n ingress-nginx --create-namespace | ingress-nginx with Helm (other clusters) |
kubectl get svc -n ingress-nginx ingress-nginx-controller | The controller's Service: EXTERNAL-IP or NodePorts |
kubectl get ingressclass | Which controllers exist (e.g. nginx) |
Rules & debugging
kubectl get ingress -A / kubectl describe ingress shop | Rules, backends, address |
kubectl get endpointslices -l kubernetes.io/service-name=shop | Pod IPs the controller will use |
kubectl logs -n ingress-nginx deploy/ingress-nginx-controller | Access log: status, upstream, latency |
curl -H 'Host: shop.example.com' http://<ingress-ip>/ | Test a host rule without DNS |
ConfigMaps
kubectl create configmap app-config --from-literal=APP_COLOR=blue | From key=value pairs |
kubectl create configmap nginx-conf --from-file=default.conf | From a file (key = file name) |
kubectl get configmap app-config -o yaml | See the contents |
Secrets
kubectl create secret generic db-cred --from-literal=username=app --from-literal=password='S3cr3t!' | Generic secret |
kubectl create secret tls web-tls --cert=tls.crt --key=tls.key | TLS certificate + key |
kubectl get secret db-cred -o jsonpath='{.data.password}' | base64 -d | Decode one value (anyone with read access can) |
Apply changes
kubectl rollout restart deployment/app | Pick up changed env vars |
kubectl exec deploy/app -- env | grep APP_ | Check what a pod actually sees |
Inspect storage
kubectl get storageclass | Available storage types (default marked) |
kubectl get pvc,pv | Claims and the volumes bound to them |
kubectl describe pvc <name> | Why a claim is Pending (events) |
Access modes
ReadWriteOnce (RWO) | Read-write by pods on ONE node |
ReadOnlyMany (ROX) | Read-only by many nodes |
ReadWriteMany (RWX) | Read-write by many nodes (NFS, CephFS, EFS) |
ReadWriteOncePod (RWOP) | Read-write by exactly one pod |
Volume types
emptyDir | Scratch space; lives and dies with the pod |
hostPath | A folder on the node; avoid for apps |
persistentVolumeClaim | Durable storage that outlives pods |
configMap / secret | Config files (lesson 05) |
Namespaces
kubectl create namespace shop | Create a namespace |
kubectl config set-context --current --namespace=shop | Make it your default for this context |
kubectl get all -n shop | Deployments, pods, Services in one list |
kubectl delete namespace shop | Delete it and everything inside |
Everyday debugging
kubectl get pods -n shop -w | Watch pods change state |
kubectl -n shop describe pod <name> | Events: scheduling, pulls, probes, mounts |
kubectl -n shop logs deploy/web | Logs from one pod of a Deployment |
kubectl -n shop get endpointslices | Which pods each Service really points to |
kubectl -n shop exec deploy/postgres -- psql -U shop -d shop -c 'select 1' | Run a command in a pod |
Node preparation
sudo swapoff -a | Disable swap now (also remove it from /etc/fstab) |
sudo modprobe overlay && sudo modprobe br_netfilter | Kernel modules containers and bridged traffic need |
containerd config default | sudo tee /etc/containerd/config.toml | Write a default containerd config (then set SystemdCgroup = true) |
sudo apt-mark hold kubelet kubeadm kubectl | Stop unattended upgrades changing versions |
Build the cluster
sudo kubeadm init --control-plane-endpoint <LB-or-DNS>:6443 --pod-network-cidr 10.244.0.0/16 --upload-certs | First control-plane node |
kubeadm token create --print-join-command | New worker join command (tokens expire after 24h) |
sudo kubeadm join <endpoint>:6443 --token … --discovery-token-ca-cert-hash sha256:… | Join a worker |
sudo kubeadm reset | Undo init/join on a node (then clean CNI config and iptables) |
Certificates
sudo kubeadm certs check-expiration | When each certificate expires |
sudo kubeadm certs renew all | Renew all kubeadm-managed certificates |
ls /etc/kubernetes/pki | Where the cluster PKI lives |
Image & boot
Base image: OS + containerd + kubelet/kubeadm/kubectl (pinned) + kernel settings | Everything that's the same on every node |
Role at boot: cloud-init user-data (init / join / join --control-plane) | One image, many roles |
DHCP reservation (MAC → IP) or static network-config | Predictable node IPs |
cloud-init clean && truncate -s0 /etc/machine-id | Make a template clonable (no duplicate identities) |
kubeadm automation
kubeadm token generate | Create a bootstrap token value in advance |
kubeadm init --config init.yaml --upload-certs | First control plane from a config file |
kubeadm init phase upload-certs --upload-certs | Re-upload certs for control-plane joins (the key expires after 2 h) |
kubeadm join --config join.yaml | Join as worker or control plane from a config file |
openssl x509 -pubkey -in /etc/kubernetes/pki/ca.crt | openssl rsa -pubin -outform der 2>/dev/null | openssl dgst -sha256 -hex | Compute the CA cert hash for discovery |
Install & configure
helm repo add metallb https://metallb.github.io/metallb && helm install metallb metallb/metallb -n metallb-system --create-namespace | Install MetalLB |
IPAddressPool (spec.addresses: [10.0.0.240-10.0.0.250]) | Which IPs MetalLB may hand out |
L2Advertisement (spec.ipAddressPools: [...]) | Announce those IPs with ARP/NDP |
BGPPeer + BGPAdvertisement | Announce them to routers with BGP |
Check
kubectl get svc -A | grep LoadBalancer | EXTERNAL-IP assigned? |
kubectl get ipaddresspools,l2advertisements,bgppeers -n metallb-system | MetalLB config |
kubectl logs -n metallb-system -l app.kubernetes.io/component=speaker | Which node announces what |
arping -I eth0 10.0.0.240 (from another host on the segment) | Who answers ARP for the VIP |
Before you start
kubectl version | Current client and server versions |
kubectl get nodes | Every node's kubelet version |
sudo kubeadm upgrade plan | What you can upgrade to, component by component |
kubectl get pdb -A | PodDisruptionBudgets that could block drains |
Control plane (first node)
sudo apt-mark unhold kubeadm && sudo apt-get install -y kubeadm='1.32.x-*' && sudo apt-mark hold kubeadm | Upgrade the kubeadm binary first |
sudo kubeadm upgrade apply v1.32.x | Upgrade control-plane components (first control-plane node only) |
sudo kubeadm upgrade node | Other control-plane nodes, and every worker |
Each node's kubelet
kubectl drain <node> --ignore-daemonsets --delete-emptydir-data | Move workloads off safely |
sudo apt-get install -y kubelet='1.32.x-*' kubectl='1.32.x-*' | Upgrade kubelet and kubectl (unhold/hold around it) |
sudo systemctl daemon-reload && sudo systemctl restart kubelet | Restart with the new version |
kubectl uncordon <node> | Allow scheduling again |
Take and check a snapshot
ETCDCTL_API=3 etcdctl --endpoints=https://127.0.0.1:2379 --cacert=/etc/kubernetes/pki/etcd/ca.crt --cert=/etc/kubernetes/pki/etcd/server.crt --key=/etc/kubernetes/pki/etcd/server.key snapshot save /backup/etcd.db | Save a snapshot (kubeadm certificate paths) |
etcdutl snapshot status /backup/etcd.db -w table | Hash, revision, key count, size |
etcdctl … endpoint status -w table | Leader, DB size, raft term per member |
etcdctl … member list -w table | Members of the etcd cluster |
Restore (single control plane, kubeadm)
sudo mv /etc/kubernetes/manifests/kube-apiserver.yaml /root/ | Stop the API server (no writes during restore) |
sudo etcdutl snapshot restore /backup/etcd.db --data-dir /var/lib/etcd-restored | Restore into a NEW data directory |
sudo vi /etc/kubernetes/manifests/etcd.yaml | Point the etcd-data hostPath at /var/lib/etcd-restored |
sudo mv /root/kube-apiserver.yaml /etc/kubernetes/manifests/ | Start the API server again |
Labels & taints
kubectl label node w1 disktype=ssd | Add a node label |
kubectl get nodes -L disktype,topology.kubernetes.io/zone | Show labels as columns |
kubectl taint node w1 dedicated=gpu:NoSchedule | Repel pods without a matching toleration |
kubectl taint node w1 dedicated=gpu:NoSchedule- | Remove that taint (trailing minus) |
kubectl describe node w1 | grep -A3 Taints | See a node's taints |
Why is it Pending?
kubectl describe pod <pod> | The FailedScheduling event lists each reason |
kubectl get events --field-selector reason=FailedScheduling | All scheduling failures |
kubectl describe node <node> | grep -A8 'Allocated resources' | How much CPU/memory is already requested |
Priority
kubectl get priorityclass | Available priorities (system-cluster-critical, …) |
Check permissions
kubectl auth can-i create deployments -n shop | Can I do this? |
kubectl auth can-i --list -n shop | Everything I can do in a namespace |
kubectl auth can-i get secrets -n shop --as=system:serviceaccount:shop:app | Test as a service account |
kubectl auth whoami | Who does the API server think I am? |
Create RBAC quickly
kubectl create role pod-reader --verb=get,list,watch --resource=pods -n shop | A namespaced Role |
kubectl create rolebinding read-pods --role=pod-reader --user=asha -n shop | Bind it to a user |
kubectl create clusterrolebinding ops-view --clusterrole=view --group=ops | Built-in read-only role, cluster-wide, for a group |
kubectl create serviceaccount app -n shop | An identity for a workload |
kubectl create token app -n shop --duration=1h | A short-lived token for that service account |
Inspect
kubectl get roles,rolebindings -n shop | Namespaced RBAC objects |
kubectl describe clusterrole edit | What a built-in role allows |
Scope it
kubectl get nodes | Is it one node or all of them? |
kubectl get pods -A -o wide | grep -v Running | Everything not Running, and where |
kubectl get events -A --sort-by=.lastTimestamp | tail -30 | What happened most recently |
Pods
kubectl describe pod <pod> | Events, state, last state, exit code |
kubectl logs <pod> --previous | Logs from the crashed container |
kubectl debug -it <pod> --image=busybox:1.36 --target=<container> | Ephemeral debug container sharing the pod's namespaces |
kubectl get pod <pod> -o jsonpath='{.status.containerStatuses[*].lastState}' | Why the last run ended |
Nodes
kubectl describe node <node> | Conditions: Ready, MemoryPressure, DiskPressure, PIDPressure |
kubectl debug node/<node> -it --image=busybox:1.36 | Shell on the node (host filesystem at /host) |
systemctl status kubelet containerd | Are the node agents running? (on the node) |
journalctl -u kubelet --since '15 min ago' | Why the kubelet is unhappy (on the node) |
crictl ps -a | Containers the runtime knows about, even if the API is down |
Build it
kubectl create namespace shop-web | A home for the app |
kubectl apply -f deploy.yaml -f svc.yaml -f ingress.yaml -n shop-web | Deployment → Service → Ingress |
kubectl rollout status deploy/podinfo -n shop-web | Wait for a rollout |
kubectl set image deploy/podinfo podinfo=ghcr.io/stefanprodan/podinfo:6.7.1 -n shop-web | Roll out a new version |
kubectl rollout undo deploy/podinfo -n shop-web | Roll back |
Debug hop by hop
kubectl get pods -n shop-web -o wide | 1. Pods Running and READY? |
kubectl get endpointslices -n shop-web | 2. Service has endpoints? |
kubectl run t --rm -it --image=busybox -- wget -qO- http://podinfo.shop-web | 3. Service works inside the cluster? |
kubectl describe ingress podinfo -n shop-web | 4. Ingress rule, class, address? |
curl -v http://podinfo.127.0.0.1.nip.io/ | 5. From outside? |
Namespaces & DNS
<service>.<namespace>.svc.cluster.local | Full DNS name of a Service |
postgres.shop-data (from another namespace) | Short form that works across namespaces |
kubectl get all -n shop-data | What's in a namespace |
Secrets and ConfigMaps are namespaced | Create them where the pods run |
Database pieces
StatefulSet + volumeClaimTemplates | Stable name (postgres-0) and its own PVC |
Service with clusterIP: None | Headless: DNS for each StatefulSet pod |
initContainer: pg_isready -h postgres.shop-data | Wait for the DB before starting the app |
NetworkPolicy (namespaceSelector + podSelector) | Only the app may reach port 5432 |
Certificates
openssl req -x509 -nodes -newkey rsa:2048 -days 30 -keyout tls.key -out tls.crt -subj '/CN=shop.127.0.0.1.nip.io' -addext 'subjectAltName=DNS:shop.127.0.0.1.nip.io' | Self-signed cert for a lab |
kubectl create secret tls shop-tls --cert=tls.crt --key=tls.key -n shop-web | Store it as a TLS Secret |
kubectl apply -f https://github.com/cert-manager/cert-manager/releases/download/v1.16.2/cert-manager.yaml | Install cert-manager (pick a current release) |
kubectl get certificate,certificaterequest -A | cert-manager status |
ingress-nginx annotations
cert-manager.io/cluster-issuer: lab-ca | Ask cert-manager for this Ingress's certificate |
nginx.ingress.kubernetes.io/ssl-redirect: "true" | HTTP → HTTPS (default when TLS is set for the host) |
nginx.ingress.kubernetes.io/ssl-passthrough: "true" | Passthrough (controller needs --enable-ssl-passthrough) |
nginx.ingress.kubernetes.io/backend-protocol: "HTTPS" | Re-encrypt to the pod |
Ingress gateway (Gateway API)
kubectl apply -f https://github.com/kubernetes-sigs/gateway-api/releases/download/v1.2.1/standard-install.yaml | Gateway API CRDs (pick a current release) |
GatewayClass → Gateway (listeners) → HTTPRoute (rules) | Platform owns Gateways, teams own Routes |
kubectl get gatewayclass,gateway,httproute -A | Status of each layer |
Egress control
NetworkPolicy podSelector: {} + policyTypes: [Egress] | Default-deny egress for a namespace |
Allow kube-dns on UDP/TCP 53 | Don't forget DNS |
CiliumNetworkPolicy toFQDNs | Allow by DNS name (Cilium) |
CiliumEgressGatewayPolicy / Istio egress gateway | Fixed exit point and source IP |
Rolling & blue-green
strategy.rollingUpdate: { maxSurge: 1, maxUnavailable: 0 } | Rolling: capacity never drops |
strategy.type: Recreate | Stop all old pods, then start new (downtime) |
kubectl patch svc shop -p '{"spec":{"selector":{"app":"shop","version":"green"}}}' | Blue-green: switch the Service |
Canary
nginx.ingress.kubernetes.io/canary: "true" + canary-weight: "10" | ingress-nginx canary Ingress |
HTTPRoute backendRefs weights 90 / 10 | Gateway API canary |
Argo Rollouts: steps [ setWeight: 20, pause, analysis ] | Automated progressive delivery |
kubectl argo rollouts promote|abort <name> | Move forward or roll back |
Join more control-plane nodes
sudo kubeadm init phase upload-certs --upload-certs | Re-upload control-plane certs; prints a new certificate key (valid 2 hours) |
kubeadm token create --print-join-command | Base join command (append --control-plane --certificate-key <key>) |
sudo kubeadm join <VIP>:6443 --token … --discovery-token-ca-cert-hash … --control-plane --certificate-key … | Join as a control-plane node |
Check HA health
kubectl get nodes -l node-role.kubernetes.io/control-plane | All control-plane nodes |
kubectl get lease -n kube-system kube-scheduler kube-controller-manager | Which node currently holds each leader lease |
etcdctl … endpoint status --cluster -w table | etcd members, leader and DB size |
curl -k https://<VIP>:6443/readyz | Is the API reachable through the load balancer? |
See usage vs requests
kubectl top nodes | Real CPU/memory usage per node (needs metrics-server) |
kubectl top pods -A --sort-by=memory | Hungriest pods |
kubectl describe node <node> | grep -A10 'Allocated resources' | Sum of requests and limits on a node |
kubectl get pod <pod> -o jsonpath='{.status.qosClass}' | A pod's QoS class |
Namespace guard-rails
kubectl get limitrange,resourcequota -n <ns> | Defaults and totals for a namespace |
kubectl describe resourcequota -n <ns> | Used vs hard limits |
Units
cpu: 500m | Half a CPU core (1000m = 1 core) |
memory: 256Mi | 256 mebibytes (Mi/Gi are powers of 2; M/G are powers of 10) |
HPA
kubectl autoscale deployment web --cpu-percent=60 --min=2 --max=10 | Quick CPU-based HPA |
kubectl get hpa -w | Watch current vs target and replica count |
kubectl describe hpa web | Conditions and scaling events (why it did or didn't scale) |
Which autoscaler?
HPA | More or fewer pods, from CPU/memory/custom metrics |
VPA | Bigger or smaller requests per pod, from observed usage |
Cluster Autoscaler | More or fewer nodes in existing node groups when pods are Pending |
Karpenter | Launches right-sized nodes directly for Pending pods; consolidates |
KEDA | Scales on events (queue length, Kafka lag, cron), including to zero |
Fleet inventory
kubectl config get-contexts -o name | Every cluster in your kubeconfig |
kubectl --context <ctx> version -o json | jq -r .serverVersion.gitVersion | One cluster's server version |
kubectl --context <ctx> get nodes -o custom-columns=NAME:.metadata.name,KUBELET:.status.nodeInfo.kubeletVersion | Kubelet versions per node |
Pre-flight checks (per cluster)
kubectl get nodes | grep -v ' Ready' | Anything not Ready? |
kubectl get pods -A | grep -vE 'Running|Completed' | Unhealthy pods before you start |
kubectl get pdb -A | Budgets that could block drains |
kubectl get --raw /metrics | grep apiserver_requested_deprecated_apis | Deprecated API usage |
sudo kubeadm certs check-expiration | Certificate expiry (kubeadm clusters) |
Health and latency metrics
etcd_disk_wal_fsync_duration_seconds | etcd write-ahead-log fsync latency (keep p99 low, ~10ms) |
etcd_disk_backend_commit_duration_seconds | etcd commit latency |
apiserver_request_duration_seconds | API latency by verb and resource |
apiserver_flowcontrol_rejected_requests_total | Requests rejected by API Priority and Fairness |
Inspect
kubectl get flowschemas,prioritylevelconfigurations | How the API server prioritises clients |
kubectl get --raw /metrics | grep apiserver_request_total | Request counts by verb/resource/code |
sudo cat /var/lib/kubelet/config.yaml | The kubelet's configuration (kubeadm nodes) |
Velero
velero backup create shop-$(date +%F) --include-namespaces shop | Back up one namespace (objects + volumes as configured) |
velero backup describe <backup> --details | What was captured, warnings and errors |
velero restore create --from-backup <backup> | Restore it |
velero schedule create daily-shop --schedule='0 2 * * *' --include-namespaces shop --ttl 720h | Nightly backup kept for 30 days |
velero backup-location get | Is the object storage target available? |
Words to agree on
RPO | Recovery Point Objective: how much data you can afford to lose (time) |
RTO | Recovery Time Objective: how long until service is back |
Cluster API (clusterctl)
clusterctl init --infrastructure <provider> | Turn a cluster into a management cluster |
clusterctl generate cluster <name> --kubernetes-version <v> > cluster.yaml | Generate a workload cluster definition |
kubectl apply -f cluster.yaml | Create it (declaratively) |
clusterctl describe cluster <name> | Tree view of the cluster's machines and conditions |
clusterctl get kubeconfig <name> > <name>.kubeconfig | Get access to the new cluster |
kubectl get clusters,machinedeployments,machines | Fleet objects in the management cluster |
Fleet GitOps
argocd cluster add <context> | Register a cluster with Argo CD |
kubectl get applicationsets -n argocd | Templates that fan out apps to clusters |
Tenant kit (per namespace)
kubectl create namespace team-a | The tenant boundary |
kubectl label ns team-a pod-security.kubernetes.io/enforce=restricted | Enforce the 'restricted' Pod Security Standard |
kubectl create rolebinding team-a-edit --clusterrole=edit --group=team-a -n team-a | The team may manage its own namespace |
kubectl apply -f quota.yaml -f limitrange.yaml -f netpol-default-deny.yaml | Budget, defaults, and network isolation |
Verify isolation
kubectl auth can-i list secrets -n team-b --as-group=team-a --as=someone | Can team A read team B's secrets? (should be no) |
kubectl get networkpolicy -A | Which namespaces are isolated |
Platform building blocks
Golden path | An opinionated, supported way to build and ship a service end to end |
Platform API | A simple, declarative interface (CRD, Terraform module, template) hiding infrastructure detail |
Developer portal | A catalog of services, owners, docs and self-service actions (e.g. Backstage) |
Guard-rails | Policies that prevent unsafe choices automatically, instead of manual approvals |
Measure it
Deployment frequency / lead time | DORA: how fast changes reach production |
Change failure rate / time to restore | DORA: how safely |
Time to first deploy (new service) | How long a golden path takes, end to end |
Developer satisfaction | Survey the platform's users regularly |
Definitions
SLI | A measured ratio of good events: good ÷ total |
SLO | The target for that SLI over a window, e.g. 99.9% over 30 days |
Error budget | What the SLO allows to fail: 0.1% of 30 days ≈ 43 minutes |
Burn rate | How fast the budget is being used (1 = exactly on budget) |
Example PromQL
sum(rate(apiserver_request_total{code=~"5.."}[5m])) / sum(rate(apiserver_request_total[5m])) | API server error ratio |
histogram_quantile(0.99, sum(rate(apiserver_request_duration_seconds_bucket{verb!~"WATCH|CONNECT"}[5m])) by (le)) | p99 API latency (excluding long-lived watches) |
Design document outline
1. Context & requirements | Who, what, scale, constraints, SLOs, RPO/RTO |
2. Cluster topology | How many clusters, where, why (with an ADR) |
3. Tenancy & access | Model, tenant kit, identity, RBAC |
4. Lifecycle | Provisioning, upgrades, waves, add-ons |
5. Resilience & DR | Failure domains, backups, drills, runbooks |
6. Observability & SLOs | SLIs, SLOs, probes, alerting |
7. Cost & risks | Estimate, top risks, open questions |
Questions reviewers always ask
What happens when X fails? | For every component and dependency |
How do you upgrade it? | Without downtime, repeatedly, for years |
How do you know it's working? | SLIs, probes, alerts |
What does it cost, and what could we drop? | Trade-offs made explicit |