Incident Handling cheat sheet
151 commands from every lesson of Incident Handling — On-Call Playbook & Real Scenarios, on one page.
First 5 minutes
Acknowledge the page | Means 'I own this now', not 'it's fixed' |
Read the alert + its runbook link | What fired, since when, which service |
Check user impact (SLO dashboard, edge errors, support) | Who is hurt, how many, how badly |
What changed? (deploys, config, infra, certificates, time) | Most incidents follow a change |
Decide severity; declare if SEV1/SEV2 | Get help early |
Post the first update | Impact · what we know · next update time |
Stabilise options
kubectl rollout undo deploy/<name> -n <ns> | Roll back a bad release |
git revert <sha> (GitOps) | Roll back through Git |
Disable the feature flag | Turn off new behaviour |
Scale up / fail over / shed load | Buy time while you diagnose |
Kubernetes first look
kubectl get events -A --sort-by=.lastTimestamp | tail -30 | What just happened, everywhere |
kubectl get pods -A | grep -vE 'Running|Completed' | Anything unhealthy |
kubectl get nodes -o wide / kubectl top nodes | Node status and load |
kubectl describe pod <p> -n <ns> | Events, restarts, last state |
kubectl logs <p> -n <ns> --previous | Logs of the crashed container |
Method
Scope: user → pod → node → zone → everything | Where does it stop being broken? |
Compare a good and a bad instance | The difference is the clue |
One hypothesis at a time, cheapest test first | Write down what's ruled out |
Escalate when
No mitigation after 15–30 minutes | Don't struggle alone |
Impact is growing or severe (SEV1/SEV2) | More hands, clear roles |
It needs access or knowledge you don't have | Service owners, DBAs, network, vendor |
You're tired or unsure | Fresh eyes are a feature |
A good ask
Context: service, impact, since when | One line |
What's done and ruled out | Saves them repeating it |
The specific ask + where to join | "Check the payments DB, bridge link…" |
Runbook sections
What the alert means + user impact | Why you were woken |
Quick checks (copy-paste commands) | Confirm and scope in minutes |
Mitigations (safest first) | Stop the bleeding |
Escalation (who, how) | When to call whom |
Links: dashboard, logs, past incidents | One click away |
Known-issue entry
Symptom (exact error text) | What people will search for |
Cause | Why it happens |
Action / fix / workaround | What to do |
Prevention status + links | Is it fixed for good? |
RCA template
Summary (3 sentences) | What happened, impact, fix |
Impact (numbers) | Users, duration, errors, SLO budget used |
Timeline (UTC) | Detection, response, mitigation, resolution |
Root cause + contributing factors | Mechanism, and why it was possible |
What went well / badly / lucky | Honest response review |
Action items (owner, date) | Specific and verifiable |
Useful tools
5 whys | Ask 'why' until you reach something you can change |
Time to detect / mitigate / resolve | Measure the response |
Four kinds of fixes (strongest first)
Eliminate | Design it away (automation, safer defaults, remove the dependency) |
Detect earlier | Alert on leading indicators (80% full, 30 days to expiry) |
Limit blast radius | Canaries, quotas, limits, staged rollouts, isolation |
Recover faster | Runbooks, tested rollback/restore, automation |
Follow-through
Track actions like features (owner, due date) | Not in a forgotten doc |
Verify: test, drill or chaos experiment | Prove the fix works |
Monthly review of repeats and overdue actions | Keep it honest |
Inspect (on a control-plane node, kubeadm paths)
export ETCDCTL_API=3 ETCD='--endpoints=https://127.0.0.1:2379 --cacert=/etc/kubernetes/pki/etcd/ca.crt --cert=/etc/kubernetes/pki/etcd/server.crt --key=/etc/kubernetes/pki/etcd/server.key' | Connection flags |
etcdctl $ETCD endpoint status --cluster -w table | DB size, leader, raft index per member |
etcdctl $ETCD alarm list | memberID:… alarm:NOSPACE ? |
df -h /var/lib/etcd | Is the disk itself full? |
Recover
etcdctl $ETCD snapshot save /var/backups/etcd-$(date +%F-%H%M).db | Snapshot first, always |
rev=$(etcdctl $ETCD endpoint status -w json | jq '.[0].Status.header.revision') | Current revision |
etcdctl $ETCD compact $rev | Drop old revisions |
etcdctl $ETCD defrag --endpoints=<one member> | Return free space, one member at a time |
etcdctl $ETCD alarm disarm | Clear NOSPACE after space is back |
Find the culprit
kubectl top pods -A --sort-by=cpu | head / --sort-by=memory | Biggest consumers |
kubectl describe node <n> | sed -n '/Conditions/,/Events/p' | Pressure conditions, allocated resources |
kubectl get pods -A -o wide --field-selector spec.nodeName=<n> | Who shares the node |
crictl stats (on the node) | Per-container usage |
iostat -x 2 / pidstat -d 2 | Disk hogs |
Contain & prevent
kubectl scale deploy/<hog> --replicas=0 -n <ns> (or cordon + move) | Immediate relief |
resources.requests + limits on every container | Fair scheduling and caps |
LimitRange (defaults) + ResourceQuota (per namespace) | Guard rails for teams |
Taints/tolerations, dedicated node pools | Isolate heavy or sensitive workloads |
Diagnose
sudo kubeadm certs check-expiration | Every kubeadm-managed cert and CA |
echo | openssl s_client -connect 127.0.0.1:6443 2>/dev/null | openssl x509 -noout -dates | What the API server presents |
sudo openssl x509 -in /var/lib/kubelet/pki/kubelet-client-current.pem -noout -enddate | Kubelet client cert on a node |
kubectl get certificate -A (cert-manager) | Ingress/app certificates |
Renew (kubeadm control plane)
sudo kubeadm certs renew all | Renew leaf certs (on each control-plane node) |
Restart control-plane static pods (move manifests out and back, or restart the containers) | They must reload the new certs |
sudo cp /etc/kubernetes/admin.conf ~/.kube/config | Refresh your admin kubeconfig |
Diagnose
ip route / ip -br addr | Look for 172.17.0.0/16 dev docker0 (or br-…) overlapping site ranges |
ip route get 172.17.5.40 | Which interface would traffic to that IP use? |
docker network ls / docker network inspect bridge | Docker networks and their subnets |
kubectl cluster-info dump | grep -m1 -E 'cluster-cidr|service-cluster-ip-range' | Pod and Service CIDRs |
Fix Docker ranges (/etc/docker/daemon.json)
"bip": "192.168.254.1/24" | Move the default bridge |
"default-address-pools": [ { "base": "192.168.240.0/20", "size": 24 } ] | Ranges for user-defined networks |
sudo systemctl restart docker | Apply (recreate containers/networks as needed) |
Check (from the console / BMC)
ip -br link / ip -br addr | Interfaces up? Addresses present? |
ip route / ip route show table local | Default route? local routes incl. 127.0.0.0/8 dev lo? |
ping -c1 127.0.0.1 / ping -c1 <gateway> | Loopback and gateway reachable? |
journalctl -u systemd-networkd -u NetworkManager --since -1h | What the network service did |
systemctl status kubelet containerd | Did Kubernetes components survive? |
Recover
sudo ip link set lo up | Bring the loopback back |
sudo netplan apply / nmcli con up <name> | Re-apply the persistent config |
sudo systemctl restart kubelet | After the network is right |
Diagnose
kubectl describe pod <p> -n <ns> | sed -n '/Events/,$p' | FailedAttachVolume / FailedMount / Multi-Attach |
kubectl -n longhorn-system get volumes.longhorn.io,replicas.longhorn.io | Volume state/robustness, replicas |
kubectl get volumeattachments | grep <pv> | Stale attachment to another node? |
systemctl status iscsid (on the node) | iSCSI initiator running? |
multipath -ll / lsblk | Is multipathd holding Longhorn's devices? |
Fix
sudo systemctl enable --now iscsid | Start iSCSI and keep it enabled |
/etc/multipath.conf: blacklist { devnode "^sd[a-z0-9]+" } | Stop multipathd grabbing Longhorn devices (check Longhorn's KB for your setup) |
kubectl delete volumeattachment <name> (only when the old node is really gone) | Clear a stale attachment |
Check
timedatectl | System clock synchronized? NTP service active? |
chronyc tracking | Offset from the reference, stratum, last update |
chronyc sources -v | Which servers, reachable? (^* = selected) |
for n in node1 node2 node3; do ssh $n date -u +%s; done | Compare node clocks quickly |
Fix
Allow UDP 123 to your NTP servers / run a local NTP server | Restore sync |
sudo chronyc makestep | Step the clock now (jumps time; see cautions) |
node_timex_offset_seconds > 0.05 (node-exporter) | Alert on drift |
Diagnose
kubectl -n kube-system get pods -l k8s-app=kube-dns -o wide | CoreDNS pods healthy? Where? |
kubectl -n kube-system logs -l k8s-app=kube-dns --tail=50 | Errors, upstream timeouts |
kubectl run dns --rm -it --image=busybox:1.36 -- nslookup kubernetes.default | Test from a pod |
cat /etc/resolv.conf (inside a pod) | search domains, ndots:5 |
coredns_dns_requests_total, coredns_dns_request_duration_seconds | CoreDNS metrics |
Fix
Scale CoreDNS (replicas / cluster-proportional-autoscaler) | More capacity |
NodeLocal DNSCache | Per-node cache, fewer conntrack issues |
dnsConfig: options: [ { name: ndots, value: "2" } ] | Fewer search-path lookups for external names |
Use FQDNs with a trailing dot for external hosts (api.example.com.) | Skip search expansion |
Diagnose
kubectl get validatingwebhookconfigurations,mutatingwebhookconfigurations | Which webhooks exist |
kubectl get validatingwebhookconfiguration <name> -o yaml | grep -E 'failurePolicy|timeoutSeconds|service:' -A3 | Fail closed? Which Service? |
kubectl -n <ns> get pods,endpoints <webhook-svc> | Is the webhook backend running? |
Recover (carefully)
Restore the webhook's pods (scale up, fix the node, fix its certificate) | Best fix |
kubectl patch validatingwebhookconfiguration <name> --type=json -p '[{"op":"replace","path":"/webhooks/0/failurePolicy","value":"Ignore"}]' | Temporary fail-open (record it, revert later) |
Back up the config, then delete it only as last resort | Escape hatch |
Diagnose (on the node)
kubectl describe node <n> | grep -A6 Conditions | DiskPressure=True? |
df -h / /var/lib/containerd /var/log | Which filesystem is full |
sudo du -xh --max-depth=2 /var | sort -h | tail -15 | Biggest directories |
sudo crictl images / sudo crictl ps -a | Images and (exited) containers |
sudo journalctl --disk-usage | journald size |
Free space safely
sudo crictl rmi --prune | Remove images not used by any container |
sudo journalctl --vacuum-size=500M | Shrink the journal |
kubectl delete pods -A --field-selector=status.phase=Failed | Clean up evicted pod records |
resources.limits.ephemeral-storage | Cap a pod's local disk use |
Diagnose
kubectl describe pod <p> | grep -i 'assign an IP' | The CNI error |
aws ec2 describe-subnets --subnet-ids <ids> --query 'Subnets[].[SubnetId,AvailableIpAddressCount]' | Free IPs per subnet |
kubectl -n kube-system logs -l k8s-app=aws-node -c aws-node --tail=50 | VPC CNI (ipamd) errors |
kubectl get node <n> -o jsonpath='{.status.allocatable.pods}' | Max pods on the node |
kubectl -n kube-system get ds aws-node -o yaml | grep -A1 -E 'WARM_|PREFIX' | CNI settings |
Fix / plan
ENABLE_PREFIX_DELEGATION=true (Nitro instances) | Assign /28 prefixes per ENI slot: more pods per node |
Tune WARM_IP_TARGET / MINIMUM_IP_TARGET | Don't hoard IPs on every node |
Add subnets / a secondary CIDR (e.g. 100.64.0.0/16) with custom networking | More address space for pods |
Diagnose
aws sts get-caller-identity | Which IAM identity am I using? |
aws eks describe-cluster --name <c> --query 'cluster.accessConfig' | Authentication mode (CONFIG_MAP / API_AND_CONFIG_MAP / API) |
aws eks list-access-entries --cluster-name <c> | Access entries (if enabled) |
kubectl -n kube-system get configmap aws-auth -o yaml | The mapping (once you have access) |
Recover & harden
Use the cluster creator identity or an existing admin access entry | A path that doesn't depend on aws-auth |
aws eks create-access-entry + associate-access-policy (AmazonEKSClusterAdminPolicy) | Grant admin via the EKS API |
aws eks update-cluster-config --access-config authenticationMode=API_AND_CONFIG_MAP | Enable access entries |
Diagnose
ALB metrics: HTTPCode_ELB_502_Count, HTTPCode_ELB_504_Count, TargetResponseTime | Errors align with deploy times? |
kubectl get targetgroupbindings -A | Which Services the controller manages |
kubectl get pod <p> -o jsonpath='{.status.conditions}' | Readiness gate condition present? |
ALB access logs: elb_status_code vs target_status_code | ALB-generated vs app-generated errors |
Fix
lifecycle.preStop: sleep ~15–30s | Keep serving while the ALB deregisters the target |
terminationGracePeriodSeconds > preStop + drain time | Don't get killed mid-drain |
namespace label elbv2.k8s.aws/pod-readiness-gate-inject=enabled | Pods Ready only when healthy in the target group |
alb.ingress.kubernetes.io/target-group-attributes: deregistration_delay.timeout_seconds=30 | Align deregistration delay |