Incident Handling — On-Call Playbook & Real Scenarios
How to handle an incident well as the engineer on call: acknowledge and triage, know where to start, bring in the right people, use and grow a knowledge base, write the RCA and stop it happening again. Then practise on realistic Kubernetes, edge and AWS/EKS incidents, each with symptoms, commands, fix and prevention.
What you'll be able to do
- Take a page calmly: acknowledge, assess impact, stabilise, then diagnose
- Start troubleshooting in the right place and narrow down by evidence
- Escalate and communicate clearly, and leave a knowledge trail for next time
- Recover from classic Kubernetes, edge and EKS incidents with confidence
Before you start
Kubernetes Administration Level 2, Linux Level 2.
How it works
Each lesson: plain-language idea → how it really works → hands-on. Each section ends with a cheat sheet & self-check.
Curriculum
Lessons marked “Read” are ready; the rest are on the way.
The on-call playbook: from page to prevention
- 01Answering the pageAcknowledge, assess impact, stabilise firstRead →
- 02Where to start troubleshootingScope, what changed, outside-in, hypothesesRead →
- 03Engaging people & communicatingWhen and how to escalate, updates, vendors, handoversRead →
- 04Runbooks & the knowledge baseA wiki of known issues and actions that actually gets usedRead →
- 05Writing the RCAA practical template, with a worked exampleRead →
- 06Preventing the repeatFrom action items to fixes that remove the class of failureRead →
- 📋Cheat sheet & self-checkEvery command from this section on one page, then 18 questions to check yourself.Open →
Real-world incident scenarios
- 07etcd out of space at 2 a.m.Quota alarm vs full disk, compact, defrag, disarmRead →
- 08The noisy neighbourOne workload starves a node; limits, QoS and isolationRead →
- 09Certificates expired after a failed rotationkubeadm certs, kubelet client certs, cert-managerRead →
- 10Docker network collides with the site networkEdge sites where 172.17.0.0/16 was already takenRead →
- 11Node unreachable after a network restartLost routes, a downed loopback, and remote handsRead →
- 12Longhorn volumes won't attach after a rebootiscsid, multipathd, stale attachments, degraded replicasRead →
- 13Clock drift: NTP sync failedx509 'not yet valid', etcd clock warnings, chronyRead →
- 14Cluster-wide DNS timeoutsCoreDNS load, ndots, conntrack and node-local cachingRead →
- 15A broken admission webhook blocks every deployfailurePolicy, scope and escape hatchesRead →
- 16Node disk full: images and logsDiskPressure evictions, image GC, log rotationRead →
- 17EKS: pods can't get IP addressesVPC CNI IP exhaustion, prefix delegation, subnet planningRead →
- 18EKS: locked out after editing aws-authBroken identity mapping, recovery, access entriesRead →
- 19AWS: ALB 502s during every deploymentTarget deregistration, preStop and readiness gatesRead →
- 📋Cheat sheet & self-checkEvery command from this section on one page, then 39 questions to check yourself.Open →
Real-world scenarios
Work through each one: symptom → misleading signal → evidence → root cause → prevention.
Writes fail cluster-wide. Recover without losing state.
No SSH, no console, customer site.
One bad YAML edit in aws-auth.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.