Incident Handling — On-Call Playbook & Real Scenarios
Learning Hub / Incident Handling

Incident Handling — On-Call Playbook & Real Scenarios

Practitioner → Advanced19 lessonsAvailable

How to handle an incident well as the engineer on call: acknowledge and triage, know where to start, bring in the right people, use and grow a knowledge base, write the RCA and stop it happening again. Then practise on realistic Kubernetes, edge and AWS/EKS incidents, each with symptoms, commands, fix and prevention.

You'll meeton-calltriagetroubleshootingescalationrunbooksRCAetcdcertificatesLonghornNTPEKS
Start lesson 01 →

What you'll be able to do

  • Take a page calmly: acknowledge, assess impact, stabilise, then diagnose
  • Start troubleshooting in the right place and narrow down by evidence
  • Escalate and communicate clearly, and leave a knowledge trail for next time
  • Recover from classic Kubernetes, edge and EKS incidents with confidence

Before you start

Kubernetes Administration Level 2, Linux Level 2.

How it works

Each lesson: plain-language idea → how it really works → hands-on. Each section ends with a cheat sheet & self-check.

Curriculum

Lessons marked “Read” are ready; the rest are on the way.

Real-world scenarios

Work through each one: symptom → misleading signal → evidence → root cause → prevention.

Its 2 a.m. and etcd is full

Writes fail cluster-wide. Recover without losing state.

The edge node vanished after 'just restarting the network'

No SSH, no console, customer site.

Everyone lost access to the EKS cluster

One bad YAML edit in aws-auth.

This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.