The on-call playbook: from page to prevention · wrap-up
Cheat sheet & self-check
18 questions across 6 lessons. Each answer links back to the lesson it came from.
Pick an answer to see if you got it, and why.
Q1. What does acknowledging a page communicate?
Show answer
B. Acknowledge fast, even before you know anything; then assess.
From lesson 01 · Answering the pageQ2. The error rate jumped right after a deploy 10 minutes ago. What's usually the best first action?
Show answer
B. Stabilise first. Root cause can wait; users can't.
From lesson 01 · Answering the pageQ3. Why start a timeline immediately?
Show answer
B. A few lines per action in the incident channel is enough.
From lesson 01 · Answering the pageQ4. Only pods on one node are failing. What does that suggest?
Show answer
B. Scoping by node, zone or version points you at the right layer immediately.
From lesson 02 · Where to start troubleshootingQ5. Why walk the request path outside-in (edge → ingress → service → pods → nodes)?
Show answer
B. The first hop that fails tells you where to dig.
From lesson 02 · Where to start troubleshootingQ6. You have three theories. Which do you test first?
Show answer
B. Quick eliminations shrink the search space fast. Change one thing at a time.
From lesson 02 · Where to start troubleshootingQ7. You've been investigating alone for 25 minutes, impact is growing, no mitigation yet. What now?
Show answer
B. Escalating early is good practice, not a failure. Most organisations wish people escalated sooner.
From lesson 03 · Engaging people & communicatingQ8. What makes a request for help effective?
Show answer
B. People can act in seconds instead of asking ten questions.
From lesson 03 · Engaging people & communicatingQ9. Why open a cloud/vendor support case early during a suspected provider issue?
Show answer
B. You can always close the case if it turns out to be your side.
From lesson 03 · Engaging people & communicatingQ10. Where should a runbook be linked from?
Show answer
B. The runbook is only useful if it's one click from the page.
From lesson 04 · Runbooks & the knowledge baseQ11. What makes a known-issue entry findable during an incident?
Show answer
B. People search for what's on their screen. Put the literal error text in the entry.
From lesson 04 · Runbooks & the knowledge baseQ12. When should a runbook step be automated?
Show answer
B. Automate the boring, proven steps first; keep judgement where it matters.
From lesson 04 · Runbooks & the knowledge baseQ13. What's wrong with the root cause 'engineer forgot to renew the certificate'?
Show answer
B. Blameless RCAs find system fixes: automation, monitoring, checks.
From lesson 05 · Writing the RCAQ14. Which is a good action item?
Show answer
B. Specific, owned, dated and verifiable.
From lesson 05 · Writing the RCAQ15. Why record 'what went well' too?
Show answer
B. Learning includes recognising what to keep doing.
From lesson 05 · Writing the RCAQ16. etcd filled up because of event churn. Which fix is 'eliminate' rather than 'detect'?
Show answer
B. Detection is valuable too, but removing the cause beats catching it every time.
From lesson 06 · Preventing the repeatQ17. What's a leading indicator?
Show answer
B. Leading indicators turn incidents into planned work.
From lesson 06 · Preventing the repeatQ18. Why verify a fix with a test, drill or chaos experiment?
Show answer
B. An untested fix is a hope.
From lesson 06 · Preventing the repeat