Incident Handling — On-Call Playbook & Real Scenarios›03 · Engaging people & communicating

Lesson 03 of 19 · The on-call playbook: from page to prevention

Engaging people & communicating

Incidents are team sports: when to escalate (earlier than you think), how to ask for help so people can act immediately, keeping stakeholders informed with predictable updates, working with cloud and vendor support, and handing over cleanly.

Practitioner
Key wordsescalationengaging expertsincident commanderstatus updatesstakeholdersvendor supporthandoverclear asks

Nobody handles a big incident alone

The most common incident anti-pattern isn't a technical mistake; it's one engineer quietly struggling for an hour before telling anyone. Escalating is part of doing the job well.

If you're lost on a school trip, the smart move isn't wandering around alone until dark. You ask a grown-up early, say clearly where you've already looked, and stay where they can find you. Asking early is being sensible, not being weak.

When to escalate

  • No mitigation within 15–30 minutes (set a timer).
  • Impact is severe or growing.
  • It needs access or knowledge you don't have (database, network, a service you don't own, the cloud provider).
  • You're tired, unsure or overwhelmed.

Who to engage

Need Who
Coordination on a big incident Incident commander (see SRE & Production Incident Response, lesson 02)
The affected service's internals Its owning team (from your service catalogue)
Specialist layers DBA, network, security, storage teams
Provider or product issues Cloud/vendor support
Business and customer communication Communications lead / incident manager

Keep an up-to-date ownership catalogue (service → team → on-call rota); during an incident is the worst time to find out who owns something.

How to ask

@payments-oncall SEV2: checkout 5xx ~6% since 02:01 (EU only).
Done: rolled back payments v2.3 at 02:15, no change. DB CPU normal, no pod restarts.
Ruled out: ingress, DNS.
Ask: can you check the payment-provider connector? Bridge: <link>. Timeline in #inc-0927.

Keeping people informed

  • Fixed rhythm: every 30 minutes for SEV1, 60 for SEV2, even when there's no news ("still investigating, next update 03:30").
  • Same structure: impact, what's known, what's being done, next update.
  • One channel for responders; stakeholders get summaries, so engineers aren't interrupted.
  • Be factual; avoid guessing causes in customer-facing messages.

Working with cloud and vendor support

  • Open the case early with the right severity.
  • Include: account/region, resource IDs (cluster ARN, instance IDs, load balancer names), exact timestamps (UTC), error messages and request IDs, what you've tried.
  • Keep mitigating in parallel; don't wait for their answer.

Handover

When a shift ends mid-incident, hand over explicitly: current impact, hypothesis, what's ruled out, mitigations in place (and how to undo them), next steps, who's doing what, next update time. The new owner confirms "I have it".

Try it: practise the human side

  1. Write your team's escalation rules on one page: timers, who to call for what, how to page them.
  2. Check your ownership catalogue for your top 10 services; fix any missing owners.
  3. Draft a help request and a status update for a fictional incident using the templates above.
  4. Prepare a "vendor case" template with the fields support always asks for.
  5. Run a 20-minute tabletop with a mid-incident handover.

Going deeper: people and pressure

  • Thank people who were paged; being woken to help should feel valued.
  • Avoid "hero culture"; reliable systems don't depend on one person being awake.
  • After big incidents, check on responders' workload and rest.

Recap

  • Escalate early: timers, growing impact, missing access, tiredness.
  • Engage the right owners from an up-to-date catalogue.
  • Ask with context, what's ruled out and a specific ask.
  • Predictable updates for stakeholders; early, detailed vendor cases; explicit handovers.

This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.