Operations · wrap-up
Cheat sheet & self-check
12 questions across 4 lessons. Each answer links back to the lesson it came from.
Pick an answer to see if you got it, and why.
Q1. Why should the monitoring backends not live only inside the production cluster they monitor?
Show answer
B. Observability must survive the failures it's meant to explain.
From lesson 10 · Observability & SLOsQ2. All internal health checks pass but users fail. Which SLO was likely missing from the design?
Show answer
B. Internal checks can't see DNS, CDN/WAF, TLS or routing failures before traffic reaches the service.
From lesson 10 · Observability & SLOsQ3. What must every alert in the catalogue have?
Show answer
B. Alerts without owners and runbooks get ignored or cause confused incident response.
From lesson 10 · Observability & SLOsQ4. In an FMEA, what does the RPN (risk priority number) represent?
Show answer
B. RPN is a prioritisation aid, not an absolute measure; also look at any single high severity score.
From lesson 11 · Resilience & failure modesQ5. What's a steady-state hypothesis in a chaos experiment?
Show answer
B. The experiment tests whether the system keeps its steady state under the fault.
From lesson 11 · Resilience & failure modesQ6. Why does a high 'Detection' score (hard to detect) raise priority?
Show answer
B. Better alerts and synthetic checks reduce D, and therefore RPN.
From lesson 11 · Resilience & failure modesQ7. Why must the old primary database be fenced before promoting the secondary?
Show answer
B. Fencing can be power-off, network isolation, or revoking access; it must be certain before promotion.
From lesson 12 · Primary/secondary & failoverQ8. What keeps the secondary site's clusters in the same configuration as the primary?
Show answer
B. Drift between sites is a classic reason failovers fail.
From lesson 12 · Primary/secondary & failoverQ9. Why is an untested DR plan risky?
Show answer
B. The first real failover shouldn't be the first failover.
From lesson 12 · Primary/secondary & failoverQ10. What's the purpose of an NFR traceability table in a design document?
Show answer
B. Reviewers look for requirements with no design answer or no evidence.
From lesson 13 · Capstone: design reviewQ11. A reviewer asks, 'What happens if DC-B's management cluster is down during a DC-A failure?' A good answer…
Show answer
B. Credible designs acknowledge limits honestly and show they were considered.
From lesson 13 · Capstone: design reviewQ12. Why include 'open questions' in the design document?
Show answer
B. Hidden uncertainty surfaces later as incidents; visible uncertainty gets managed.
From lesson 13 · Capstone: design review