Cluster Design — Architect Track›Requirements & Proof · Cheat sheet & self-check

Requirements & Proof · wrap-up

Cheat sheet & self-check

12 questions across 4 lessons. Each answer links back to the lesson it came from.

Pick an answer to see if you got it, and why.

  1. Q1. The customer says 'the tracking API must be fast'. What's a good NFR?

    Show answer

    B. A good NFR states the metric, target, percentile, measurement point, window and load.

    From lesson 01 · Requirements → NFRs
  2. Q2. Two components in series each have 99.95% availability. What's the best case for the combined system?

    Show answer

    B. 0.9995 × 0.9995 ≈ 0.9990. Every serial dependency lowers the ceiling.

    From lesson 01 · Requirements → NFRs
  3. Q3. What's the difference between RPO and RTO?

    Show answer

    B. RPO drives replication/backup frequency; RTO drives automation and failover design.

    From lesson 01 · Requirements → NFRs
  4. Q4. Why write PoC success criteria before building anything?

    Show answer

    B. A PoC is an experiment: define the hypothesis and measurements first.

    From lesson 02 · PoC from scratch
  5. Q5. Why measure fsync latency on the disks planned for etcd?

    Show answer

    B. Disk latency is the most common hidden cause of unstable control planes.

    From lesson 02 · PoC from scratch
  6. Q6. What's the purpose of a weighted scoring matrix?

    Show answer

    B. Weights are agreed with stakeholders; scores come from PoC evidence.

    From lesson 02 · PoC from scratch
  7. Q7. What does a 72-hour soak test reveal that a 10-minute load test doesn't?

    Show answer

    B. Many production incidents come from things that accumulate over hours or days.

    From lesson 03 · Proving stability
  8. Q8. Why run an upgrade while the load test is running?

    Show answer

    B. An upgrade that only works on an idle cluster hasn't been proven.

    From lesson 03 · Proving stability
  9. Q9. What belongs on the executive one-pager?

    Show answer

    B. Executives decide on trade-offs; give them the decision, the evidence summary and the risks.

    From lesson 03 · Proving stability
  10. Q10. Why size from requests rather than observed average usage?

    Show answer

    B. Measure usage to set sensible requests; then size capacity from requests (plus headroom).

    From lesson 04 · Sizing & capacity math
  11. Q11. You need 8 nodes' worth of capacity and must survive the loss of 1 of 3 racks. How many nodes?

    Show answer

    B. After losing a rack, the remaining two racks must hold all 8 nodes' worth of work.

    From lesson 04 · Sizing & capacity math
  12. Q12. Why not plan for 100% utilisation of allocatable?

    Show answer

    B. Planning at 100% means the first deploy or node drain fails to schedule.

    From lesson 04 · Sizing & capacity math