Cluster Design cheat sheet
84 commands from every lesson of Cluster Design — Architect Track, on one page.
Availability math (30-day month)
99.9% → 43.2 min/month downtime | 43,200 min × 0.001 |
99.95% → 21.6 min/month | 43,200 × 0.0005 |
99.99% → 4.32 min/month | 43,200 × 0.0001 |
Serial: A = A1 × A2 | 0.999 × 0.999 ≈ 0.998 |
Parallel (independent): A = 1 − (1 − A1)(1 − A2) | two 99% → 99.99% (only if truly independent) |
The NFR sheet
ID · category · requirement · metric · target · measurement · owner | One measurable row per requirement |
Categories: availability, latency, throughput, scalability, durability (RPO/RTO), security, operability, cost | Cover them all |
Baseline tools
sonobuoy run --mode=certified-conformance | Kubernetes conformance tests |
kube-burner init -c <config> | Control-plane / object churn load |
iperf3 -c <pod-ip> -P 4 -t 30 | Pod-to-pod network throughput |
fio --name=etcd --rw=write --ioengine=sync --fdatasync=1 --bs=2300 --size=22m --directory=/var/lib/etcd-test | fsync latency like etcd's WAL |
k6 run load.js | Application load test |
PoC discipline
Success criteria written before starting | Pass/fail per criterion |
Weighted scoring matrix | Criteria × weight × score per option |
Time-box (e.g. 3–4 weeks) | A PoC is not a pilot |
Stability tests
Soak: realistic load for 48–72 h | Leaks, slow degradation, log/disk growth |
Peak + burst profile (k6 stages) | Behaviour at NFR-03 load |
Kill a node / pod / etcd member under load | Recovery time and error budget spent |
Minor upgrade under load | No failed requests (NFR-07) |
Deliverables
Evidence pack | Method, environment, raw results, graphs, issues found and fixed |
Exec one-pager | Recommendation, confidence, risks, cost, next steps |
Risk register | Risk · likelihood · impact · mitigation · owner |
The chain
Demand = Σ (replicas × requests) at peak (+ daemonsets per node) | What workloads reserve |
Allocatable = capacity − system/kube reserved − eviction threshold | What a node can offer |
Usable per node = allocatable × target utilisation (e.g. 70%) | Room for bin-packing and bursts |
Nodes = ceil(demand ÷ usable per node), for CPU and memory; take the larger | Base count |
Survive losing 1 of R racks: nodes × R ÷ (R − 1) | Failure-domain headroom |
Limits to check
max pods per node (kubelet default 110) | Pod density |
pod CIDR per node (e.g. /24 → 256 addresses) | IPAM (lesson 06) |
Upstream tested limits: 5,000 nodes, 150,000 pods per cluster, 110 pods per node | Far above most designs, but real |
Control plane options
3 stacked control-plane nodes (etcd on the same nodes) | Default for most clusters |
5 members | Survives 2 failures; more write latency and cost |
External etcd cluster | Isolates etcd; more machines to run |
Spread members across racks/rooms | One rack loss keeps quorum |
ADR template
Title · Status · Context · Decision · Consequences · Alternatives considered | One decision per record |
docs/adr/ADR-00N-<slug>.md in the platform repo | Versioned with the code |
Sizing ranges
/16 pod CIDR with /24 per node → 256 nodes × 256 addresses | kubeadm/controller-manager: --cluster-cidr, --node-cidr-mask-size |
Service CIDR /16 → 65,536 ClusterIPs | --service-cluster-ip-range (API server) |
Calico default IPAM block /26; Cilium cluster-pool default /24 per node | CNI-specific allocation |
Ranges must not overlap: clusters, DCs, VPN users, partners | Future routing and peering |
Routing
Overlay (VXLAN/Geneve) | Works anywhere; ~50 bytes overhead; pods not routable outside |
Native routing + BGP to ToR | Pods routable; no encapsulation; needs network team |
Layers
L4: MetalLB (BGP/L2), kube-vip, or hardware/virtual LB | Stable IPs for ingress and the API server |
L7: ingress controller or Gateway API implementation | Host/path routing, TLS, retries, rate limits |
externalTrafficPolicy: Local | Keep client source IP; only nodes with local endpoints receive traffic |
PROXY protocol | Pass client IP through an L4 LB that terminates TCP |
TLS options
Terminate at external LB/WAF | Central certs; LB sees plaintext |
Terminate at ingress (cert-manager) | Kubernetes-native certs; common default |
Re-encrypt to pods / mesh mTLS | Encrypted end to end inside the DC |
Passthrough (SNI routing) | App holds the key; no L7 features at ingress |
Edge layers
DNS/GSLB → CDN → WAF + DDoS → L4 VIP → ingress | Outside-in order |
Rate limits: per client/API key at the gateway or ingress | Protect backends from floods and abuse |
API gateway: auth, quotas, versioning for partners | Partner-facing APIs |
Isolation layers (namespace model)
RBAC + NetworkPolicy (default deny) + ResourceQuota/LimitRange | The minimum set |
Pod Security Admission (restricted) + policy engine | Workload hardening |
Dedicated node pools / taints for sensitive tenants | Kernel and noisy-neighbour isolation |
Choices per data class
Operator-managed DB on local NVMe (app-level replication) | Fast; the database replicates, not the storage |
Replicated block storage (Ceph RBD, Longhorn) | Storage-level replication; easier moves |
External database service / DBA-run cluster | Existing expertise, outside the cluster lifecycle |
Replication across sites
Synchronous | RPO ≈ 0; needs low round-trip latency (metro distances); writes slower |
Asynchronous | RPO = replication lag; any distance; possible data loss on failover |
Backups (PITR) off-site | Protect against corruption and mistakes, not just hardware loss |
Architecture
Per-cluster collection (Prometheus/agents, log and trace collectors) | Close to the workloads |
Central storage/query in the mgmt cluster or another site | Survives a production cluster failure |
Meta-monitoring + external dead man's switch | Know when monitoring itself is down |
Synthetic checks from outside both DCs | The user's view |
Alert catalogue columns
Name · signal/SLO · condition · severity · owner · runbook · notification route | One row per alert |
FMEA
Component · failure mode · effect · S · O · D · RPN = S×O×D · mitigation · owner | One row per failure mode |
S, O, D scored 1–10 (10 = worst: severe, frequent, undetectable) | Scoring scale |
Re-score after mitigation | Show risk reduced |
Chaos experiments
Steady-state hypothesis → inject → observe → abort conditions | Experiment structure |
Chaos Mesh / LitmusChaos (PodChaos, NetworkChaos, IOChaos…) | Kubernetes-native fault injection |
Start in nonprod, small blast radius, business hours | Safety rules |
Patterns
Active-passive (warm standby) | Secondary runs the platform + replicas; takes traffic on failover |
Active-active | Both sites serve traffic; needs multi-site data design |
GSLB / DNS with health checks (low TTL) | Switch traffic between sites |
Failover essentials
Decide → fence old primary → promote → switch traffic → verify | Order matters |
Measure RPO (lag at failover) and RTO (time to healthy) | Compare with NFR-04 |
Failback is a second, planned failover | Resync, then switch back |
Document structure
1 Summary · 2 Requirements (NFR/SLO sheet) · 3 Architecture overview | Front |
4 Compute & sizing · 5 Network & IPAM · 6 Traffic & edge · 7 Storage & data | Core design |
8 Security & tenancy · 9 Observability · 10 Resilience (FMEA, chaos) · 11 DR | Operability |
12 Cost · 13 Risks & open questions · Appendix: ADR index, evidence pack | Close |
Traceability
NFR → design section → ADR → evidence (test/drill) → status | One row per NFR |