Cheat Sheets / Kubernetes & Platform

Cluster Design cheat sheet

84 commands from every lesson of Cluster Design — Architect Track, on one page.

01 · Requirements → NFRs

Availability math (30-day month)

99.9% → 43.2 min/month downtime43,200 min × 0.001
99.95% → 21.6 min/month43,200 × 0.0005
99.99% → 4.32 min/month43,200 × 0.0001
Serial: A = A1 × A20.999 × 0.999 ≈ 0.998
Parallel (independent): A = 1 − (1 − A1)(1 − A2)two 99% → 99.99% (only if truly independent)

The NFR sheet

ID · category · requirement · metric · target · measurement · ownerOne measurable row per requirement
Categories: availability, latency, throughput, scalability, durability (RPO/RTO), security, operability, costCover them all

02 · PoC from scratch

Baseline tools

sonobuoy run --mode=certified-conformanceKubernetes conformance tests
kube-burner init -c <config>Control-plane / object churn load
iperf3 -c <pod-ip> -P 4 -t 30Pod-to-pod network throughput
fio --name=etcd --rw=write --ioengine=sync --fdatasync=1 --bs=2300 --size=22m --directory=/var/lib/etcd-testfsync latency like etcd's WAL
k6 run load.jsApplication load test

PoC discipline

Success criteria written before startingPass/fail per criterion
Weighted scoring matrixCriteria × weight × score per option
Time-box (e.g. 3–4 weeks)A PoC is not a pilot

03 · Proving stability

Stability tests

Soak: realistic load for 48–72 hLeaks, slow degradation, log/disk growth
Peak + burst profile (k6 stages)Behaviour at NFR-03 load
Kill a node / pod / etcd member under loadRecovery time and error budget spent
Minor upgrade under loadNo failed requests (NFR-07)

Deliverables

Evidence packMethod, environment, raw results, graphs, issues found and fixed
Exec one-pagerRecommendation, confidence, risks, cost, next steps
Risk registerRisk · likelihood · impact · mitigation · owner

04 · Sizing & capacity math

The chain

Demand = Σ (replicas × requests) at peak (+ daemonsets per node)What workloads reserve
Allocatable = capacity − system/kube reserved − eviction thresholdWhat a node can offer
Usable per node = allocatable × target utilisation (e.g. 70%)Room for bin-packing and bursts
Nodes = ceil(demand ÷ usable per node), for CPU and memory; take the largerBase count
Survive losing 1 of R racks: nodes × R ÷ (R − 1)Failure-domain headroom

Limits to check

max pods per node (kubelet default 110)Pod density
pod CIDR per node (e.g. /24 → 256 addresses)IPAM (lesson 06)
Upstream tested limits: 5,000 nodes, 150,000 pods per cluster, 110 pods per nodeFar above most designs, but real

05 · Control plane & node topology

Control plane options

3 stacked control-plane nodes (etcd on the same nodes)Default for most clusters
5 membersSurvives 2 failures; more write latency and cost
External etcd clusterIsolates etcd; more machines to run
Spread members across racks/roomsOne rack loss keeps quorum

ADR template

Title · Status · Context · Decision · Consequences · Alternatives consideredOne decision per record
docs/adr/ADR-00N-<slug>.md in the platform repoVersioned with the code

06 · Cluster networking

Sizing ranges

/16 pod CIDR with /24 per node → 256 nodes × 256 addresseskubeadm/controller-manager: --cluster-cidr, --node-cidr-mask-size
Service CIDR /16 → 65,536 ClusterIPs--service-cluster-ip-range (API server)
Calico default IPAM block /26; Cilium cluster-pool default /24 per nodeCNI-specific allocation
Ranges must not overlap: clusters, DCs, VPN users, partnersFuture routing and peering

Routing

Overlay (VXLAN/Geneve)Works anywhere; ~50 bytes overhead; pods not routable outside
Native routing + BGP to ToRPods routable; no encapsulation; needs network team

07 · Load balancing & ingress

Layers

L4: MetalLB (BGP/L2), kube-vip, or hardware/virtual LBStable IPs for ingress and the API server
L7: ingress controller or Gateway API implementationHost/path routing, TLS, retries, rate limits
externalTrafficPolicy: LocalKeep client source IP; only nodes with local endpoints receive traffic
PROXY protocolPass client IP through an L4 LB that terminates TCP

TLS options

Terminate at external LB/WAFCentral certs; LB sees plaintext
Terminate at ingress (cert-manager)Kubernetes-native certs; common default
Re-encrypt to pods / mesh mTLSEncrypted end to end inside the DC
Passthrough (SNI routing)App holds the key; no L7 features at ingress

08 · Customer-facing edge

Edge layers

DNS/GSLB → CDN → WAF + DDoS → L4 VIP → ingressOutside-in order
Rate limits: per client/API key at the gateway or ingressProtect backends from floods and abuse
API gateway: auth, quotas, versioning for partnersPartner-facing APIs

Isolation layers (namespace model)

RBAC + NetworkPolicy (default deny) + ResourceQuota/LimitRangeThe minimum set
Pod Security Admission (restricted) + policy engineWorkload hardening
Dedicated node pools / taints for sensitive tenantsKernel and noisy-neighbour isolation

09 · Storage & the stateful tier

Choices per data class

Operator-managed DB on local NVMe (app-level replication)Fast; the database replicates, not the storage
Replicated block storage (Ceph RBD, Longhorn)Storage-level replication; easier moves
External database service / DBA-run clusterExisting expertise, outside the cluster lifecycle

Replication across sites

SynchronousRPO ≈ 0; needs low round-trip latency (metro distances); writes slower
AsynchronousRPO = replication lag; any distance; possible data loss on failover
Backups (PITR) off-siteProtect against corruption and mistakes, not just hardware loss

10 · Observability & SLOs

Architecture

Per-cluster collection (Prometheus/agents, log and trace collectors)Close to the workloads
Central storage/query in the mgmt cluster or another siteSurvives a production cluster failure
Meta-monitoring + external dead man's switchKnow when monitoring itself is down
Synthetic checks from outside both DCsThe user's view

Alert catalogue columns

Name · signal/SLO · condition · severity · owner · runbook · notification routeOne row per alert

11 · Resilience & failure modes

FMEA

Component · failure mode · effect · S · O · D · RPN = S×O×D · mitigation · ownerOne row per failure mode
S, O, D scored 1–10 (10 = worst: severe, frequent, undetectable)Scoring scale
Re-score after mitigationShow risk reduced

Chaos experiments

Steady-state hypothesis → inject → observe → abort conditionsExperiment structure
Chaos Mesh / LitmusChaos (PodChaos, NetworkChaos, IOChaos…)Kubernetes-native fault injection
Start in nonprod, small blast radius, business hoursSafety rules

12 · Primary/secondary & failover

Patterns

Active-passive (warm standby)Secondary runs the platform + replicas; takes traffic on failover
Active-activeBoth sites serve traffic; needs multi-site data design
GSLB / DNS with health checks (low TTL)Switch traffic between sites

Failover essentials

Decide → fence old primary → promote → switch traffic → verifyOrder matters
Measure RPO (lag at failover) and RTO (time to healthy)Compare with NFR-04
Failback is a second, planned failoverResync, then switch back

13 · Capstone: design review

Document structure

1 Summary · 2 Requirements (NFR/SLO sheet) · 3 Architecture overviewFront
4 Compute & sizing · 5 Network & IPAM · 6 Traffic & edge · 7 Storage & dataCore design
8 Security & tenancy · 9 Observability · 10 Resilience (FMEA, chaos) · 11 DROperability
12 Cost · 13 Risks & open questions · Appendix: ADR index, evidence packClose

Traceability

NFR → design section → ADR → evidence (test/drill) → statusOne row per NFR