Modules · wrap-up
Cheat sheet & self-check
30 questions across 10 lessons. Each answer links back to the lesson it came from.
Pick an answer to see if you got it, and why.
Q1. You created a ServiceMonitor, but Prometheus doesn't scrape the target. What's the classic cause with kube-prometheus-stack?
Show answer
B. Check Prometheus's serviceMonitorSelector/NamespaceSelector. The chart value serviceMonitorSelectorNilUsesHelmValues=false makes it select all.
From lesson 01 · kube-prometheus-stack baselineQ2. What does kube-state-metrics provide?
Show answer
B. node-exporter covers the machine; kube-state-metrics covers object state from the API; cAdvisor (kubelet) covers container usage.
From lesson 01 · kube-prometheus-stack baselineQ3. Why use rate() on a counter before summing?
Show answer
B. Always rate first, then sum: sum(rate(x[5m])), never rate(sum(x)).
From lesson 01 · kube-prometheus-stack baselineQ4. What mainly drives Prometheus memory usage?
Show answer
B. Each active series has in-memory structures, recent samples and index entries. Double the series ≈ much more memory.
From lesson 02 · Why Prometheus doesn't scaleQ5. A metric has labels method (5 values), status (10), endpoint (50) and pod (40). Roughly how many series can it create?
Show answer
B. Cardinality multiplies across labels. Adding one unbounded label (user ID) makes it effectively infinite.
From lesson 02 · Why Prometheus doesn't scaleQ6. Why isn't 'just keep a year of data locally' a good plan for one Prometheus?
Show answer
B. Long-term storage and a global view are exactly what Thanos and Mimir add (lessons 04–07).
From lesson 02 · Why Prometheus doesn't scaleQ7. Why do Thanos and Mimir store long-term metrics in object storage?
Show answer
B. Immutable TSDB blocks fit object storage perfectly; compute (queriers, store gateways) scales separately.
From lesson 03 · MinIO object storageQ8. What does erasure coding give an S3-compatible cluster like MinIO or Ceph?
Show answer
B. For example, 8 data + 4 parity tolerates 4 lost drives with 1.5× overhead instead of 3× for triple replication.
From lesson 03 · MinIO object storageQ9. Your team plans to use MinIO's open-source edition. What should you check first?
Show answer
B. Thanos and Mimir only need an S3-compatible API, so choose the store you can support long term.
From lesson 03 · MinIO object storageQ10. Why are external labels (like cluster) essential with Thanos?
Show answer
B. Without unique external labels, series from different sources collide in the global view.
From lesson 04 · Thanos: sidecar, store, querierQ11. Two Prometheus replicas scrape the same targets. How does the Thanos querier show one clean series?
Show answer
B. Pass --query.replica-label with the replica label name (prometheus_replica for the Operator).
From lesson 04 · Thanos: sidecar, store, querierQ12. What does the store gateway do?
Show answer
B. Sidecars cover recent hours; store gateways cover everything already in the bucket.
From lesson 04 · Thanos: sidecar, store, querierQ13. How many compactors should run against one bucket?
Show answer
B. The compactor is a singleton by design. Run it as a single-replica StatefulSet with enough disk.
From lesson 05 · Thanos: compactor, downsampling, hot/coldQ14. Does downsampling reduce storage usage?
Show answer
B. Keep raw data short, downsampled data long: that's where the savings come from.
From lesson 05 · Thanos: compactor, downsampling, hot/coldQ15. When are blocks downsampled by default?
Show answer
B. Downsampling needs big enough compacted blocks; that's why it happens after these ages.
From lesson 05 · Thanos: compactor, downsampling, hot/coldQ16. You need 2 years of request-rate trends per service, but raw per-pod series are only useful for 2 weeks. What do you do?
Show answer
B. Aggregates have far fewer series; keeping them long costs a fraction of keeping raw data.
From lesson 06 · Retention by data priorityQ17. What's the cheapest place to get rid of a metric nobody uses?
Show answer
B. Dropping early saves memory, network, storage and query time everywhere downstream.
From lesson 06 · Retention by data priorityQ18. How do you know which metrics are unused?
Show answer
B. Usage analysis often finds that a large share of stored series are never queried.
From lesson 06 · Retention by data priorityQ19. How does data get into Mimir?
Show answer
B. Mimir is push-based. Local Prometheus/agents do the scraping; Mimir stores and serves.
From lesson 07 · Grafana MimirQ20. What does the ingester replication factor of 3 mean?
Show answer
B. Recent (not yet uploaded) data lives in ingesters; replication protects it. Object storage has its own durability.
From lesson 07 · Grafana MimirQ21. Why are per-tenant limits important in Mimir?
Show answer
B. Multi-tenancy without limits is a shared failure domain.
From lesson 07 · Grafana MimirQ22. Which is a strong reason to pick Thanos with sidecars?
Show answer
B. Sidecars bolt on to what you have. Mimir changes the architecture to push-based central storage.
From lesson 08 · Thanos vs MimirQ23. Which is a strong reason to pick Mimir?
Show answer
B. Multi-tenancy, limits and horizontal scale of every component are Mimir's core strengths.
From lesson 08 · Thanos vs MimirQ24. What's a real risk of the sidecar model for remote sites?
Show answer
B. Push models (Receive/Mimir with agents) buffer and send; pull models need connectivity at query time.
From lesson 08 · Thanos vs MimirQ25. What is Prometheus federation good for?
Show answer
B. Federating raw series at scale overloads both sides. Use it for aggregates; use remote write or global query for the rest.
From lesson 09 · Federation & multi-clusterQ26. Why run Prometheus in agent mode on edge clusters?
Show answer
B. Edge clusters then only need to scrape and forward; queries and alerting happen centrally.
From lesson 09 · Federation & multi-clusterQ27. An edge site loses its uplink for 6 hours. With remote write, what's at risk?
Show answer
B. Remote write reads from the WAL. Plan buffering for your realistic outage lengths, and alert on growing pending samples.
From lesson 09 · Federation & multi-clusterQ28. 500,000 active series scraped every 30 seconds produce roughly how many samples per second?
Show answer
B. 500,000 / 30 ≈ 16,667 samples per second.
From lesson 10 · Capstone: cost, HA & the platform viewQ29. Why run a separate, small 'meta-monitoring' Prometheus?
Show answer
B. Monitoring can't reliably report its own death. Combine with a dead man's switch (Watchdog) to an external service.
From lesson 10 · Capstone: cost, HA & the platform viewQ30. What makes a capacity estimate trustworthy in a design review?
Show answer
B. Reviewers can challenge assumptions; measurements anchor them.
From lesson 10 · Capstone: cost, HA & the platform view