Learning Hub / Observability & Reliability
Scaling Prometheus to Production
Advanced → Architect10 lessonsAvailable
From a single kube-prometheus-stack to a federated, multi-tier, cost-aware metrics platform, comparing Thanos and Grafana Mimir on MinIO object storage.
You'll meetPrometheusTSDBcardinalityThanosMimirMinIOdownsamplingremote-writefederation
Start lesson 01 →
What you'll be able to do
- Explain why a single Prometheus doesn't scale (TSDB internals)
- Run Thanos or Mimir with hot/cold retention
- Choose between federation, remote-write and global query
Before you start
Prometheus & Grafana basics.
How it works
Each lesson: plain-language idea → how it really works → hands-on. Each section ends with a cheat sheet & self-check.
Curriculum
Lessons marked “Read” are ready; the rest are on the way.
Modules
- 01kube-prometheus-stack baselineOperator, ServiceMonitors, rulesRead →
- 02Why Prometheus doesn't scaleTSDB internals, cardinalityRead →
- 03MinIO object storageThe foundation for long-term metricsRead →
- 04Thanos: sidecar, store, querierGlobal viewRead →
- 05Thanos: compactor, downsampling, hot/coldRetention that you can affordRead →
- 06Retention by data priorityNot all metrics are equalRead →
- 07Grafana MimirThe other modelRead →
- 08Thanos vs MimirA head-to-head decision guideRead →
- 09Federation & multi-clusterRemote-write vs global queryRead →
- 10Capstone: cost, HA & the platform viewA 500k active-series designRead →
- 📋Cheat sheet & self-checkEvery command from this section on one page, then 30 questions to check yourself.Open →
Real-world scenarios
Work through each one: symptom → misleading signal → evidence → root cause → prevention.
The cardinality explosion
One label that took down monitoring.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.