Learning Hub / Observability & Reliability
Observability with OpenTelemetry
Practitioner → Advanced7 lessonsAvailable
Metrics tell you what is broken; observability tells you why. Correlate metrics, logs and traces through OpenTelemetry into one investigation flow.
You'll meetmetricslogstracesOpenTelemetryCollectorTempoLokiexemplarsSLOsburn rate
Start lesson 01 →
What you'll be able to do
- Deploy an OpenTelemetry Collector pipeline
- Jump from a metric spike to the exact failing trace
- Write SLOs with burn-rate alerting
Before you start
Kubernetes basics, Prometheus fundamentals.
How it works
Each lesson: plain-language idea → how it really works → hands-on. Each section ends with a cheat sheet & self-check.
Curriculum
Lessons marked “Read” are ready; the rest are on the way.
Modules
- 01Observability fundamentalsThree pillars, one collectorRead →
- 02OpenTelemetry Collector deep diveReceivers, processors, exportersRead →
- 03Distributed tracingTempo/Jaeger, samplingRead →
- 04Structured logging & LokiLabels, cardinalityRead →
- 05Signal correlationExemplars, trace IDs in logsRead →
- 06SLOs, SLIs & burn-rate alertingAlerts people trustRead →
- 07Production observability patternsCost, retention, multi-tenancyRead →
- 📋Cheat sheet & self-checkEvery command from this section on one page, then 21 questions to check yourself.Open →
Real-world scenarios
Work through each one: symptom → misleading signal → evidence → root cause → prevention.
The dashboard is green while users see errors
What your health checks aren't checking.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.