Platform engineering, learned the way production teaches it.
Structured tracks on Kubernetes, security, edge, AWS/EKS, delivery and observability, written for engineers who operate real systems. Every track moves from concepts to hands-on labs to real-world failure scenarios.
- 01ConceptsThe mental model: how it works and why.
- 02Hands-on labBuild it on a real cluster, step by step.
- 03ScenariosRealistic failures to reason through.
- 04PlaybookChecklists and runbooks to reuse.
- 05Q&AArchitect-level questions to test depth.
Linux & Scripting
The foundation under every platform: Linux, Bash and Python.
Kubernetes & Platform
Run, secure and design Kubernetes, from the first cluster to a fleet.
-
Available
Kubernetes Administration — Level by Level
Five levels from first cluster to fleet architect.
-
Available
Kubernetes Security & Hardening
Certificates, OIDC login, Ingress/Gateway TLS, policy and hardening.
-
Available
Cluster Design — Architect Track
From requirements to a defended design, one decision record at a time.
-
Available
Networking Deep Dive
One request, every layer, from Ethernet to Gateway.
-
Available
Kubernetes Storage & Data Protection
CSI, Longhorn, Rook-Ceph and backups you have actually restored.
-
Available
Service Mesh — Istio & Linkerd
mTLS, traffic management, and when not to use a mesh.
Bare Metal & Edge
Provision and operate infrastructure where there is no cloud to lean on.
Cloud — OpenStack, AWS & EKS
Private cloud with OpenStack, public cloud with AWS, and managed Kubernetes on EKS.
-
Available
OpenStack Private Cloud
How a private cloud really works: identity, compute, network and storage.
-
Available
AWS for Platform Engineers
The AWS building blocks every container platform relies on.
-
Available
Amazon EKS in Production with Terraform
Build, secure and operate EKS the way you'd run it for real.
Delivery & Infrastructure as Code
Ship infrastructure and apps declaratively, safely and repeatably.
Observability & Reliability
See what production is doing, and respond well when it breaks.
-
Available
Observability with OpenTelemetry
Metrics, logs and traces correlated into one workflow.
-
Available
Scaling Prometheus to Production
From one Prometheus to a federated, multi-tier metrics platform.
-
Available
Centralized Logging with EFK
Elasticsearch, Fluent Bit and Kibana on Kubernetes, done properly.
-
Available
SRE & Production Incident Response
Real-world failure scenarios and how to reason through them.
Incident Handling
Take the page, find the cause, fix it calmly, and stop it happening again.
AI-Assisted Engineering
Use AI where it genuinely helps operations, with guardrails.
Hands-on Projects
Guided end-to-end builds that combine several courses, the way interviews and real jobs ask for them.
No tracks match that search.
Where to start
Pick the path closest to your role. Each one strings tracks together in a sensible order.
Start here: foundations
Kubernetes Administrator → Platform Engineer
Cloud Platform on AWS
Edge & Bare-Metal Infrastructure
SRE & Production Reliability
Interview-ready: build it end to end
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.