Kubernetes Administration — Level by Level
A single path that takes you from how Kubernetes works to operating it at production scale. Each level ends with a checkpoint lab you must pass before moving on, so the skills stack instead of scattering.
What you'll be able to do
- Explain every control-plane and node component and what breaks when each one fails
- Install, upgrade, back up and restore clusters with kubeadm and RKE2
- Tune scheduling, resources and autoscaling for real workloads
- Design multi-cluster fleets with Cluster API and GitOps
Before you start
Linux command line, containers (Docker/Podman), basic networking.
How it works
Each lesson: plain-language idea → how it really works → hands-on. Each section ends with a cheat sheet & self-check.
Curriculum
Lessons marked “Read” are ready; the rest are on the way.
Level 1 — Foundations
- 01Set up your practice labkind on your laptop, kubectl, and your first cluster in 5 minutesRead →
- 02Architecture & the control planeAPI server, etcd, scheduler, controllers, kubelet: the reconcile loopRead →
- 03Pods, Deployments & workload typesDeployments, StatefulSets, DaemonSets, Jobs and when to use eachRead →
- 04Services, DNS & basic networkingClusterIP, NodePort, LoadBalancer, CoreDNSRead →
- 05Ingress: getting traffic inInstall ingress-nginx, expose it, and follow a request to the right pod and backRead →
- 06Configuration & SecretsConfigMaps, Secrets, env vs volume mountsRead →
- 07Storage basicsVolumes, PV/PVC, StorageClassesRead →
- 08Checkpoint: deploy a 3-tier appShip, expose and debug a real applicationRead →
- 📋Cheat sheet & self-checkEvery command from this section on one page, then 35 questions to check yourself.Open →
Level 2 — Operator
- 09Build a cluster with kubeadmHA-ready topology, certificates, join flowRead →
- 10kubeadm at scale: immutable imagesGolden images per role, boot, IP assignment and automatic joinRead →
- 11Bare-metal load balancing with MetalLBLoadBalancer Services without a cloud: L2 and BGP modesRead →
- 12Upgrades & version skewControl plane first, drain/cordon, skew policyRead →
- 13etcd backup & restoreSnapshots, restore drills, quorum lossRead →
- 14Scheduling in depthTaints, tolerations, affinity, topology spread, priorityRead →
- 15RBAC & service accountsRoles, bindings, least privilege for humans and botsRead →
- 16Troubleshooting: nodes, pods, networkingA repeatable method, not a list of commandsRead →
- 📋Cheat sheet & self-checkEvery command from this section on one page, then 30 questions to check yourself.Open →
Level 3 — Application delivery
- 17Deploy an app end to endDeployment → Service → Ingress, step by stepRead →
- 18Namespaces & an app with a databaseApps across namespaces, cross-namespace DNS, a PostgreSQL backendRead →
- 19HTTPS & TLS terminationCertificates, cert-manager, edge vs passthrough vs re-encryptRead →
- 20Ingress & egress gatewaysControlled entry and exit points: Gateway API, egress policies and gatewaysRead →
- 21Release strategies: rolling, blue-green, canaryShip changes with a small blast radius and a fast way backRead →
- 📋Cheat sheet & self-checkEvery command from this section on one page, then 15 questions to check yourself.Open →
Level 4 — Production
- 22HA control plane & load balancingStacked vs external etcd, API VIPs, failure domainsRead →
- 23Resource management & QoSRequests, limits, QoS classes, eviction, LimitRanges, quotasRead →
- 24AutoscalingHPA, VPA, Cluster Autoscaler, Karpenter, KEDARead →
- 25Cluster lifecycle at scaleRolling upgrades across many clusters, maintenance windowsRead →
- 26Performance tuningAPI server and etcd tuning, kubelet flags, large-cluster limitsRead →
- 27Disaster recoveryVelero, etcd, GitOps rebuild, RPO/RTO you can proveRead →
- 📋Cheat sheet & self-checkEvery command from this section on one page, then 23 questions to check yourself.Open →
Level 5 — Architect
- 28Multi-cluster & fleet managementCluster API, Rancher, fleet GitOpsRead →
- 29Multi-tenancy modelsNamespaces vs vClusters vs clusters-per-tenantRead →
- 30Internal Developer Platform designGolden paths, self-service, platform APIsRead →
- 31SLOs for the platform itselfWhat 'the cluster is healthy' really meansRead →
- 32Capstone: platform design reviewDefend a design against failure modes and costRead →
- 📋Cheat sheet & self-checkEvery command from this section on one page, then 16 questions to check yourself.Open →
Real-world scenarios
Work through each one: symptom → misleading signal → evidence → root cause → prevention.
Half the nodes run the new version and cross-version traffic fails. Diagnose before you roll back.
Writes are failing cluster-wide. Recover without losing state.
Scheduling constraints, fragmentation, or quota? Prove which one.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.