Edge Kubernetes & Zero-Touch Provisioning cheat sheet
88 commands from every lesson of Edge Kubernetes & Zero-Touch Provisioning, on one page.
The ZTP promise
Rack, cable, power on | The only hands-on work at the site |
BMC/PXE → OS → Kubernetes → workloads | Everything else automated and remote |
Site definition in Git | Hardware inventory + cluster spec + workloads |
Lifecycle
Day 0 | Design, inventory, images, site definition |
Day 1 | Provision: bare metal → cluster → apps |
Day 2 | Upgrade, scale, repair, rebuild, retire |
Planes
Management plane (central) | Git, registries, cluster lifecycle controllers, observability |
Provisioning at the site | DHCP/TFTP/HTTP boot + BMC access: needs L2 or DHCP relay |
Workload clusters (per site) | Run the applications; pull config and images |
Site networks
BMC / OOB VLAN | Redfish/IPMI; isolated, no internet |
Provisioning VLAN | PXE/iPXE, DHCP; can be the node network |
Node / cluster network | Kubernetes API VIP, node IPs |
Workload / uplink | Services, backhaul to the centre |
Shapes
1 node | No HA; cheapest; a node failure = site outage |
3 nodes, compact (CP + workloads on all) | Tolerates 1 node failure; common edge default |
3 CP + N workers | More isolation and capacity; more hardware |
2 nodes | No etcd fault tolerance without an external witness |
Settings
kubectl taint nodes <n> node-role.kubernetes.io/control-plane:NoSchedule- | Allow workloads on control-plane nodes (compact) |
kubelet: systemReserved / kubeReserved / evictionHard | Protect the OS and kubelet from workloads |
etcd quorum = floor(n/2) + 1 | 3 members → survives 1 failure |
Redfish (curl)
curl -sku user:pass https://<bmc>/redfish/v1/Systems | List systems (IDs vary by vendor: 1, System.Embedded.1, …) |
PATCH …/Systems/<id> {"Boot":{"BootSourceOverrideTarget":"Pxe","BootSourceOverrideEnabled":"Once"}} | Boot from network once |
POST …/Systems/<id>/Actions/ComputerSystem.Reset {"ResetType":"ForceRestart"} | Power cycle |
POST …/Managers/<id>/VirtualMedia/<cd>/Actions/VirtualMedia.InsertMedia {"Image":"http://…/boot.iso"} | Mount an ISO remotely |
IPMI & PXE
ipmitool -I lanplus -H <bmc> -U user -P pass chassis bootdev pxe options=efiboot | Legacy: boot from network (UEFI) |
ipmitool -I lanplus -H <bmc> -U user -P pass power cycle | Legacy: power cycle |
DHCP option 93 (client arch) → ipxe.efi vs undionly.kpxe | Serve the right iPXE binary |
Tinkerbell
Smee | DHCP, TFTP and iPXE scripts (formerly Boots) |
HookOS | In-memory OS that runs workflow actions (formerly Hook) |
Tink server/controller + worker | Workflow engine; workers run actions as containers |
Tootles | Metadata service for cloud-init (formerly Hegel) |
Rufio | BMC control as Kubernetes resources (Machine, Job, Task) |
kubectl get hardware,templates,workflows -A | Inspect provisioning state |
Metal3
Bare Metal Operator + Ironic | Manage hosts through BareMetalHost resources |
kubectl get baremetalhosts -A | States: registering → inspecting → available → provisioning → provisioned |
bmc.address: redfish-virtualmedia://<bmc>/redfish/v1/Systems/1 | BMC driver + address |
CAPM3 | Cluster API provider for Metal3 |
Formats
RAW (.raw.gz / .raw.xz) | Streamed straight onto the disk (e.g. image2disk); bare metal |
QCOW2 | VMs, and Ironic can convert/write it |
ISO | Installer or live boot via virtual media/USB |
qemu-img convert -f qcow2 -O raw in.qcow2 out.raw | Convert between formats |
Build & first boot
image-builder (kubernetes-sigs) / Packer | Reproducible node images with Kubernetes components |
cloud-init NoCloud: user-data, meta-data, network-config | First-boot configuration |
Ignition (Flatcar, Fedora CoreOS) | First-boot provisioning for those OSes |
sha256sum image.raw.gz > image.raw.gz.sha256; cosign sign-blob … | Checksum and sign images |
Moving artifacts
skopeo copy --all docker://registry.k8s.io/pause:3.10 docker://harbor.local/mirror/pause:3.10 | Copy an image (all architectures) |
skopeo copy --all docker://… oci-archive:bundle/pause.tar | Image to a file for transfer |
helm pull oci://… --version X && helm push chart-X.tgz oci://harbor.local/charts | Mirror a Helm chart |
oras copy <src> <dst> | Copy any OCI artifact (SBOMs, signatures, files) |
Pointing nodes at mirrors
/etc/rancher/rke2/registries.yaml (mirrors + configs) | RKE2/K3s mirror configuration |
registryMirrorConfiguration (EKS Anywhere cluster spec) | EKS-A mirror + CA |
containerd: /etc/containerd/certs.d/<registry>/hosts.toml | Plain containerd mirror config |
Create
eksctl anywhere generate clusterconfig site042 --provider tinkerbell > site042.yaml | Start a cluster spec |
hardware.csv: hostname,bmc_ip,bmc_username,bmc_password,mac,ip_address,netmask,gateway,nameservers,labels,disk | Hardware inventory (labels like type=cp) |
eksctl anywhere create cluster --hardware-csv hardware.csv -f site042.yaml | Provision machines and form the cluster |
Day 2
eksctl anywhere upgrade plan cluster -f site042.yaml | See available component upgrades |
eksctl anywhere upgrade cluster -f site042.yaml | Upgrade (Kubernetes version and components) |
eksctl anywhere generate hardware -z hardware.csv > hardware.yaml | Hardware objects for adding machines later |
kubectl get clusters.anywhere.eks.amazonaws.com,machines -A | Cluster and CAPI machine status |
RKE2
curl -sfL https://get.rke2.io | sh - && systemctl enable --now rke2-server | Install and start a server (online) |
/etc/rancher/rke2/config.yaml | Node configuration (token, tls-san, profile, server…) |
server: https://<vip-or-first-server>:9345 | Join an existing cluster (supervisor port) |
/etc/rancher/rke2/rke2.yaml + /var/lib/rancher/rke2/bin/kubectl | Admin kubeconfig and bundled kubectl |
rke2 etcd-snapshot save --name pre-upgrade | Manual etcd snapshot |
Rancher & Elemental
Cluster Management → Create → Custom | Register existing machines with a registration command |
MachineRegistration / MachineInventory / SeedImage (Elemental) | Onboard and manage edge OS + nodes |
Fleet: GitRepo + cluster labels | GitOps across many clusters (built into Rancher) |
Targeting
Cluster labels: site=042, size=s, region=eu, wave=canary | Describe clusters; select by label |
Fleet: GitRepo targets[].clusterSelector | Rancher Fleet targeting |
Argo CD: ApplicationSet cluster generator selector | Argo CD targeting |
Flux: a Kustomization per cluster from its own path | Flux per-cluster entry point |
Rollout
wave=lab → canary → early → all | Staged rollout by label |
Pin revisions per wave (tags/branches) | Promote by moving a pointer |
Fleet/Argo status per cluster | Which revision each site runs |
Networking
kube-vip (ARP/BGP) | API server VIP and/or Service LoadBalancer IPs |
MetalLB (L2 or BGP) | LoadBalancer Services on bare metal |
Local DNS forwarder + NTP server (or GPS/PTP clock) | Sites keep working offline |
Multus + SR-IOV device plugin | Extra, high-performance pod interfaces (telecom) |
Storage & security
local-path / TopoLVM (LVM-backed local PVs) | Single-node or node-local storage |
Longhorn (3 nodes) | Replicated block storage for small clusters |
LUKS + TPM2 (e.g. systemd-cryptenroll / Clevis) | Encrypted disks that unlock only on the original hardware |
Secure Boot + measured boot | Only signed boot chains; tampering is detectable |
Backups
rke2 etcd-snapshot save / automatic snapshots (+ S3 upload) | RKE2 etcd backups |
etcdctl snapshot save (kubeadm-based clusters, e.g. EKS-A control plane) | Generic etcd snapshot |
velero backup create site042-daily --include-namespaces shop | Kubernetes objects + volume data |
Compliance
kube-bench run --targets master,node | CIS Kubernetes Benchmark checks |
Policy engine reports (Kyverno/Gatekeeper) | Continuous config compliance |
Per-site inventory: OS image, K8s, bundle, firmware, Secure Boot, encryption | Evidence and drift detection |