Lesson 05 of 7 · Level 2 — Running it
Drift, rollback and recovery
Work with the reconcile loop instead of against it: what drift is and when to self-heal, pruning safely, rolling back with Git, pausing automation during incidents, and rebuilding an entire cluster from Git (and finding what was never in Git).
Drift: detect, then decide
Drift is any difference between the cluster and Git. Sources:
- People:
kubectl edit,kubectl scale, hotfixes in the console. - Other controllers: an HPA changing
replicas, a mutating webhook adding sidecars or defaults. - Objects nobody declared: created by hand, so they aren't in Git at all.
| Setting | Behaviour | Use for |
|---|---|---|
| Report only | Shows OutOfSync; a human decides | Early adoption, sensitive platform components |
| Self-heal | Reverts manual changes automatically | Applications once the team works through Git |
| Ignore differences | Excludes fields another controller owns | replicas under an HPA, webhook-injected fields |
The thermostat again. If someone opens a window, a good thermostat heats the room back up (self-heal). But if the cat's heat lamp is meant to be on, you tell the thermostat to ignore that corner (ignore differences) instead of fighting it all day.
Pruning: deleting from Git deletes from the cluster
With prune on, removing a manifest from Git removes the object from the cluster. That keeps clusters clean, and it's also how a bad commit can delete important things. Protect stateful objects:
metadata:
annotations:
argocd.argoproj.io/sync-options: Prune=false,Delete=false # Argo CD: never prune/delete this
# Flux equivalent: kustomize.toolkit.fluxcd.io/prune: disabled
Keep PVCs, namespaces with data, and CRDs (whose deletion removes every custom resource of that type) protected, and make deletions stand out in review.
Rolling back
Roll back in Git, not in the cluster:
$ git revert 7c1d2e3 # the promotion or config change that caused the problem
$ git push # fast-track PR if the branch is protected
The agent applies the previous state. Rolling back in the cluster (kubectl rollout undo, helm rollback) only lasts until the next reconcile, because Git still describes the bad version.
During an incident
Sometimes you need the agent to stop while you work:
- Pause reconciliation for the affected app only (Argo CD: disable automated sync; Flux:
flux suspend). - Make the emergency change, recorded in the incident channel.
- Put the fix into Git immediately after (same values the cluster now has).
- Resume reconciliation and confirm the app shows Synced with no diff.
A paused app that's forgotten becomes permanent drift; alert on apps that stay paused or OutOfSync longer than an agreed time.
Disaster recovery: rebuild from Git
If a cluster is lost, GitOps makes recovery mostly mechanical:
- Create a new cluster (Terraform).
- Install the agent and point it at the cluster's entry point in the config repo (the bootstrap:
flux bootstrap, or Argo CD + a root application). - The agent installs platform add-ons, then applications, in the configured order.
- Restore data separately: databases from backups or replicas, volumes with Velero or snapshots.
Test this regularly by rebuilding a staging cluster from scratch. Every manual step you need is something that isn't in Git yet: a secret created by hand, a CRD installed manually, a DNS record, a cluster-scoped role. Fix each one until the rebuild is fully automatic.
Try it: break, roll back, rebuild
- With an app synced by Argo CD or Flux on kind, make a manual
kubectl editchange with and without self-heal and observe both behaviours. - Add an HPA and watch
replicasfight with Git; fix it withignoreDifferencesor by removingreplicasfrom Git. - Push a bad change, then
git revertit and watch the agent recover. - Delete the kind cluster, create a new one, bootstrap the agent, and time how long until everything is back. List anything you had to do by hand.
Recap
- Drift comes from people, other controllers and undeclared objects; report, self-heal or ignore deliberately.
- Prune keeps clusters clean; protect stateful objects and review deletions.
- Roll back with git revert; pause reconciliation only briefly during incidents, then put the fix in Git.
- Rebuild from Git regularly; every manual step found is a gap to close.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.