Lesson 07 of 7 · Level 2 — Running it
Progressive delivery
Make production rollouts safe automatically: canary and blue-green releases with Argo Rollouts or Flagger, traffic shifting through the ingress or a service mesh, and metric-driven analysis that promotes a healthy release or rolls back a bad one without a human.
From "deploy" to "release gradually"
A normal rolling update replaces pods quickly; within a minute every user is on the new version. Progressive delivery releases in controlled steps:
canary: 5% ──check──► 25% ──check──► 50% ──check──► 100%
✗ at any check → all traffic back to stable, automatically
Testing a new recipe in a restaurant: serve it to one table first and watch their faces. If they like it, serve it to a few more tables, then the whole room. If the first table grimaces, you stop, and only one table had a bad meal.
Strategies
| Strategy | How | Needs |
|---|---|---|
| Canary | Shift a growing percentage of traffic to the new version | Traffic splitting (ingress, service mesh or Gateway API), or replica-ratio approximation |
| Blue-green | Run the full new version beside the old, switch traffic at once after checks | Double capacity briefly; instant switch-back |
| Feature flags | Deploy code dark, enable per user/percentage | A flag service; complements the above |
Argo Rollouts
A Rollout replaces a Deployment and adds steps and analysis:
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata: { name: orders-api, namespace: shop }
spec:
replicas: 6
selector: { matchLabels: { app: orders-api } }
template:
metadata: { labels: { app: orders-api } }
spec:
containers:
- name: api
image: ghcr.io/acme/orders-api@sha256:3f1c9a0b7e2d4c5a8b9e1f2d3c4b5a6978e1d2c3b4a5968778695a4b3c2d1e0f
strategy:
canary:
canaryService: orders-api-canary
stableService: orders-api-stable
trafficRouting:
nginx: { stableIngress: orders-api } # or a service mesh / Gateway API plugin
steps:
- setWeight: 5
- analysis: { templates: [ { templateName: error-rate } ] }
- setWeight: 25
- pause: { duration: 10m }
- analysis: { templates: [ { templateName: error-rate } ] }
- setWeight: 50
- pause: { duration: 10m }
---
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata: { name: error-rate, namespace: shop }
spec:
metrics:
- name: error-rate
interval: 1m
count: 5
failureLimit: 1
successCondition: result[0] < 0.01
provider:
prometheus:
address: http://prometheus.monitoring:9090
query: |
sum(rate(http_requests_total{app="orders-api",version="canary",code=~"5.."}[2m]))
/
sum(rate(http_requests_total{app="orders-api",version="canary"}[2m]))
If the error rate of the canary exceeds 1%, the analysis fails, the rollout aborts and traffic returns to the stable version.
Flagger
Flagger (a Flux project, also usable with Argo CD) takes a different approach: you keep your normal Deployment, and a Canary object tells Flagger to manage the progressive rollout, generating the canary and primary objects and routing traffic through the ingress or mesh, with built-in request-success-rate and latency checks.
| Argo Rollouts | Flagger | |
|---|---|---|
| Object | Rollout replaces Deployment |
Canary wraps an existing Deployment |
| Ecosystem | Argo family, strong UI/CLI | Flux family, works with many meshes and ingresses |
| Analysis | AnalysisTemplate with many providers |
Built-in metrics + custom metric templates |
How it fits GitOps
- Git holds the Rollout/Canary spec, including the steps and analysis.
- A promotion PR changes the image digest; the GitOps agent applies it; the rollout controller performs the steps.
- If the rollout aborts, Git still names the new version. Revert the promotion in Git so the declared and running versions agree again.
Good analysis
- Measure user-facing signals (error rate, latency) of the canary, ideally compared to stable at the same time.
- Make sure there's enough traffic at small weights for the numbers to mean something (or use longer steps).
- Include a business metric where possible (orders placed, payments succeeded).
- Test the rollback path: deploy a deliberately broken version to staging and watch it abort.
Try it: an automatic rollback
- On kind with ingress-nginx and Prometheus, install Argo Rollouts and convert a Deployment to the Rollout above.
- Deploy version A, then promote version B through Git; watch the canary steps with
kubectl argo rollouts get rollout --watch. - Deploy a version C that returns 500 errors for 10% of requests; watch the analysis fail and the rollout abort.
- Revert the promotion in Git so Git and the cluster agree again.
Recap
- Progressive delivery = small exposure, measured steps, automatic rollback.
- Canary (weights), blue-green (switch), plus feature flags.
- Argo Rollouts (
Rollout+AnalysisTemplate) or Flagger (Canaryaround a Deployment). - Git declares version and strategy; the controller executes it; revert in Git after an abort.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.