GitOps Principles & Practice›07 · Progressive delivery

Lesson 07 of 7 · Level 2 — Running it

Progressive delivery

Make production rollouts safe automatically: canary and blue-green releases with Argo Rollouts or Flagger, traffic shifting through the ingress or a service mesh, and metric-driven analysis that promotes a healthy release or rolls back a bad one without a human.

Advanced
Key wordsprogressive deliverycanaryblue-greenArgo RolloutsRolloutAnalysisTemplateFlaggerCanarytraffic shiftingservice meshingressautomated rollbackSLO

From "deploy" to "release gradually"

A normal rolling update replaces pods quickly; within a minute every user is on the new version. Progressive delivery releases in controlled steps:

canary:  5% ──check──► 25% ──check──► 50% ──check──► 100%
                 ✗ at any check → all traffic back to stable, automatically

Testing a new recipe in a restaurant: serve it to one table first and watch their faces. If they like it, serve it to a few more tables, then the whole room. If the first table grimaces, you stop, and only one table had a bad meal.

Strategies

Strategy How Needs
Canary Shift a growing percentage of traffic to the new version Traffic splitting (ingress, service mesh or Gateway API), or replica-ratio approximation
Blue-green Run the full new version beside the old, switch traffic at once after checks Double capacity briefly; instant switch-back
Feature flags Deploy code dark, enable per user/percentage A flag service; complements the above

Argo Rollouts

A Rollout replaces a Deployment and adds steps and analysis:

apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata: { name: orders-api, namespace: shop }
spec:
  replicas: 6
  selector: { matchLabels: { app: orders-api } }
  template:
    metadata: { labels: { app: orders-api } }
    spec:
      containers:
        - name: api
          image: ghcr.io/acme/orders-api@sha256:3f1c9a0b7e2d4c5a8b9e1f2d3c4b5a6978e1d2c3b4a5968778695a4b3c2d1e0f
  strategy:
    canary:
      canaryService: orders-api-canary
      stableService: orders-api-stable
      trafficRouting:
        nginx: { stableIngress: orders-api }     # or a service mesh / Gateway API plugin
      steps:
        - setWeight: 5
        - analysis: { templates: [ { templateName: error-rate } ] }
        - setWeight: 25
        - pause: { duration: 10m }
        - analysis: { templates: [ { templateName: error-rate } ] }
        - setWeight: 50
        - pause: { duration: 10m }
---
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata: { name: error-rate, namespace: shop }
spec:
  metrics:
    - name: error-rate
      interval: 1m
      count: 5
      failureLimit: 1
      successCondition: result[0] < 0.01
      provider:
        prometheus:
          address: http://prometheus.monitoring:9090
          query: |
            sum(rate(http_requests_total{app="orders-api",version="canary",code=~"5.."}[2m]))
            /
            sum(rate(http_requests_total{app="orders-api",version="canary"}[2m]))

If the error rate of the canary exceeds 1%, the analysis fails, the rollout aborts and traffic returns to the stable version.

Flagger

Flagger (a Flux project, also usable with Argo CD) takes a different approach: you keep your normal Deployment, and a Canary object tells Flagger to manage the progressive rollout, generating the canary and primary objects and routing traffic through the ingress or mesh, with built-in request-success-rate and latency checks.

Argo Rollouts Flagger
Object Rollout replaces Deployment Canary wraps an existing Deployment
Ecosystem Argo family, strong UI/CLI Flux family, works with many meshes and ingresses
Analysis AnalysisTemplate with many providers Built-in metrics + custom metric templates

How it fits GitOps

  • Git holds the Rollout/Canary spec, including the steps and analysis.
  • A promotion PR changes the image digest; the GitOps agent applies it; the rollout controller performs the steps.
  • If the rollout aborts, Git still names the new version. Revert the promotion in Git so the declared and running versions agree again.

Good analysis

  • Measure user-facing signals (error rate, latency) of the canary, ideally compared to stable at the same time.
  • Make sure there's enough traffic at small weights for the numbers to mean something (or use longer steps).
  • Include a business metric where possible (orders placed, payments succeeded).
  • Test the rollback path: deploy a deliberately broken version to staging and watch it abort.

Try it: an automatic rollback

  1. On kind with ingress-nginx and Prometheus, install Argo Rollouts and convert a Deployment to the Rollout above.
  2. Deploy version A, then promote version B through Git; watch the canary steps with kubectl argo rollouts get rollout --watch.
  3. Deploy a version C that returns 500 errors for 10% of requests; watch the analysis fail and the rollout abort.
  4. Revert the promotion in Git so Git and the cluster agree again.

Recap

  • Progressive delivery = small exposure, measured steps, automatic rollback.
  • Canary (weights), blue-green (switch), plus feature flags.
  • Argo Rollouts (Rollout + AnalysisTemplate) or Flagger (Canary around a Deployment).
  • Git declares version and strategy; the controller executes it; revert in Git after an abort.

This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.