Argo CD — Level by Level›08 · Troubleshooting Argo CD

Lesson 08 of 8 · Argo CD in practice

Troubleshooting Argo CD

Fix the Argo CD problems that come up every week: apps that never become Synced, syncs that fail or hang, Degraded or stuck Progressing health, diffs that never go away, manifest-generation errors, and slow or overloaded Argo CD components.

Practitioner → Advanced
Key wordsArgo CD troubleshootingOutOfSyncDegradedProgressingMissingsync failedComparisonErrorignoreDifferenceshard refreshrepo-serverapplication-controllersync wavesfinalizers

Two statuses, read separately

Every application has two independent states:

Meaning Values
Sync status Does the cluster match Git? Synced, OutOfSync, Unknown
Health status Are the resources working? Healthy, Progressing, Degraded, Suspended, Missing

"Synced but Degraded" means Argo CD applied exactly what Git says, and what Git says doesn't work. "OutOfSync but Healthy" means something runs fine but doesn't match Git. The fix is in different places.

A delivery app with two ticks: "delivered to the right address" (sync) and "the customer is happy with the meal" (health). A meal can arrive at the right door and still be cold.

OutOfSync that won't go away

Cause Check Fix
Another controller owns a field (HPA replicas, webhook-injected sidecars, defaults) argocd app diff shows the same field every time Remove the field from Git, or ignoreDifferences for it
API server normalises values (cpu: 1000m vs 1, empty vs omitted) Diff shows equivalent values Write the normalised form in Git; server-side diff helps
Resource created by hand exists in the namespace Shows as extra (with prune off) Add it to Git or delete it
Wrong target revision or path argocd app get source section Fix the Application spec

Example: let the HPA own replicas:

spec:
  ignoreDifferences:
    - group: apps
      kind: Deployment
      jsonPointers: [ /spec/replicas ]

Sync fails or hangs

Error Cause Fix
ComparisonError / failed to generate manifests Helm/Kustomize can't render: bad values, missing chart version, private repo credentials Repo-server logs; argocd app manifests; render locally with the same versions
could not find the requested resource CRD not installed yet Sync waves (CRDs first), a separate CRD app, or SkipDryRunOnMissingResource=true
namespaces "x" not found Namespace doesn't exist CreateNamespace=true sync option, or declare it in Git with an earlier wave
forbidden Argo CD (or the AppProject) isn't allowed that resource kind or namespace AppProject whitelists, destination namespaces, cluster RBAC
Immutable field error (e.g. a Job's template, a selector) Kubernetes won't update that field Replace=true for that resource, or delete and recreate it deliberately
Sync stuck "Running" A hook or wave waiting on a resource that never becomes healthy argocd app get --show-operation; fix the resource; argocd app terminate-op
Deleting an app hangs Finalizers waiting on resources that can't be deleted Check the resources' own finalizers first; remove the app's finalizer only as a last resort

Degraded or Progressing forever

Health comes from the resources. Go to Kubernetes:

$ argocd app resources orders-api | grep -v Healthy
$ kubectl -n shop rollout status deploy/orders-api
$ kubectl -n shop get pods
$ kubectl -n shop describe pod <pod>        # events: probes, image pulls, scheduling

Common causes: failing readiness probes, image pull errors (wrong digest, registry auth), insufficient resources, a PVC that can't bind. Custom resources without a health check show as Healthy or Progressing according to Argo CD's built-in rules; add custom health checks (Lua) for CRDs you depend on.

Argo CD itself is slow or struggling

Symptom Likely component Tuning
"Refreshing" for minutes, manifest timeouts repo-server More replicas, bigger resources, cache; reduce huge Helm dependencies
Status updates lag, high CPU application-controller Shard the controller across clusters, raise processors; reduce resource tracking of noisy objects
Git provider rate limits Polling many repos often Use webhooks from the Git server and a longer polling interval
UI/API slow argocd-server More replicas; check SSO/Redis

Monitor Argo CD's own Prometheus metrics (sync failures, reconciliation duration, repo-server errors) and alert on apps OutOfSync or Degraded for longer than an agreed time.

Try it: break it five ways

On a kind cluster with Argo CD and a test app:

  1. Add an HPA without removing replicas from Git; fix the permanent OutOfSync with ignoreDifferences.
  2. Add a custom resource whose CRD isn't installed; fix it with a sync wave that installs the CRD first.
  3. Reference a Helm chart version that doesn't exist; find the error in the repo-server logs.
  4. Deploy an image tag that doesn't exist; follow Degraded down to the pod's ImagePullBackOff event.
  5. Change a Job's pod template in Git; resolve the immutable-field error with Replace=true.

Recap

  • Read sync and health separately; they're fixed in different places.
  • Persistent OutOfSync: two owners for one field or normalisation; use ignoreDifferences or fix Git.
  • Sync errors: rendering (repo-server), ordering (CRDs, namespaces), permissions (AppProject), immutable fields.
  • Health problems live in the workload; Argo CD performance lives in repo-server, controller and polling.

This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.