Lesson 05 of 13 · Level 2 — Docker Compose
Compose for real environments
Make a Compose stack start in the right order, recover from crashes, stay inside resource limits and keep passwords out of environment variables. Use override files and profiles to run the same stack in dev, CI and on a server.
Health checks and start order
depends_on: [db] only waits for the db container to start, not for PostgreSQL to be ready, so the API often crashes on its first connection attempt. Add a health check and wait for it:
services:
db:
image: postgres:16
healthcheck:
test: ["CMD-SHELL", "pg_isready -U app -d shop"]
interval: 5s
timeout: 3s
retries: 10
start_period: 20s # failures during start-up don't count
api:
image: shop/api:1.4
depends_on:
db:
condition: service_healthy
restart: true # restart api if db is restarted by Compose
healthcheck:
test: ["CMD", "curl", "-fsS", "http://localhost:8000/healthz"]
interval: 10s
retries: 3
$ docker compose up -d --wait
$ docker compose ps
NAME SERVICE STATUS
shop-db-1 db Up 25 seconds (healthy)
shop-api-1 api Up 8 seconds (healthy)
The health check command runs inside the container, so the tool it uses (curl, pg_isready, wget) must exist in the image. Minimal images often need a tiny health endpoint client or a wget-based check instead.
depends_on alone is "start cooking the main course once the oven is switched on". The health check is "once the oven has actually reached 200 degrees".
Restart policies
| Policy | Restarts when |
|---|---|
no (default) |
Never |
on-failure[:5] |
The process exits non-zero (optionally at most 5 times) |
always |
Always, even after you docker stop it, once the daemon restarts |
unless-stopped |
Always, except after you stopped it yourself. Usually what you want on a server. |
A restart policy brings containers back after a crash and after a host reboot, as long as the Docker service itself is enabled (systemctl enable docker).
Limits and logs
Unlimited containers can take the whole host down. Set limits, and cap log files so they can't fill the disk:
services:
api:
image: shop/api:1.4
restart: unless-stopped
deploy:
resources:
limits:
cpus: "1.0"
memory: 512M
logging:
driver: json-file
options:
max-size: "10m"
max-file: "3"
read_only: true
tmpfs: [ /tmp ]
If the process goes over its memory limit, the kernel kills it (exit code 137, OOMKilled: true in docker inspect). The log settings can also be set once for the whole host in daemon.json (Level 4).
Secrets: keep passwords out of the environment
Environment variables leak: they show in docker inspect, in /proc/<pid>/environ, and in crash reports. Compose can mount a secret as a file instead:
services:
db:
image: postgres:16
environment:
POSTGRES_PASSWORD_FILE: /run/secrets/db_password
secrets: [ db_password ]
api:
image: shop/api:1.4
secrets: [ db_password ] # the app reads /run/secrets/db_password
secrets:
db_password:
file: ./secrets/db_password.txt # chmod 600, git-ignored
Many official images (PostgreSQL, MySQL, MariaDB) accept a *_FILE variant of their password variables for exactly this. On a single host these secrets are plain files, not encrypted; the gain is that the value never appears in the container's configuration or environment.
One stack, several environments
Keep the common setup in compose.yaml and put differences in override files:
# compose.override.yaml: loaded automatically, for local development
services:
api:
build: ./api
volumes:
- ./api:/app # live code reload
environment:
LOG_LEVEL: debug
# compose.prod.yaml: used explicitly on servers
services:
api:
image: registry.lab.local:5000/shop/api:1.4
restart: unless-stopped
$ docker compose up -d # base + compose.override.yaml
$ docker compose -f compose.yaml -f compose.prod.yaml up -d # base + prod only
Lists such as ports are merged (added together) by default, so an override file can't remove a port the base file publishes. Recent Compose v2 releases add the !override and !reset YAML tags for that; check the merged result with docker compose config before you rely on it.
Profiles switch optional services on and off:
services:
pgadmin:
image: dpage/pgadmin4
profiles: [ debug ]
docker compose --profile debug up -d starts it; a plain up -d doesn't.
Try it: break the start order, then fix it
- Take the stack from the previous lesson and remove any retry logic from the API's database connection.
docker compose down -v && docker compose up -d, thendocker compose logs api: connection refused on the first start.- Add the
dbhealth check andcondition: service_healthy, repeat, and watchdocker compose psgo from(health: starting)to(healthy)before the API starts. - Add
mem_limit-style limits (deploy.resources.limits.memory: 64M) to the API, generate load, and find the 137 exit code withdocker inspect -f '{{.State.OOMKilled}}'.
Going deeper: Compose on a server
Compose is a good fit for a single host: a small internal tool, an edge appliance, a lab, a CI environment. Things to add when it runs for real:
- Pin images by tag and digest in
compose.prod.yamlso a re-pull can't change what runs. - Run it from systemd (a unit that does
docker compose up -d/down) or rely onrestart: unless-stoppedplus an enabled Docker service. - Keep the files in Git and deploy with a pipeline, not by editing on the server.
- For several hosts, self-healing and rolling updates across machines, you have outgrown Compose: that's what Kubernetes is for (next track).
Recap
- healthcheck +
condition: service_healthygives a real start order;depends_onalone doesn't wait for readiness. restart: unless-stoppedon servers; set memory/CPU limits and log rotation.- Secrets as files under
/run/secrets/, using*_FILEvariables where images support them. - Override files for environment differences, profiles for optional services. Verify with
docker compose config.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.