Lesson 11 of 18 · Part 3 — Operate
Observability: Prometheus, Grafana & CloudWatch
See what the platform is doing: EKS control-plane logs, metrics with Prometheus or CloudWatch Container Insights, logs with Fluent Bit, traces with OpenTelemetry, and alerts that page on symptoms, with the trade-offs between self-run and managed options.
What to observe, layer by layer
| Layer | Signals | Where they come from |
|---|---|---|
| Control plane | API server latency and errors, audit trail, authentication | EKS control-plane logs (CloudWatch Logs), API server /metrics |
| Nodes | CPU, memory, disk, network, kubelet health | node-exporter, CloudWatch agent |
| Kubernetes objects | Pod restarts, pending pods, deployment status, HPA state | kube-state-metrics, Container Insights |
| Applications | Request rate, errors, latency (RED), business metrics, traces | App instrumentation (OpenTelemetry), service mesh |
| AWS resources | ALB 5xx and latency, NAT bytes, EBS throughput | CloudWatch metrics |
Running a platform without observability is driving at night without dashboard lights. The car works, but you'll only discover you're out of fuel, overheating or speeding when something stops.
Control-plane logs: turn them on
EKS can send five log types to CloudWatch Logs: api, audit, authenticator, controllerManager, scheduler. They're off by default. In production enable at least audit and authenticator; set a retention period on the log group, because audit logs are large.
Metrics: two good paths
| Prometheus + Grafana (self-run or managed) | CloudWatch Container Insights | |
|---|---|---|
| Setup | kube-prometheus-stack via Helm, or collectors writing to Amazon Managed Service for Prometheus, dashboards in Amazon Managed Grafana | amazon-cloudwatch-observability managed add-on |
| Strengths | PromQL, huge ecosystem of exporters and dashboards, portable across clouds and on-prem | Zero servers, native AWS console, integrates with CloudWatch alarms |
| Costs | Your compute and storage (self-run) or per-sample pricing (managed) | Per metric, per GB of logs |
| Choose when | You already run Prometheus, need custom metrics at scale, or multi-cloud | Small teams, AWS-only, want minimal operations |
Many platforms combine them: Prometheus (managed or self-run) for workload and Kubernetes metrics, CloudWatch for AWS-resource metrics and control-plane logs, one Grafana over both. Long-term Prometheus storage at scale is covered in the "Scaling Prometheus to Production" track.
Logs
Run Fluent Bit as a DaemonSet (the CloudWatch observability add-on includes it) to ship container logs. Decide per destination:
- CloudWatch Logs: simple, but priced per GB ingested; set retention and filter out debug noise.
- OpenSearch / an EFK stack: full-text search at volume (see "Centralized Logging with EFK").
- S3: cheapest for bulk and compliance archives; query with Athena when needed.
Log in JSON with a request ID so logs, metrics and traces can be joined.
Traces
Instrument services with OpenTelemetry SDKs and run a collector (the AWS Distro for OpenTelemetry, ADOT, is available as an EKS add-on). Send traces to AWS X-Ray, Tempo, Jaeger or a vendor; the OTel data model keeps you portable. See "Observability with OpenTelemetry".
Alerts that deserve to page someone
| Page on | Keep on dashboards |
|---|---|
| SLO burn rate (errors, latency) for user-facing services | CPU and memory per node |
Pods stuck Pending for more than N minutes (capacity or quota) |
Individual pod restarts |
Node NotReady in more than one AZ |
Single-node blips |
| Certificate expiry within 14 days | Disk usage under 80% |
| ALB 5xx rate, target group with zero healthy targets | Request volume |
| Control-plane API error rate | Scheduler latency |
Every alert needs a runbook link and an owner. An alert nobody acts on is noise, and noise trains people to ignore the real page.
Try it: first signals (sandbox)
- Enable
auditandauthenticatorlogs on a lab cluster; run a fewkubectlcommands and find them in CloudWatch Logs Insights withfilter @logStream like /authenticator/. - Install kube-prometheus-stack; open Grafana and find the "Kubernetes / Compute Resources / Namespace (Pods)" dashboard.
- Deploy a pod that crash-loops and write a PromQL alert on
increase(kube_pod_container_status_restarts_total[10m]) > 3. - Set a 7-day retention on the log group, then delete the lab resources.
Going deeper: cost-aware observability
Observability can become one of the largest line items on a platform. Control it deliberately: drop high-cardinality labels (pod UID, request IDs) from metrics, sample traces (tail-based sampling keeps errors and slow requests), set log retention per namespace class, route debug logs to cheap storage, and review the top ten metric and log producers monthly.
Recap
- Enable audit + authenticator control-plane logs with retention.
- Metrics via Prometheus/Grafana (self-run or managed) and/or Container Insights; logs via Fluent Bit; traces via OpenTelemetry.
- Page on symptoms and SLOs, keep causes on dashboards, attach runbooks.
- Watch observability cost: cardinality, log volume, retention.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.