Production EKS Platform — From Zero to Production›11 · Observability: Prometheus, Grafana & CloudWatch

Lesson 11 of 18 · Part 3 — Operate

Observability: Prometheus, Grafana & CloudWatch

See what the platform is doing: EKS control-plane logs, metrics with Prometheus or CloudWatch Container Insights, logs with Fluent Bit, traces with OpenTelemetry, and alerts that page on symptoms, with the trade-offs between self-run and managed options.

Intermediate → Advanced
Key wordsobservabilitycontrol plane logsCloudWatchContainer InsightsPrometheuskube-prometheus-stackAmazon Managed Service for PrometheusAmazon Managed GrafanaFluent BitOpenTelemetryADOTalertingSLO
apps SDK / auto-instrumented nodes & logs hostmetrics, filelog Prometheus targets scrape OpenTelemetry Collector receivers otlp prometheus filelog processors memory_limiter k8sattributes batch exporters otlp remote write otlphttp Traces Tempo / Jaeger Metrics Prometheus / Mimir Logs Loki / Elasticsearch
Metrics, logs and traces leave the cluster through collectors; you choose self-run or managed back ends.

What to observe, layer by layer

Layer Signals Where they come from
Control plane API server latency and errors, audit trail, authentication EKS control-plane logs (CloudWatch Logs), API server /metrics
Nodes CPU, memory, disk, network, kubelet health node-exporter, CloudWatch agent
Kubernetes objects Pod restarts, pending pods, deployment status, HPA state kube-state-metrics, Container Insights
Applications Request rate, errors, latency (RED), business metrics, traces App instrumentation (OpenTelemetry), service mesh
AWS resources ALB 5xx and latency, NAT bytes, EBS throughput CloudWatch metrics

Running a platform without observability is driving at night without dashboard lights. The car works, but you'll only discover you're out of fuel, overheating or speeding when something stops.

Control-plane logs: turn them on

EKS can send five log types to CloudWatch Logs: api, audit, authenticator, controllerManager, scheduler. They're off by default. In production enable at least audit and authenticator; set a retention period on the log group, because audit logs are large.

Metrics: two good paths

Prometheus + Grafana (self-run or managed) CloudWatch Container Insights
Setup kube-prometheus-stack via Helm, or collectors writing to Amazon Managed Service for Prometheus, dashboards in Amazon Managed Grafana amazon-cloudwatch-observability managed add-on
Strengths PromQL, huge ecosystem of exporters and dashboards, portable across clouds and on-prem Zero servers, native AWS console, integrates with CloudWatch alarms
Costs Your compute and storage (self-run) or per-sample pricing (managed) Per metric, per GB of logs
Choose when You already run Prometheus, need custom metrics at scale, or multi-cloud Small teams, AWS-only, want minimal operations

Many platforms combine them: Prometheus (managed or self-run) for workload and Kubernetes metrics, CloudWatch for AWS-resource metrics and control-plane logs, one Grafana over both. Long-term Prometheus storage at scale is covered in the "Scaling Prometheus to Production" track.

Logs

Run Fluent Bit as a DaemonSet (the CloudWatch observability add-on includes it) to ship container logs. Decide per destination:

  • CloudWatch Logs: simple, but priced per GB ingested; set retention and filter out debug noise.
  • OpenSearch / an EFK stack: full-text search at volume (see "Centralized Logging with EFK").
  • S3: cheapest for bulk and compliance archives; query with Athena when needed.

Log in JSON with a request ID so logs, metrics and traces can be joined.

Traces

Instrument services with OpenTelemetry SDKs and run a collector (the AWS Distro for OpenTelemetry, ADOT, is available as an EKS add-on). Send traces to AWS X-Ray, Tempo, Jaeger or a vendor; the OTel data model keeps you portable. See "Observability with OpenTelemetry".

Alerts that deserve to page someone

Page on Keep on dashboards
SLO burn rate (errors, latency) for user-facing services CPU and memory per node
Pods stuck Pending for more than N minutes (capacity or quota) Individual pod restarts
Node NotReady in more than one AZ Single-node blips
Certificate expiry within 14 days Disk usage under 80%
ALB 5xx rate, target group with zero healthy targets Request volume
Control-plane API error rate Scheduler latency

Every alert needs a runbook link and an owner. An alert nobody acts on is noise, and noise trains people to ignore the real page.

Try it: first signals (sandbox)

  1. Enable audit and authenticator logs on a lab cluster; run a few kubectl commands and find them in CloudWatch Logs Insights with filter @logStream like /authenticator/.
  2. Install kube-prometheus-stack; open Grafana and find the "Kubernetes / Compute Resources / Namespace (Pods)" dashboard.
  3. Deploy a pod that crash-loops and write a PromQL alert on increase(kube_pod_container_status_restarts_total[10m]) > 3.
  4. Set a 7-day retention on the log group, then delete the lab resources.

Going deeper: cost-aware observability

Observability can become one of the largest line items on a platform. Control it deliberately: drop high-cardinality labels (pod UID, request IDs) from metrics, sample traces (tail-based sampling keeps errors and slow requests), set log retention per namespace class, route debug logs to cheap storage, and review the top ten metric and log producers monthly.

Recap

  • Enable audit + authenticator control-plane logs with retention.
  • Metrics via Prometheus/Grafana (self-run or managed) and/or Container Insights; logs via Fluent Bit; traces via OpenTelemetry.
  • Page on symptoms and SLOs, keep causes on dashboards, attach runbooks.
  • Watch observability cost: cardinality, log volume, retention.

This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.