Lesson 11 of 18 · Part 3 — Operate
Observability: Cloud Logging, Cloud Monitoring & Managed Prometheus
See what GKE and its workloads are doing: what Cloud Logging and Cloud Monitoring collect by default, Managed Service for Prometheus for your own metrics, control-plane and kube-state metrics, traces with OpenTelemetry, alerts and SLOs, and how to keep the logging bill under control.
What you get by default
| Signal | Collected by default | Where |
|---|---|---|
| Container stdout/stderr | Yes (WORKLOAD logging) |
Cloud Logging, k8s_container resource |
| System component logs (kubelet, runtime, add-ons) | Yes (SYSTEM) |
Cloud Logging |
| Control-plane logs (API server, scheduler, controller manager) | Optional | Cloud Logging |
| Kubernetes audit logs | Admin activity always; data access optional | Cloud Audit Logs |
| System metrics (node, Pod, container CPU/memory) | Yes | Cloud Monitoring |
| Control-plane metrics, kube-state metrics | Optional packages | Cloud Monitoring (PromQL-queryable) |
| Your app's metrics | Through Managed Service for Prometheus | Cloud Monitoring |
GKE's dashboards in the console (cluster, workload, node views) build on these.
Observability is a car dashboard. Google already fitted the engine lights (system metrics and logs). Managed Prometheus lets you add your own gauges for what matters to your app. Alerts are the warning beep, and SLOs tell you whether the journey is still on time.
Your metrics: Managed Service for Prometheus
Apps expose /metrics as usual. With managed collection enabled, you declare what to scrape:
apiVersion: monitoring.googleapis.com/v1
kind: PodMonitoring
metadata:
name: shop
namespace: shop
spec:
selector:
matchLabels: { app: shop }
endpoints:
- port: metrics
interval: 30s
ClusterPodMonitoringdoes the same across namespaces;Rules/ClusterRuleshold recording and alerting rules in Prometheus format.- Query with PromQL in Cloud Monitoring, or point Grafana at it (Google provides a data-source syncer for authentication).
- Data is stored by Google with long retention and global querying across clusters and projects (metrics scopes).
- Self-deployed collection (your own Prometheus that writes to the managed backend) is an option when you need full Prometheus Operator compatibility.
- Cost follows samples ingested: watch high-cardinality labels and scrape intervals.
Traces
Instrument apps with OpenTelemetry and export to Cloud Trace (directly or through an OpenTelemetry Collector in the cluster, using Workload Identity for credentials). Load balancer logs and trace IDs let you follow a request from the edge into the app. The Observability with OpenTelemetry track covers instrumentation in depth.
Alerts and SLOs
- Alert on symptoms users feel: error rate, latency, saturation of the critical path, SLO burn rate.
- Cloud Monitoring supports SLOs on request-based or windows-based indicators, with burn-rate alert policies.
- Platform alerts worth having from day one:
| Alert | Why |
|---|---|
| Pods Pending for more than N minutes | Capacity, quota, IP or scheduling problems |
| Node NotReady, many Pod restarts (CrashLoopBackOff) | Node or app failures |
| PVC usage above 80 % | Disks fill quietly |
| Cloud NAT dropped packets / port exhaustion | Egress failures look like app timeouts |
| Upgrade or security bulletin notifications | Know when GKE is about to change your cluster |
| Certificate expiry (Certificate Manager), load balancer 5xx | Edge problems |
Route alerts to the owning team through notification channels (PagerDuty, Slack, email), with runbook links in the alert documentation.
Keeping logging affordable
Ingestion is usually the biggest observability cost. In order:
- Reduce at the source: app log levels, no per-request debug logs in production.
- Exclusion filters on the
_Defaultsink for noisy, low-value logs (health checks, chatty sidecars). - Sinks to route logs: security logs to a long-retention bucket or BigQuery, the rest with short retention.
- Sample access logs at the load balancer if full logs aren't required.
gcloud logging sinks update _Default \
--add-exclusion=name=healthchecks,filter='resource.type="k8s_container" AND httpRequest.requestUrl:"/healthz"'
Try it: from metric to alert
- Enable managed Prometheus on a lab cluster and deploy an app that exposes
/metrics(any Prometheus client example). - Add a
PodMonitoringresource and query the metric with PromQL in Metrics Explorer. - Create an alerting policy on the 5xx rate and send yourself a test notification.
- Find the noisiest log source with a Logs Analytics or Log Explorer query grouped by container, and add an exclusion filter for it.
- Enable control-plane metrics and chart API server request latency.
Going deeper: one view across many clusters
- Use a metrics scope (a monitoring project that sees many projects) for a fleet-wide view, and keep alert policies as code (Terraform).
- Decide early whether teams use Cloud Monitoring dashboards or Grafana, and keep one source of truth for SLOs.
- Treat telemetry pipelines as production systems: monitor dropped samples, export errors and agent restarts.
Recap
- GKE collects system and workload logs and system metrics by default; control-plane and kube-state metrics are optional packages.
- Managed Service for Prometheus: you declare
PodMonitoring, Google stores and serves PromQL. - Traces with OpenTelemetry to Cloud Trace.
- Alert on symptoms and SLO burn rate, with platform alerts for pending Pods, NAT, disks and the edge.
- Control cost with log levels, exclusions, sinks and retention.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.