Lesson 14 of 18 · Part 3 — Operate
Terraform, GitOps & fleets across cloud and on-prem
Run GKE as code and as a fleet: Terraform for networks, clusters and node pools, GitOps with Argo CD or Config Sync for everything inside the clusters, and fleets that manage GKE, on-prem and other-cloud clusters together with Connect gateway and multi-cluster Gateways.
Who manages what
| Layer | Tool | Examples |
|---|---|---|
| Organisation, folders, projects, Shared VPC, NAT, firewall policies | Terraform (landing zone) | Lesson 03 |
| Clusters, node pools, IAM, KMS keys, Artifact Registry | Terraform (platform stack) | This lesson |
| Inside clusters: namespaces, RBAC, quotas, policies, add-ons, apps | GitOps (Argo CD or Config Sync) | The Argo CD — Level by Level track |
Terraform is the builder who puts up the building and changes its walls when the plans change. GitOps is the caretaker who walks the corridors all day, making sure every room still matches the floor plan, and puts things back when someone moves them.
GKE in Terraform
resource "google_container_cluster" "prod" {
name = "prod"
project = var.project_id
location = "europe-west1" # a region: regional cluster
network = var.network_self_link # Shared VPC network and subnet
subnetwork = var.subnet_self_link
ip_allocation_policy {
cluster_secondary_range_name = "pods"
services_secondary_range_name = "services"
}
release_channel { channel = "REGULAR" }
maintenance_policy {
recurring_window {
start_time = "2026-01-06T02:00:00Z"
end_time = "2026-01-06T06:00:00Z"
recurrence = "FREQ=WEEKLY;BYDAY=TU,WE,TH"
}
}
private_cluster_config {
enable_private_nodes = true
}
workload_identity_config { workload_pool = "${var.project_id}.svc.id.goog" }
datapath_provider = "ADVANCED_DATAPATH" # Dataplane V2
database_encryption {
state = "ENCRYPTED"
key_name = var.secrets_kms_key
}
remove_default_node_pool = true
initial_node_count = 1
deletion_protection = true
}
resource "google_container_node_pool" "apps" {
name = "apps"
cluster = google_container_cluster.prod.id
location = "europe-west1"
autoscaling {
total_min_node_count = 3
total_max_node_count = 30
location_policy = "BALANCED"
}
management {
auto_repair = true
auto_upgrade = true
}
upgrade_settings {
max_surge = 1
max_unavailable = 0
}
node_config {
machine_type = "n2-standard-8"
image_type = "COS_CONTAINERD"
service_account = var.node_service_account
oauth_scopes = ["https://www.googleapis.com/auth/cloud-platform"]
workload_metadata_config { mode = "GKE_METADATA" }
shielded_instance_config {
enable_secure_boot = true
enable_integrity_monitoring = true
}
labels = { pool = "apps" }
}
}
remove_default_node_pool+ separategoogle_container_node_poolresources: pools can change without touching the cluster.deletion_protectionstops an accidentaldestroy.- The
cloud-platformscope plus a minimal service account is the recommended pattern: IAM, not scopes, limits access. - The terraform-google-modules/kubernetes-engine modules wrap these resources with good defaults (including private and Autopilot variants); pin the module version and read its changelog before upgrading it.
- Remote state in a Cloud Storage bucket (versioned, in a separate project), applied from CI with Workload Identity Federation, one state per environment and stack. The Manage EKS with Terraform and Terraform & IaC tracks cover layout, state and pipelines; the same patterns apply here.
GitOps inside the clusters
- Argo CD: runs in a management cluster (or each cluster), pulls from Git, rich UI, ApplicationSets across many clusters. Register GKE clusters with least-privilege credentials, or use Connect gateway for fleet clusters.
- Config Sync: Google's GitOps agent in each cluster, configured per fleet (all clusters get the same "root sync" plus team "repo syncs"), integrates with Policy Controller. Simpler and agent-per-cluster; less UI.
- Config Connector manages Google Cloud resources (buckets, Pub/Sub, IAM) as Kubernetes objects, so app teams can declare their cloud dependencies next to their apps through GitOps.
Pick one GitOps tool per platform, and decide what's platform-owned (add-ons, policies, namespaces) versus team-owned (apps).
Fleets: many clusters, one management plane
A fleet groups clusters for shared management:
| Fleet feature | What it does |
|---|---|
| Connect gateway | kubectl access through Google Cloud IAM to any fleet cluster, including on-prem and other clouds, without VPNs |
| Config Sync / Policy Controller | The same config and policies everywhere |
| Multi-cluster Gateway | One load balancer in front of Services in several clusters (failover, regional routing) |
| Multi-cluster Services | A Service exported from one cluster reachable from the others |
| Fleet-level settings | Default cluster configurations, rollout sequencing, team scopes |
Treat fleet members as interchangeable where possible: the same namespace and ServiceAccount names mean the same thing across the fleet (namespace sameness), which is what makes multi-cluster networking and identity work.
Hybrid: on-prem and other clouds
| Option | Runs where | Managed by |
|---|---|---|
| Google Distributed Cloud (software only, for bare metal or VMware) | Your data centre or edge sites | Installed and operated by you, with Google's lifecycle tooling; fleet member |
| Google Distributed Cloud connected / air-gapped | Google-provided hardware at your site; an air-gapped variant for disconnected environments | Google-supported appliance |
| GKE attached clusters | EKS, AKS or any conformant cluster | Registered to the fleet; you keep operating it |
For a platform team running Kubernetes across cloud and on-prem, the pattern is the same everywhere: one Git source of truth, a fleet (or an Argo CD hub) that reaches every cluster, consistent policies, and observability that sees all clusters in one place. Product names and packaging in this area change often; check the current documentation before designing around a specific offering.
Try it: one cluster from code, then a fleet
- Write the Terraform above for a lab project (a single zonal cluster is fine), with state in a versioned bucket; run plan and apply.
- Change
max_surgeon the pool and confirm the plan only touches the pool. - Install Argo CD (or enable Config Sync) and sync a repository with a namespace, a RoleBinding and an app.
- Register the cluster to a fleet and connect with
gcloud container fleet memberships get-credentials. - If you have a kind cluster on your laptop, attach it to the fleet (follow the attached-clusters guide for your version) and see both in the console.
Going deeper: platform as a product
- Give teams a self-service path: a pull request that adds a namespace, quotas, RBAC and a Gateway route, reviewed and applied by GitOps.
- Keep cluster blueprints (Terraform modules plus GitOps bootstrap) so a new cluster is a variable file and one apply.
- Test Terraform and GitOps changes on a non-production fleet first, with policy checks in CI.
Recap
- Terraform for Google Cloud resources (network, clusters, separate node pools, IAM, KMS); remote state in Cloud Storage;
deletion_protectionon. - GitOps (Argo CD or Config Sync) for everything inside clusters; Config Connector for cloud resources declared by teams.
- Fleets manage many clusters together: Connect gateway, shared config and policy, multi-cluster Gateways and Services.
- Hybrid: Google Distributed Cloud on-prem and attached clusters elsewhere join the same fleet.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.