Production GKE Platform — From Zero to Production›14 · Terraform, GitOps & fleets across cloud and on-prem

Lesson 14 of 18 · Part 3 — Operate

Terraform, GitOps & fleets across cloud and on-prem

Run GKE as code and as a fleet: Terraform for networks, clusters and node pools, GitOps with Argo CD or Config Sync for everything inside the clusters, and fleets that manage GKE, on-prem and other-cloud clusters together with Connect gateway and multi-cluster Gateways.

Advanced
Key wordsTerraform GKEgoogle_container_clustergoogle_container_node_poolterraform-google-modulesremote stateGitOpsArgo CDConfig SyncConfig Connectorfleetsfleet membershipConnect gatewaymulti-cluster Gatewaymulti-cluster ServicesGoogle Distributed Cloudattached clustershybridon-prem

Who manages what

Layer Tool Examples
Organisation, folders, projects, Shared VPC, NAT, firewall policies Terraform (landing zone) Lesson 03
Clusters, node pools, IAM, KMS keys, Artifact Registry Terraform (platform stack) This lesson
Inside clusters: namespaces, RBAC, quotas, policies, add-ons, apps GitOps (Argo CD or Config Sync) The Argo CD — Level by Level track

Terraform is the builder who puts up the building and changes its walls when the plans change. GitOps is the caretaker who walks the corridors all day, making sure every room still matches the floor plan, and puts things back when someone moves them.

GKE in Terraform

resource "google_container_cluster" "prod" {
  name     = "prod"
  project  = var.project_id
  location = "europe-west1" # a region: regional cluster

  network    = var.network_self_link # Shared VPC network and subnet
  subnetwork = var.subnet_self_link
  ip_allocation_policy {
    cluster_secondary_range_name  = "pods"
    services_secondary_range_name = "services"
  }

  release_channel { channel = "REGULAR" }
  maintenance_policy {
    recurring_window {
      start_time = "2026-01-06T02:00:00Z"
      end_time   = "2026-01-06T06:00:00Z"
      recurrence = "FREQ=WEEKLY;BYDAY=TU,WE,TH"
    }
  }

  private_cluster_config {
    enable_private_nodes = true
  }
  workload_identity_config { workload_pool = "${var.project_id}.svc.id.goog" }
  datapath_provider = "ADVANCED_DATAPATH" # Dataplane V2
  database_encryption {
    state    = "ENCRYPTED"
    key_name = var.secrets_kms_key
  }

  remove_default_node_pool = true
  initial_node_count       = 1
  deletion_protection      = true
}

resource "google_container_node_pool" "apps" {
  name     = "apps"
  cluster  = google_container_cluster.prod.id
  location = "europe-west1"

  autoscaling {
    total_min_node_count = 3
    total_max_node_count = 30
    location_policy      = "BALANCED"
  }
  management {
    auto_repair  = true
    auto_upgrade = true
  }
  upgrade_settings {
    max_surge       = 1
    max_unavailable = 0
  }
  node_config {
    machine_type    = "n2-standard-8"
    image_type      = "COS_CONTAINERD"
    service_account = var.node_service_account
    oauth_scopes    = ["https://www.googleapis.com/auth/cloud-platform"]
    workload_metadata_config { mode = "GKE_METADATA" }
    shielded_instance_config {
      enable_secure_boot          = true
      enable_integrity_monitoring = true
    }
    labels = { pool = "apps" }
  }
}
  • remove_default_node_pool + separate google_container_node_pool resources: pools can change without touching the cluster.
  • deletion_protection stops an accidental destroy.
  • The cloud-platform scope plus a minimal service account is the recommended pattern: IAM, not scopes, limits access.
  • The terraform-google-modules/kubernetes-engine modules wrap these resources with good defaults (including private and Autopilot variants); pin the module version and read its changelog before upgrading it.
  • Remote state in a Cloud Storage bucket (versioned, in a separate project), applied from CI with Workload Identity Federation, one state per environment and stack. The Manage EKS with Terraform and Terraform & IaC tracks cover layout, state and pipelines; the same patterns apply here.

GitOps inside the clusters

  • Argo CD: runs in a management cluster (or each cluster), pulls from Git, rich UI, ApplicationSets across many clusters. Register GKE clusters with least-privilege credentials, or use Connect gateway for fleet clusters.
  • Config Sync: Google's GitOps agent in each cluster, configured per fleet (all clusters get the same "root sync" plus team "repo syncs"), integrates with Policy Controller. Simpler and agent-per-cluster; less UI.
  • Config Connector manages Google Cloud resources (buckets, Pub/Sub, IAM) as Kubernetes objects, so app teams can declare their cloud dependencies next to their apps through GitOps.

Pick one GitOps tool per platform, and decide what's platform-owned (add-ons, policies, namespaces) versus team-owned (apps).

Fleets: many clusters, one management plane

A fleet groups clusters for shared management:

Fleet feature What it does
Connect gateway kubectl access through Google Cloud IAM to any fleet cluster, including on-prem and other clouds, without VPNs
Config Sync / Policy Controller The same config and policies everywhere
Multi-cluster Gateway One load balancer in front of Services in several clusters (failover, regional routing)
Multi-cluster Services A Service exported from one cluster reachable from the others
Fleet-level settings Default cluster configurations, rollout sequencing, team scopes

Treat fleet members as interchangeable where possible: the same namespace and ServiceAccount names mean the same thing across the fleet (namespace sameness), which is what makes multi-cluster networking and identity work.

Hybrid: on-prem and other clouds

Option Runs where Managed by
Google Distributed Cloud (software only, for bare metal or VMware) Your data centre or edge sites Installed and operated by you, with Google's lifecycle tooling; fleet member
Google Distributed Cloud connected / air-gapped Google-provided hardware at your site; an air-gapped variant for disconnected environments Google-supported appliance
GKE attached clusters EKS, AKS or any conformant cluster Registered to the fleet; you keep operating it

For a platform team running Kubernetes across cloud and on-prem, the pattern is the same everywhere: one Git source of truth, a fleet (or an Argo CD hub) that reaches every cluster, consistent policies, and observability that sees all clusters in one place. Product names and packaging in this area change often; check the current documentation before designing around a specific offering.

Try it: one cluster from code, then a fleet

  1. Write the Terraform above for a lab project (a single zonal cluster is fine), with state in a versioned bucket; run plan and apply.
  2. Change max_surge on the pool and confirm the plan only touches the pool.
  3. Install Argo CD (or enable Config Sync) and sync a repository with a namespace, a RoleBinding and an app.
  4. Register the cluster to a fleet and connect with gcloud container fleet memberships get-credentials.
  5. If you have a kind cluster on your laptop, attach it to the fleet (follow the attached-clusters guide for your version) and see both in the console.

Going deeper: platform as a product

  • Give teams a self-service path: a pull request that adds a namespace, quotas, RBAC and a Gateway route, reviewed and applied by GitOps.
  • Keep cluster blueprints (Terraform modules plus GitOps bootstrap) so a new cluster is a variable file and one apply.
  • Test Terraform and GitOps changes on a non-production fleet first, with policy checks in CI.

Recap

  • Terraform for Google Cloud resources (network, clusters, separate node pools, IAM, KMS); remote state in Cloud Storage; deletion_protection on.
  • GitOps (Argo CD or Config Sync) for everything inside clusters; Config Connector for cloud resources declared by teams.
  • Fleets manage many clusters together: Connect gateway, shared config and policy, multi-cluster Gateways and Services.
  • Hybrid: Google Distributed Cloud on-prem and attached clusters elsewhere join the same fleet.

This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.