Production GKE Platform — From Zero to Production›09 · Storage: Persistent Disk, Hyperdisk, Filestore & Cloud Storage

Lesson 09 of 18 · Part 2 — Build the platform

Storage: Persistent Disk, Hyperdisk, Filestore & Cloud Storage

Give GKE workloads the right storage: Persistent Disk and Hyperdisk through the managed CSI driver, regional disks that survive a zone failure, Filestore for shared files, Cloud Storage FUSE for objects, StorageClasses, StatefulSets, snapshots, and Backup for GKE.

Practitioner → Advanced
Key wordsGKE storagePersistent Disk CSIpd-balancedpd-ssdHyperdiskregional persistent diskFilestore CSIReadWriteManyCloud Storage FUSEStorageClassWaitForFirstConsumerStatefulSetVolumeSnapshotBackup for GKECMEK

The building blocks

Storage Access Through Use for
Persistent Disk (pd-balanced, pd-ssd, pd-standard) ReadWriteOnce, one zone Compute Engine PD CSI driver (managed) Databases, queues, most stateful apps
Hyperdisk (Balanced, Throughput, Extreme) ReadWriteOnce Same driver, on supported machine series Tunable IOPS and throughput independent of size
Regional Persistent Disk ReadWriteOnce, replicated across two zones replication-type: regional-pd Stateful apps that must survive a zone failure without app-level replication
Filestore ReadWriteMany (NFS) Filestore CSI driver Shared files, legacy apps, content stores
Cloud Storage Object storage, mountable as files Cloud Storage FUSE CSI driver ML datasets, read-heavy content; not a database disk
Local SSD Node-local, lost with the node Ephemeral storage or local PVs Caches, scratch space

GKE manages the CSI drivers for you: Persistent Disk is always on; Filestore, Cloud Storage FUSE and Backup for GKE are add-ons you enable.

Storage choices are like places to keep your work. A Persistent Disk is your own desk drawer: fast, but in one room. A regional disk is a drawer with an automatic copy in the room next door. Filestore is the shared filing cabinet the whole team can open. Cloud Storage is the warehouse: huge and cheap, but you fetch boxes, you don't scribble in them.

StorageClasses

GKE ships standard-rwo (pd-balanced) and premium-rwo (pd-ssd), both with WaitForFirstConsumer. Define your own for other needs:

apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: regional-balanced
provisioner: pd.csi.storage.gke.io
parameters:
  type: pd-balanced
  replication-type: regional-pd
volumeBindingMode: WaitForFirstConsumer
allowVolumeExpansion: true
reclaimPolicy: Retain
---
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: shared-files
provisioner: filestore.csi.storage.gke.io
parameters:
  tier: standard
  network: projects/net-host-prod/global/networks/prod-vpc
volumeBindingMode: WaitForFirstConsumer
allowVolumeExpansion: true
  • Retain on production data classes: deleting a PVC by mistake shouldn't delete the disk.
  • Encryption: disks are encrypted by default; set a CMEK key (disk-encryption-kms-key parameter) where compliance requires customer-managed keys.
  • Filestore on a Shared VPC needs the host network named, and its own IP range (private services access); plan it with the IP plan.

StatefulSets in a multi-zone cluster

  • Each replica gets its own PVC from volumeClaimTemplates; with zonal disks, each replica is pinned to the zone its disk landed in.
  • Spread replicas across zones (topology spread constraints) and let the application replicate data (PostgreSQL streaming, Kafka replicas) so losing a zone loses one replica, not the data.
  • Use regional disks when the app can't replicate itself and a restart in another zone is acceptable.
  • Managed databases (Cloud SQL, AlloyDB, Memorystore, Spanner) are often the better home for important data than a StatefulSet; the cluster then stays easier to rebuild.

Snapshots

The PD CSI driver supports VolumeSnapshots:

apiVersion: snapshot.storage.k8s.io/v1
kind: VolumeSnapshotClass
metadata:
  name: pd-snapshots
driver: pd.csi.storage.gke.io
deletionPolicy: Retain
---
apiVersion: snapshot.storage.k8s.io/v1
kind: VolumeSnapshot
metadata:
  name: orders-db-2026-10-07
  namespace: shop
spec:
  volumeSnapshotClassName: pd-snapshots
  source:
    persistentVolumeClaimName: data-orders-db-0

Restore by creating a PVC with dataSource pointing at the snapshot. Snapshots are crash-consistent; quiesce or use the database's own backup for application consistency.

Backup for GKE

Backup for GKE backs up Kubernetes resources and volume data together, on a schedule, with retention, and restores them into the same or another cluster, including in another region:

  • A backup plan per cluster: which namespaces (or all), whether to include volume data and Secrets, schedule (cron), retention, optional encryption key.
  • A restore plan: target cluster, which namespaces, how to handle conflicts, and transformation rules (for example changing a StorageClass).
  • Test restores regularly (lesson 15): a backup you've never restored is a hope, not a plan.

Try it: zone failure for a stateful app

  1. Create a regional cluster and the regional-balanced StorageClass above.
  2. Deploy a single-replica StatefulSet (a small PostgreSQL) with a PVC from that class; write some rows.
  3. Cordon and drain the node it runs on, and taint all nodes in that zone so the Pod can't return there.
  4. Watch the Pod start in the other zone with its data intact (kubectl get pod -o wide, then query the rows).
  5. Enable Backup for GKE, create a backup plan for the namespace, run a backup, delete the namespace, and restore it.

Going deeper: performance and cost

  • Persistent Disk performance scales with size and machine type; Hyperdisk lets you provision IOPS and throughput separately. Measure with fio on the real machine type before sizing.
  • Watch out for the attach limit per node (it depends on the machine type) when packing many small stateful Pods.
  • Unused disks keep costing money: alert on PVs in Released state and unattached disks.

Recap

  • Persistent Disk / Hyperdisk (RWO, zonal), regional PD (two zones), Filestore (RWX), Cloud Storage FUSE (objects), all through managed CSI drivers.
  • StorageClasses with WaitForFirstConsumer, Retain for important data, CMEK where required.
  • In multi-zone clusters, replicate at the application level or use regional disks; prefer managed databases for critical data.
  • VolumeSnapshots for disks, Backup for GKE for resources plus volumes, and test restores.

This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.