Lesson 09 of 18 · Part 2 — Build the platform
Storage: Persistent Disk, Hyperdisk, Filestore & Cloud Storage
Give GKE workloads the right storage: Persistent Disk and Hyperdisk through the managed CSI driver, regional disks that survive a zone failure, Filestore for shared files, Cloud Storage FUSE for objects, StorageClasses, StatefulSets, snapshots, and Backup for GKE.
The building blocks
| Storage | Access | Through | Use for |
|---|---|---|---|
| Persistent Disk (pd-balanced, pd-ssd, pd-standard) | ReadWriteOnce, one zone | Compute Engine PD CSI driver (managed) | Databases, queues, most stateful apps |
| Hyperdisk (Balanced, Throughput, Extreme) | ReadWriteOnce | Same driver, on supported machine series | Tunable IOPS and throughput independent of size |
| Regional Persistent Disk | ReadWriteOnce, replicated across two zones | replication-type: regional-pd |
Stateful apps that must survive a zone failure without app-level replication |
| Filestore | ReadWriteMany (NFS) | Filestore CSI driver | Shared files, legacy apps, content stores |
| Cloud Storage | Object storage, mountable as files | Cloud Storage FUSE CSI driver | ML datasets, read-heavy content; not a database disk |
| Local SSD | Node-local, lost with the node | Ephemeral storage or local PVs | Caches, scratch space |
GKE manages the CSI drivers for you: Persistent Disk is always on; Filestore, Cloud Storage FUSE and Backup for GKE are add-ons you enable.
Storage choices are like places to keep your work. A Persistent Disk is your own desk drawer: fast, but in one room. A regional disk is a drawer with an automatic copy in the room next door. Filestore is the shared filing cabinet the whole team can open. Cloud Storage is the warehouse: huge and cheap, but you fetch boxes, you don't scribble in them.
StorageClasses
GKE ships standard-rwo (pd-balanced) and premium-rwo (pd-ssd), both with WaitForFirstConsumer. Define your own for other needs:
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: regional-balanced
provisioner: pd.csi.storage.gke.io
parameters:
type: pd-balanced
replication-type: regional-pd
volumeBindingMode: WaitForFirstConsumer
allowVolumeExpansion: true
reclaimPolicy: Retain
---
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: shared-files
provisioner: filestore.csi.storage.gke.io
parameters:
tier: standard
network: projects/net-host-prod/global/networks/prod-vpc
volumeBindingMode: WaitForFirstConsumer
allowVolumeExpansion: true
- Retain on production data classes: deleting a PVC by mistake shouldn't delete the disk.
- Encryption: disks are encrypted by default; set a CMEK key (
disk-encryption-kms-keyparameter) where compliance requires customer-managed keys. - Filestore on a Shared VPC needs the host network named, and its own IP range (private services access); plan it with the IP plan.
StatefulSets in a multi-zone cluster
- Each replica gets its own PVC from
volumeClaimTemplates; with zonal disks, each replica is pinned to the zone its disk landed in. - Spread replicas across zones (topology spread constraints) and let the application replicate data (PostgreSQL streaming, Kafka replicas) so losing a zone loses one replica, not the data.
- Use regional disks when the app can't replicate itself and a restart in another zone is acceptable.
- Managed databases (Cloud SQL, AlloyDB, Memorystore, Spanner) are often the better home for important data than a StatefulSet; the cluster then stays easier to rebuild.
Snapshots
The PD CSI driver supports VolumeSnapshots:
apiVersion: snapshot.storage.k8s.io/v1
kind: VolumeSnapshotClass
metadata:
name: pd-snapshots
driver: pd.csi.storage.gke.io
deletionPolicy: Retain
---
apiVersion: snapshot.storage.k8s.io/v1
kind: VolumeSnapshot
metadata:
name: orders-db-2026-10-07
namespace: shop
spec:
volumeSnapshotClassName: pd-snapshots
source:
persistentVolumeClaimName: data-orders-db-0
Restore by creating a PVC with dataSource pointing at the snapshot. Snapshots are crash-consistent; quiesce or use the database's own backup for application consistency.
Backup for GKE
Backup for GKE backs up Kubernetes resources and volume data together, on a schedule, with retention, and restores them into the same or another cluster, including in another region:
- A backup plan per cluster: which namespaces (or all), whether to include volume data and Secrets, schedule (cron), retention, optional encryption key.
- A restore plan: target cluster, which namespaces, how to handle conflicts, and transformation rules (for example changing a StorageClass).
- Test restores regularly (lesson 15): a backup you've never restored is a hope, not a plan.
Try it: zone failure for a stateful app
- Create a regional cluster and the
regional-balancedStorageClass above. - Deploy a single-replica StatefulSet (a small PostgreSQL) with a PVC from that class; write some rows.
- Cordon and drain the node it runs on, and taint all nodes in that zone so the Pod can't return there.
- Watch the Pod start in the other zone with its data intact (
kubectl get pod -o wide, then query the rows). - Enable Backup for GKE, create a backup plan for the namespace, run a backup, delete the namespace, and restore it.
Going deeper: performance and cost
- Persistent Disk performance scales with size and machine type; Hyperdisk lets you provision IOPS and throughput separately. Measure with
fioon the real machine type before sizing. - Watch out for the attach limit per node (it depends on the machine type) when packing many small stateful Pods.
- Unused disks keep costing money: alert on PVs in
Releasedstate and unattached disks.
Recap
- Persistent Disk / Hyperdisk (RWO, zonal), regional PD (two zones), Filestore (RWX), Cloud Storage FUSE (objects), all through managed CSI drivers.
- StorageClasses with WaitForFirstConsumer, Retain for important data, CMEK where required.
- In multi-zone clusters, replicate at the application level or use regional disks; prefer managed databases for critical data.
- VolumeSnapshots for disks, Backup for GKE for resources plus volumes, and test restores.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.