Cheat Sheets / Vocabulary

Engineering vocabulary

The words you hear in design reviews, incident calls and platform discussions, each with a plain meaning and a Kubernetes / cloud example. 90 terms.

Reliability 12

TermSimple meaningExample
Blast radiusHow much is affected when something failsOne node failure takes down 20 pods; one bad config push affects every clusterFind in lessons →
Failure domainA boundary within which a single failure can take things down togetherA node, a rack, an availability zone, a regionFind in lessons →
SPOF (single point of failure)One component whose failure breaks the whole serviceAll replicas scheduled on one node; a single NAT gatewayFind in lessons →
Fault toleranceKeep working while a component is failingReplicas spread across three AZs keep serving when one AZ is downFind in lessons →
ResilienceWithstand failure and recover from itPods recreated on healthy nodes after a node diesFind in lessons →
FailoverSwitch to a standby when the primary failsDatabase primary in AZ-1 fails; replica in AZ-2 takes overFind in lessons →
Graceful degradationThe core keeps working with reduced features when a part failsRecommendations are down, but checkout still worksFind in lessons →
Cascading failureOne failure triggers more failures downstreamSlow database → client retries → CPU exhaustion → more services failFind in lessons →
Noisy neighbourOne workload hurts others sharing the same resourcesA CPU-hungry pod without limits slows every pod on the nodeFind in lessons →
Thundering herdA huge number of clients act at the same momentThousands of pods reconnect to a database the second it comes backFind in lessons →
Retry stormAggressive retries make an outage worseEvery pod retries a failing API immediately and in a tight loopFind in lessons →
Split brainTwo parts of a system both believe they're in charge after losing contactTwo database nodes both accept writes during a network partition; quorum-based systems like etcd prevent thisFind in lessons →

Kubernetes operations 9

TermSimple meaningExample
Pod churnPods are created and destroyed over and overCrash → restart → reschedule, repeatedlyFind in lessons →
Pod evictionKubernetes removes a pod from a nodeThe node is under memory pressure; lower-priority pods are evictedFind in lessons →
Resource starvationA workload can't get the resources it needsA pod needs CPU but the node has none leftFind in lessons →
Resource contentionWorkloads compete for the same resourcesSeveral pods fighting for CPU on one nodeFind in lessons →
Hot spotOne resource gets a disproportionate share of the loadOne node or one partition receives most of the trafficFind in lessons →
Cluster sprawlSo many clusters that they become hard to manageHundreds of clusters on different versions with different add-onsFind in lessons →
Orphaned resourceA resource left behind after its owner is goneA load balancer or EBS volume still billing after the cluster was deletedFind in lessons →
Image pull stormMany nodes pull the same images at the same timeA large deployment starts on 100 new nodes at onceFind in lessons →
Control-plane saturationThe Kubernetes control plane is overloadedA controller floods the API server with list/watch requestsFind in lessons →

Lifecycle 3

TermSimple meaningExample
Day 0Design and initial provisioningPlan and create the VPC, cluster and node groupsFind in lessons →
Day 1Initial configuration and onboardingInstall add-ons, ingress, monitoring; onboard the first appsFind in lessons →
Day 2Ongoing operations for the rest of the system's lifeUpgrades, scaling, patching, incidents, cost reviewsFind in lessons →

Scaling 5

TermSimple meaningExample
Horizontal scaling (scale-out)Add more instances3 pods → 10 pods; 5 nodes → 8 nodesFind in lessons →
Vertical scaling (scale-up)Give an existing instance more resourcesA pod from 2 CPU to 8 CPU; a bigger EC2 instance typeFind in lessons →
Capacity planningEstimating future resource needsHow many nodes and IPs will we need in six months?Find in lessons →
OverprovisioningKeeping more capacity than currently needed30% spare capacity to absorb spikesFind in lessons →
UnderprovisioningNot enough capacity for the loadPods stay Pending; requests time out at peakFind in lessons →

Networking 7

TermSimple meaningExample
North-south trafficTraffic entering or leaving the systemInternet → load balancer → podFind in lessons →
East-west trafficInternal service-to-service trafficorders service → payments serviceFind in lessons →
IngressThe HTTP/HTTPS entry point into a clusterALB → Ingress → Service → podsFind in lessons →
EgressTraffic leaving the clusterPod → NAT gateway → internetFind in lessons →
Network partitionParts of a system can't reach each otherAn AZ loses connectivity to the othersFind in lessons →
Network bottleneckA network component limits throughputInstance bandwidth or ENI limits; one overloaded NAT gatewayFind in lessons →
Traffic hairpinningTraffic takes a detour out and back in to reach something nearbyPod calls another service through the public load balancer instead of the internal ServiceFind in lessons →

Observability 7

TermSimple meaningExample
MonitoringTells you that something is wrongAlert: CPU above 90% for 10 minutesFind in lessons →
ObservabilityHelps you understand why it's wrongMetrics + logs + traces to follow one slow requestFind in lessons →
Golden signalsLatency, traffic, errors and saturationAPI latency and error rate on the main dashboardFind in lessons →
SLI (service level indicator)A measured number describing service quality99.95% of requests succeeded this monthFind in lessons →
SLO (service level objective)The target for an SLI99.9% of requests succeed over 30 daysFind in lessons →
SLA (service level agreement)A contractual promise, usually with penalties99.9% monthly uptime or service creditsFind in lessons →
Error budgetHow much unreliability the SLO allows0.1% of requests may fail this month; spend it on releases or save itFind in lessons →

Incident management 6

TermSimple meaningExample
RCA (root cause analysis)Finding the underlying cause, not just the symptomWhy did the node run out of memory?Find in lessons →
MTTRMean time to recover (restore service)Service restored in 20 minutes on averageFind in lessons →
MTBFMean time between failuresOn average 200 days between failuresFind in lessons →
Five whysAsking 'why?' repeatedly to get past the obvious causeApp errors → pod restarts → OOM → no memory limit → no default LimitRangeFind in lessons →
PostmortemA written review of an incident with corrective actionsTimeline, impact, causes, actions with owners and datesFind in lessons →
Blameless postmortemFocuses on systems and processes, not individualsAdd a guard-rail to the pipeline instead of blaming whoever pressed deployFind in lessons →

Disaster recovery 3

TermSimple meaningExample
RTO (recovery time objective)Maximum acceptable time to restore serviceBack online within 30 minutesFind in lessons →
RPO (recovery point objective)Maximum acceptable data loss, measured in timeLose at most 5 minutes of dataFind in lessons →
Backup and restoreKeeping copies of data and configuration, and proving you can bring them backRestore an EBS snapshot or database backup in a drillFind in lessons →

Security 6

TermSimple meaningExample
Least privilegeGive only the permissions actually neededA node role that can only pull from ECRFind in lessons →
Defence in depthSeveral independent layers of protectionIAM + security groups + RBAC + network policies + image scanningFind in lessons →
Zero trustNever trust based on network location; verify every accessEvery service call is authenticated, even inside the VPCFind in lessons →
Attack surfaceAll the points where an attacker could get inA public Kubernetes API endpoint; an open SSH portFind in lessons →
Secret sprawlSecrets scattered across many placesPasswords in YAML, Git, scripts and CI variablesFind in lessons →
Credential rotationReplacing credentials on a schedule or after exposureRotate database passwords and certificates automaticallyFind in lessons →

Platform engineering 7

TermSimple meaningExample
IaC (infrastructure as code)Infrastructure defined in code, reviewed and versionedTerraform creates the VPC and clusterFind in lessons →
GitOpsGit holds the desired state; an agent keeps reality matching itGit → Argo CD → clusterFind in lessons →
Immutable infrastructureReplace instead of modifying in placeReplace a broken node with a fresh one instead of repairing itFind in lessons →
Configuration driftActual state differs from the intended (coded) stateTerraform says 3 nodes, AWS shows 5 after a console changeFind in lessons →
Paved road (golden path)The standard, supported way for teams to build and shipService template → CI/CD → cluster, with monitoring built inFind in lessons →
ToilRepetitive, manual, automatable operational work that doesn't improve anythingManually checking the health of 500 clusters every morningFind in lessons →
Single pane of glassOne interface showing many systemsOne Grafana dashboard for every clusterFind in lessons →

Delivery 8

TermSimple meaningExample
Continuous integration (CI)Every change is built and tested automatically when it's pushedPull request → build, unit tests, lint, scanFind in lessons →
Continuous delivery / deployment (CD)Delivery: always releasable; deployment: every passing change goes live automaticallyMerged to main → deployed to dev automatically, prod after approvalFind in lessons →
ArtifactThe built, versioned output that gets deployedContainer image orders-api@sha256:…Find in lessons →
PromotionMoving the same artifact from one environment to the nextSame image digest: dev → staging → prodFind in lessons →
Canary releaseSend a small share of traffic to the new version first5% → 25% → 100%, with automatic rollback on errorsFind in lessons →
Blue-green deploymentRun old and new side by side, then switch traffic at onceSwitch the load balancer from blue to green; switch back if neededFind in lessons →
RollbackReturn to the previous known-good versiongit revert the promotion commitFind in lessons →
Shift leftCatch problems earlier in the pipelineSecurity scans and policy checks in the pull request, not after deployFind in lessons →

Architecture 11

TermSimple meaningExample
StatelessDoesn't keep data locally between requestsA web or API pod that can be killed and replaced anytimeFind in lessons →
StatefulNeeds persistent dataPostgreSQL, Kafka, ElasticsearchFind in lessons →
Loose couplingComponents depend on each other as little as possibleServices talk through APIs or queues, deploy independentlyFind in lessons →
Tight couplingComponents can't work or change independentlyTwo services that must always be deployed togetherFind in lessons →
IdempotencyDoing something twice gives the same result as doing it onceterraform apply with no changes; a retried payment isn't charged twiceFind in lessons →
DeclarativeDescribe the desired end statereplicas: 3 in a DeploymentFind in lessons →
ImperativeDescribe the steps to performkubectl scale deploy/api --replicas=3Find in lessons →
Eventual consistencyState converges to correct over time rather than instantlyControllers reconcile until actual matches desiredFind in lessons →
Event-drivenComponents react to eventsA controller watches resources; a function runs when a file lands in S3Find in lessons →
Technical debtThe future cost of past shortcutsA manual deployment step nobody has time to automateFind in lessons →
Vendor lock-inHard or costly to move away from a providerHeavy use of provider-specific managed servicesFind in lessons →

Production design 6

TermSimple meaningExample
Design for failureAssume every component will fail, and plan for itMulti-AZ cluster, PDBs, retries with backoffFind in lessons →
Fail fastDetect invalid conditions early and stopReject a deployment with a bad config at admission, not at 3 a.m.Find in lessons →
Backward compatibilityA new version still works with old clientsAPI v2 still accepts v1 requestsFind in lessons →
Scalability bottleneckThe component that stops the system scaling furtherThe database hits its connection limit before the app doesFind in lessons →
Cost optimisationReducing cost while still meeting requirementsRight-size requests; Spot and consolidation with KarpenterFind in lessons →
Operational overheadThe effort needed to keep something runningManaging 500 clusters by handFind in lessons →