Cheat Sheets / Vocabulary
Engineering vocabulary
The words you hear in design reviews, incident calls and platform discussions, each with a plain meaning and a Kubernetes / cloud example. 90 terms.
Reliability 12
| Term | Simple meaning | Example | |
|---|---|---|---|
| Blast radius | How much is affected when something fails | One node failure takes down 20 pods; one bad config push affects every cluster | Find in lessons → |
| Failure domain | A boundary within which a single failure can take things down together | A node, a rack, an availability zone, a region | Find in lessons → |
| SPOF (single point of failure) | One component whose failure breaks the whole service | All replicas scheduled on one node; a single NAT gateway | Find in lessons → |
| Fault tolerance | Keep working while a component is failing | Replicas spread across three AZs keep serving when one AZ is down | Find in lessons → |
| Resilience | Withstand failure and recover from it | Pods recreated on healthy nodes after a node dies | Find in lessons → |
| Failover | Switch to a standby when the primary fails | Database primary in AZ-1 fails; replica in AZ-2 takes over | Find in lessons → |
| Graceful degradation | The core keeps working with reduced features when a part fails | Recommendations are down, but checkout still works | Find in lessons → |
| Cascading failure | One failure triggers more failures downstream | Slow database → client retries → CPU exhaustion → more services fail | Find in lessons → |
| Noisy neighbour | One workload hurts others sharing the same resources | A CPU-hungry pod without limits slows every pod on the node | Find in lessons → |
| Thundering herd | A huge number of clients act at the same moment | Thousands of pods reconnect to a database the second it comes back | Find in lessons → |
| Retry storm | Aggressive retries make an outage worse | Every pod retries a failing API immediately and in a tight loop | Find in lessons → |
| Split brain | Two parts of a system both believe they're in charge after losing contact | Two database nodes both accept writes during a network partition; quorum-based systems like etcd prevent this | Find in lessons → |
Kubernetes operations 9
| Term | Simple meaning | Example | |
|---|---|---|---|
| Pod churn | Pods are created and destroyed over and over | Crash → restart → reschedule, repeatedly | Find in lessons → |
| Pod eviction | Kubernetes removes a pod from a node | The node is under memory pressure; lower-priority pods are evicted | Find in lessons → |
| Resource starvation | A workload can't get the resources it needs | A pod needs CPU but the node has none left | Find in lessons → |
| Resource contention | Workloads compete for the same resources | Several pods fighting for CPU on one node | Find in lessons → |
| Hot spot | One resource gets a disproportionate share of the load | One node or one partition receives most of the traffic | Find in lessons → |
| Cluster sprawl | So many clusters that they become hard to manage | Hundreds of clusters on different versions with different add-ons | Find in lessons → |
| Orphaned resource | A resource left behind after its owner is gone | A load balancer or EBS volume still billing after the cluster was deleted | Find in lessons → |
| Image pull storm | Many nodes pull the same images at the same time | A large deployment starts on 100 new nodes at once | Find in lessons → |
| Control-plane saturation | The Kubernetes control plane is overloaded | A controller floods the API server with list/watch requests | Find in lessons → |
Lifecycle 3
| Term | Simple meaning | Example | |
|---|---|---|---|
| Day 0 | Design and initial provisioning | Plan and create the VPC, cluster and node groups | Find in lessons → |
| Day 1 | Initial configuration and onboarding | Install add-ons, ingress, monitoring; onboard the first apps | Find in lessons → |
| Day 2 | Ongoing operations for the rest of the system's life | Upgrades, scaling, patching, incidents, cost reviews | Find in lessons → |
Scaling 5
| Term | Simple meaning | Example | |
|---|---|---|---|
| Horizontal scaling (scale-out) | Add more instances | 3 pods → 10 pods; 5 nodes → 8 nodes | Find in lessons → |
| Vertical scaling (scale-up) | Give an existing instance more resources | A pod from 2 CPU to 8 CPU; a bigger EC2 instance type | Find in lessons → |
| Capacity planning | Estimating future resource needs | How many nodes and IPs will we need in six months? | Find in lessons → |
| Overprovisioning | Keeping more capacity than currently needed | 30% spare capacity to absorb spikes | Find in lessons → |
| Underprovisioning | Not enough capacity for the load | Pods stay Pending; requests time out at peak | Find in lessons → |
Networking 7
| Term | Simple meaning | Example | |
|---|---|---|---|
| North-south traffic | Traffic entering or leaving the system | Internet → load balancer → pod | Find in lessons → |
| East-west traffic | Internal service-to-service traffic | orders service → payments service | Find in lessons → |
| Ingress | The HTTP/HTTPS entry point into a cluster | ALB → Ingress → Service → pods | Find in lessons → |
| Egress | Traffic leaving the cluster | Pod → NAT gateway → internet | Find in lessons → |
| Network partition | Parts of a system can't reach each other | An AZ loses connectivity to the others | Find in lessons → |
| Network bottleneck | A network component limits throughput | Instance bandwidth or ENI limits; one overloaded NAT gateway | Find in lessons → |
| Traffic hairpinning | Traffic takes a detour out and back in to reach something nearby | Pod calls another service through the public load balancer instead of the internal Service | Find in lessons → |
Observability 7
| Term | Simple meaning | Example | |
|---|---|---|---|
| Monitoring | Tells you that something is wrong | Alert: CPU above 90% for 10 minutes | Find in lessons → |
| Observability | Helps you understand why it's wrong | Metrics + logs + traces to follow one slow request | Find in lessons → |
| Golden signals | Latency, traffic, errors and saturation | API latency and error rate on the main dashboard | Find in lessons → |
| SLI (service level indicator) | A measured number describing service quality | 99.95% of requests succeeded this month | Find in lessons → |
| SLO (service level objective) | The target for an SLI | 99.9% of requests succeed over 30 days | Find in lessons → |
| SLA (service level agreement) | A contractual promise, usually with penalties | 99.9% monthly uptime or service credits | Find in lessons → |
| Error budget | How much unreliability the SLO allows | 0.1% of requests may fail this month; spend it on releases or save it | Find in lessons → |
Incident management 6
| Term | Simple meaning | Example | |
|---|---|---|---|
| RCA (root cause analysis) | Finding the underlying cause, not just the symptom | Why did the node run out of memory? | Find in lessons → |
| MTTR | Mean time to recover (restore service) | Service restored in 20 minutes on average | Find in lessons → |
| MTBF | Mean time between failures | On average 200 days between failures | Find in lessons → |
| Five whys | Asking 'why?' repeatedly to get past the obvious cause | App errors → pod restarts → OOM → no memory limit → no default LimitRange | Find in lessons → |
| Postmortem | A written review of an incident with corrective actions | Timeline, impact, causes, actions with owners and dates | Find in lessons → |
| Blameless postmortem | Focuses on systems and processes, not individuals | Add a guard-rail to the pipeline instead of blaming whoever pressed deploy | Find in lessons → |
Disaster recovery 3
| Term | Simple meaning | Example | |
|---|---|---|---|
| RTO (recovery time objective) | Maximum acceptable time to restore service | Back online within 30 minutes | Find in lessons → |
| RPO (recovery point objective) | Maximum acceptable data loss, measured in time | Lose at most 5 minutes of data | Find in lessons → |
| Backup and restore | Keeping copies of data and configuration, and proving you can bring them back | Restore an EBS snapshot or database backup in a drill | Find in lessons → |
Security 6
| Term | Simple meaning | Example | |
|---|---|---|---|
| Least privilege | Give only the permissions actually needed | A node role that can only pull from ECR | Find in lessons → |
| Defence in depth | Several independent layers of protection | IAM + security groups + RBAC + network policies + image scanning | Find in lessons → |
| Zero trust | Never trust based on network location; verify every access | Every service call is authenticated, even inside the VPC | Find in lessons → |
| Attack surface | All the points where an attacker could get in | A public Kubernetes API endpoint; an open SSH port | Find in lessons → |
| Secret sprawl | Secrets scattered across many places | Passwords in YAML, Git, scripts and CI variables | Find in lessons → |
| Credential rotation | Replacing credentials on a schedule or after exposure | Rotate database passwords and certificates automatically | Find in lessons → |
Platform engineering 7
| Term | Simple meaning | Example | |
|---|---|---|---|
| IaC (infrastructure as code) | Infrastructure defined in code, reviewed and versioned | Terraform creates the VPC and cluster | Find in lessons → |
| GitOps | Git holds the desired state; an agent keeps reality matching it | Git → Argo CD → cluster | Find in lessons → |
| Immutable infrastructure | Replace instead of modifying in place | Replace a broken node with a fresh one instead of repairing it | Find in lessons → |
| Configuration drift | Actual state differs from the intended (coded) state | Terraform says 3 nodes, AWS shows 5 after a console change | Find in lessons → |
| Paved road (golden path) | The standard, supported way for teams to build and ship | Service template → CI/CD → cluster, with monitoring built in | Find in lessons → |
| Toil | Repetitive, manual, automatable operational work that doesn't improve anything | Manually checking the health of 500 clusters every morning | Find in lessons → |
| Single pane of glass | One interface showing many systems | One Grafana dashboard for every cluster | Find in lessons → |
Delivery 8
| Term | Simple meaning | Example | |
|---|---|---|---|
| Continuous integration (CI) | Every change is built and tested automatically when it's pushed | Pull request → build, unit tests, lint, scan | Find in lessons → |
| Continuous delivery / deployment (CD) | Delivery: always releasable; deployment: every passing change goes live automatically | Merged to main → deployed to dev automatically, prod after approval | Find in lessons → |
| Artifact | The built, versioned output that gets deployed | Container image orders-api@sha256:… | Find in lessons → |
| Promotion | Moving the same artifact from one environment to the next | Same image digest: dev → staging → prod | Find in lessons → |
| Canary release | Send a small share of traffic to the new version first | 5% → 25% → 100%, with automatic rollback on errors | Find in lessons → |
| Blue-green deployment | Run old and new side by side, then switch traffic at once | Switch the load balancer from blue to green; switch back if needed | Find in lessons → |
| Rollback | Return to the previous known-good version | git revert the promotion commit | Find in lessons → |
| Shift left | Catch problems earlier in the pipeline | Security scans and policy checks in the pull request, not after deploy | Find in lessons → |
Architecture 11
| Term | Simple meaning | Example | |
|---|---|---|---|
| Stateless | Doesn't keep data locally between requests | A web or API pod that can be killed and replaced anytime | Find in lessons → |
| Stateful | Needs persistent data | PostgreSQL, Kafka, Elasticsearch | Find in lessons → |
| Loose coupling | Components depend on each other as little as possible | Services talk through APIs or queues, deploy independently | Find in lessons → |
| Tight coupling | Components can't work or change independently | Two services that must always be deployed together | Find in lessons → |
| Idempotency | Doing something twice gives the same result as doing it once | terraform apply with no changes; a retried payment isn't charged twice | Find in lessons → |
| Declarative | Describe the desired end state | replicas: 3 in a Deployment | Find in lessons → |
| Imperative | Describe the steps to perform | kubectl scale deploy/api --replicas=3 | Find in lessons → |
| Eventual consistency | State converges to correct over time rather than instantly | Controllers reconcile until actual matches desired | Find in lessons → |
| Event-driven | Components react to events | A controller watches resources; a function runs when a file lands in S3 | Find in lessons → |
| Technical debt | The future cost of past shortcuts | A manual deployment step nobody has time to automate | Find in lessons → |
| Vendor lock-in | Hard or costly to move away from a provider | Heavy use of provider-specific managed services | Find in lessons → |
Production design 6
| Term | Simple meaning | Example | |
|---|---|---|---|
| Design for failure | Assume every component will fail, and plan for it | Multi-AZ cluster, PDBs, retries with backoff | Find in lessons → |
| Fail fast | Detect invalid conditions early and stop | Reject a deployment with a bad config at admission, not at 3 a.m. | Find in lessons → |
| Backward compatibility | A new version still works with old clients | API v2 still accepts v1 requests | Find in lessons → |
| Scalability bottleneck | The component that stops the system scaling further | The database hits its connection limit before the app does | Find in lessons → |
| Cost optimisation | Reducing cost while still meeting requirements | Right-size requests; Spot and consolidation with Karpenter | Find in lessons → |
| Operational overhead | The effort needed to keep something running | Managing 500 clusters by hand | Find in lessons → |
No terms match that filter.