Lesson 05 of 10 · Part 2 — Direction
Building a platform roadmap
Build a Kubernetes platform roadmap people believe: gather inputs from users, incidents, mandatory work and business plans, group them into themes with measurable outcomes, prioritise honestly, leave room for the unplanned, and communicate and revise it on a rhythm.
Inputs
| Input | Where it comes from |
|---|---|
| User pain | Interviews and surveys of application teams, ticket themes, onboarding time |
| Reliability | Incidents, SLO misses, on-call load, toil |
| Mandatory work | Kubernetes version support windows, deprecated APIs, security requirements, certificate and contract expiries, hardware refresh |
| Business direction | New regions or sites, product launches, cost targets, compliance programmes |
| Technical risk | Debt register (lesson 04), capacity forecasts, single points of failure |
A roadmap is a holiday plan. Some things are fixed (the flight leaves on Saturday, passports expire in May), some are wishes (the museum, the beach). Book the fixed things first, choose wishes that matter most to the family, and keep a free afternoon for rain.
Themes and outcomes
Group work into a few themes, each with an outcome you can measure:
| Theme | Outcome | Example initiatives |
|---|---|---|
| Developer self-service | New service to production in under one day | Golden path template, onboarding by pull request |
| Reliability | Platform SLO 99.9 %; pages per week halved | Alert clean-up, multi-zone defaults, upgrade automation |
| Security and compliance | All clusters on supported versions; signed images enforced | Upgrade waves, Binary Authorization rollout |
| Cost | 20 % lower cost per workload | Rightsizing, Spot pools, cost reports per team |
| Scale | Ready for 3 new regions | Cluster blueprints, IP plan, fleet tooling |
Prioritise honestly
- Mandatory work first, with its dates.
- Then impact versus effort, with risk reduction counted as impact.
- Keep 20–30 % unplanned capacity for incidents, requests and surprises; a roadmap at 100 % is a fiction.
- Say clearly what's not on the roadmap, and why.
Format: now, next, later
| Now (this quarter) | Next (next quarter) | Later (beyond) |
|---|---|---|
| Committed, staffed, with dates | Planned, likely to change | Direction, not promises |
This keeps near-term commitments firm and long-term plans flexible.
Communicate and revise
- One page, in plain language, shared where stakeholders actually look.
- Quarterly review with application teams, product and leadership: what shipped, what moved, why.
- Monthly check inside the team; update the page when things change, with a short changelog.
- Report outcomes (did onboarding get faster?), not just delivered items.
Try it: a one-page roadmap
- Interview three application teams: what slows them down most on the platform?
- List every dated mandatory item for the next 12 months (versions, certificates, deprecations).
- Group everything into 3–5 themes with one measurable outcome each.
- Place items into now/next/later with 25 % unplanned capacity.
- Present it to your team and one stakeholder; note what surprised them.
Going deeper: saying no
Every roadmap is also a list of things you won't do. When a request doesn't fit, explain the trade-off ("doing this means the upgrade automation slips a quarter"), offer an alternative or a date, and let the right person make the call.
Recap
- Inputs: user pain, reliability, mandatory work, business direction, technical risk.
- Themes with measurable outcomes, not feature lists.
- Mandatory work first, then impact vs effort; keep unplanned capacity.
- Now / next / later; review quarterly and report outcomes.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.