Lesson 04 of 10 · Part 2 — Direction
Technical debt versus feature delivery
Balance technical debt with feature delivery: tell debt apart from mere imperfection, make it visible with its cost, tie it to risk and speed the business cares about, reserve steady capacity for it, pay it down alongside features, and never let support deadlines become debt.
What counts as technical debt
Technical debt is a past shortcut that now has an ongoing cost: extra effort for every change, extra risk, extra toil. Not everything imperfect is debt; code that's ugly but never changes and never breaks costs nothing.
Typical platform debt: manual upgrade steps, snowflake clusters, unpinned versions, missing tests for critical tooling, alerts nobody trusts, an IP plan with no room, an ageing CNI or Kubernetes version heading for end of support.
Technical debt is like dirty dishes. Leaving one plate for later is fine. Leave them for a week, and cooking dinner (shipping a feature) means washing up first, every single time.
Make it visible, with its cost
Keep a debt register (a board or a file in Git) with, for each item:
| Field | Example |
|---|---|
| What | Node upgrades are manual on edge clusters |
| Cost now | 2 engineer-days per upgrade wave, 6 waves a year; two incidents from missed steps |
| Risk | Falling behind supported versions; security patches delayed |
| Effort to fix | 3 weeks to automate with a pipeline |
| Payback | About 4 months |
Numbers turn "we should clean this up" into a decision a product owner can weigh against features.
Get time for it
- Standing allocation: agree a share of capacity (often 20–30%) for debt, reliability and upgrades, and protect it.
- Pay it down with features: when a feature touches an area, include the fix for that area's debt in the estimate.
- Tie it to business outcomes: "automating upgrades lets us patch security issues in days instead of weeks" lands better than "cleaner code".
- Mandatory work is not debt: end-of-life versions, expiring certificates and deprecated APIs have dates. Put them on the roadmap as commitments (lesson 05).
Decide what not to fix
Some debt should stay: low cost, rarely touched, or about to be replaced. Write that down too, so nobody wastes time on it, and revisit when circumstances change.
Prevent new debt
- Definition of done includes tests, docs, monitoring and runbooks for platform changes.
- Shortcuts are allowed, but recorded with an owner and a date.
- The boy-scout rule: leave each area a little better than you found it.
Try it: a debt register in an afternoon
- List the ten things that most slow your team down or wake people up at night.
- For each, estimate cost now (hours per month, incidents), risk and effort to fix.
- Sort by payback; pick the top three.
- Write a half-page case for the top item in business terms, and propose a standing capacity share.
- Add a "debt" label to your board and review the register monthly.
Going deeper: talking with product
Product owners aren't the enemy of quality; they lack the information to weigh it. Bring the register to planning, show what the last debt fix bought (fewer incidents, faster delivery), and let them see the trade-off as clearly as you do.
Recap
- Debt is a shortcut with an ongoing cost; imperfection alone isn't debt.
- Keep a register with cost, risk, effort and payback.
- Get time through a standing allocation and by fixing debt with features.
- Mandatory work (end of life, certificates, deprecations) goes on the roadmap, not the debt list.
- Decide what not to fix, and prevent new debt with a clear definition of done.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.