Technical Leadership for Platform Engineers›10 · Measuring platform success

Lesson 10 of 10 · Part 3 — People and outcomes

Measuring platform success

Know whether the platform is working: delivery metrics (DORA), reliability (SLOs), developer experience, adoption, toil and cost, how to collect them without heavy tooling, and how to use them for decisions rather than for ranking teams.

Practitioner → Lead
Key wordsplatform metricsDORA metricsdeployment frequencylead time for changeschange failure ratetime to restoreSLOsdeveloper experienceadoptiontoilcost per workloadSPACE framework

What to measure

Area Metrics Tells you
Delivery (DORA) Deployment frequency, lead time for changes, change failure rate, time to restore Whether teams ship quickly and safely on the platform
Reliability Platform SLOs (API availability, deploy success, DNS, ingress), error-budget burn Whether the platform itself is dependable
Developer experience Time to first deploy for a new service, satisfaction survey, top friction points Whether it feels better to use
Adoption % of services on golden paths, migrations completed Whether teams choose it
Toil Hours per week on manual, repetitive platform work; tickets per team Whether automation is paying off
Cost Cost per workload or per request, idle capacity Whether it's efficient

Measuring a platform is like a school report: one grade (uptime) doesn't tell you much. You want a few subjects (speed, safety, happiness, cost), compared with last term, so you can see where to help, not to decide who sits at the front.

Collect it cheaply

  • DORA: deployment events from CI/CD and GitOps (each sync is a deploy), incident records for failures and restore time. A small script over Git history and the incident tracker is enough to start.
  • SLOs: from your monitoring system (Prometheus, Cloud Monitoring), defined for the platform's user-facing functions.
  • Experience: a three-question survey each quarter ("How easy is it to ship a change? What slows you down most? What should we fix next?").
  • Toil and tickets: label tickets and time spent; review monthly.

Use metrics for decisions

  • Set a baseline, then track trends; absolute numbers matter less than direction.
  • Pair every speed metric with a stability metric, so one isn't bought with the other.
  • Bring metrics to roadmap reviews (lesson 05): "lead time dropped from 3 days to 4 hours after the golden path" justifies the next investment.
  • Look at metrics per team to find who needs help, never to rank teams publicly.

Report simply

A monthly one-page summary: the five numbers that matter, their trend, one sentence on why they moved, and what you're doing next. Stakeholders read one page; they don't read dashboards.

Try it: a first platform scorecard

  1. Count deploys per week for the last 8 weeks from your CI or GitOps history.
  2. For the same period, count failed changes and how long each took to fix.
  3. Measure how long the last three new services took from repository creation to first production deploy.
  4. Run the three-question survey with five developers.
  5. Put it all on one page with trends and one proposed improvement.

Going deeper: beyond DORA

Frameworks such as SPACE (satisfaction, performance, activity, communication, efficiency) remind you that developer productivity has several dimensions. Pick a few signals from different dimensions rather than one number.

Recap

  • Measure delivery (DORA), reliability (SLOs), experience, adoption, toil and cost.
  • Start cheap: Git history, incident records, a short survey.
  • Track trends against a baseline; pair speed with stability.
  • Use metrics to decide and help, never to rank teams; report on one page.

This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.