How we measure our own software
Fifty-four criteria across the ten ISO/IEC 25010 characteristics, scored on every commit. The rules underneath the number matter more than the number, so they are here and the number is not.
The rubric
Weights sum to 100. Each characteristic aggregates its own criteria; the composite is a weighted mean over what was actually measured.
The rules that make it honest
Each of these was learned by getting it wrong first.
Unmeasured is not zero
A criterion nobody has measured is reported as unmeasured and excluded from the denominator. Scoring it zero would turn missing measurement into a defect and send someone hunting a bug that has not been shown to exist.
A score never travels without its coverage
Every score is published with the percentage of criteria actually measured. A high score over a third of the rubric is a different claim from the same score over all of it, and they must not read alike.
Two scores, never averaged
Release readiness is ordinal — a rung on a ladder. Quality is cardinal — a weighted index. A mean of a ladder position and a weighted index has no referent, so they are reported side by side and never combined.
A gate nobody has watched fail is not a gate
Every check has to be shown going red on a planted defect before it is trusted. This was learned the expensive way: seven coverage thresholds were declared and never evaluated, so suites ran with coverage off against floors nobody had checked.
Questions this page didn't answer
Send them and we will answer directly, including the ones with awkward answers.
Get in touch