Skip to main content
    Back to the overview
    Engineering quality

    How we measure our own software

    Fifty-four criteria across the ten ISO/IEC 25010 characteristics, scored on every commit. The rules underneath the number matter more than the number, so they are here and the number is not.

    The rubric

    Weights sum to 100. Each characteristic aggregates its own criteria; the composite is a weighted mean over what was actually measured.

    Functional Suitability16
    Reliability14
    Security14
    Maintainability13
    Process & Delivery11
    Performance Efficiency9
    Interaction Capability8
    Compatibility6
    Flexibility5
    Safety4

    The rules that make it honest

    Each of these was learned by getting it wrong first.

    Unmeasured is not zero

    A criterion nobody has measured is reported as unmeasured and excluded from the denominator. Scoring it zero would turn missing measurement into a defect and send someone hunting a bug that has not been shown to exist.

    A score never travels without its coverage

    Every score is published with the percentage of criteria actually measured. A high score over a third of the rubric is a different claim from the same score over all of it, and they must not read alike.

    Two scores, never averaged

    Release readiness is ordinal — a rung on a ladder. Quality is cardinal — a weighted index. A mean of a ladder position and a weighted index has no referent, so they are reported side by side and never combined.

    A gate nobody has watched fail is not a gate

    Every check has to be shown going red on a planted defect before it is trusted. This was learned the expensive way: seven coverage thresholds were declared and never evaluated, so suites ran with coverage off against floors nobody had checked.

    Questions this page didn't answer

    Send them and we will answer directly, including the ones with awkward answers.

    Get in touch