08

Scoring

The Clean Code Score formula, Bayesian shrinkage, confidence levels, ranking, worked example

5 min read8 sections

This document explains exactly how the Clean Code Score is computed. The code is in src/scoring/; every constant is in scoringConfig.ts.

Inputs per developer

From the attribution engine (07):

  • analyzedLines: lines they added or changed in analyzable files, in the analyzed commits;
  • weighted[category]: Σ severity weights of the violations they introduced with high/medium confidence, for duplication, structure and hygiene.

Only introduced violations count. Existing, fixed, excluded and unattributed violations never lower a score.

The formula

For each category:

text
kloc          = analyzedLines / 1000
adjusted(cat) = ( weighted(cat) + projectDensity(cat) × priorKLOC )
                / ( kloc + priorKLOC )
subScore(cat) = 100 × e^( −adjusted(cat) / scale(cat) )

Then:

text
score = subScore(duplication) × 0.30
      + subScore(structure)   × 0.45
      + subScore(hygiene)     × 0.25

The result is clamped to 0–100 and rounded. The unrounded value is kept as scoreExact.

Why a density?

Raw violation counts punish productive developers: someone who writes 20,000 lines will have more violations than someone who writes 500. Dividing by KLOC measures how clean the code they write is, not how much they write.

Why an exponential?

e^(−density / scale) maps any density ≥ 0 to 100 … 0:

  • density 0 → 100;
  • density = scale → ~37 (one "e-fold");
  • it approaches 0 smoothly and never goes negative.

Each further unit of density costs fewer points than the one before, so one very bad file cannot drive a score to zero.

Why Bayesian shrinkage?

Without it, a developer who wrote 50 clean lines would score a perfect 100, and one who wrote 50 lines with one secret would score near 0. Neither says much.

Shrinkage adds priorKLOC (= 1 KLOC) of "pseudo-code" at the project's average density to every developer:

  • a developer with little code is pulled strongly toward the project average;
  • a developer with lots of code is barely affected; their own data dominates.

The project density per category is:

text
projectDensity(cat) = Σ over all developers of weighted(cat)
                      / (total analyzed lines of all commits / 1000)

If the whole project is clean, the prior is 0 and there is no division by zero.

Constants

ConstantDefaultMeaningConfigurable via config file?
SEVERITY_WEIGHTlow 1, medium 3, high 5, critical 8Weight per violationSee note below
CATEGORY_SCALEduplication 15, structure 45, hygiene 38Density at which a sub-score falls to ~37✅ scoring.categoryScales
CATEGORY_WEIGHTduplication 0.30, structure 0.45, hygiene 0.25Share of each sub-score in the final score✅ scoring.weights
PRIOR_KLOC1Strength of the shrinkage❌ source only
DEFAULT_MIN_LINES_FOR_RANKING1000Lines needed to be ranked✅ scoring.minLinesForRanking
ENGINE_VERSION2.0.0Cache invalidation key❌ source only

The scales were calibrated on real repositories (CleanLens itself, a Django + React monorepo, two CLI tools) so that, with the default rules, a well-kept codebase scores about 70–85 and one with real debt about 40–55.

Severity weights

What the scale means in practice

Density (weighted points per KLOC) at which a sub-score reaches a given value:

Sub-scoreduplication (scale 15)structure (scale 45)hygiene (scale 38)
901.64.74.0
803.310.08.5
705.416.113.6
5010.431.226.3
37154538

Formula: density = −scale × ln(subScore / 100).

For example, structure 80 allows about 10 weighted points per 1,000 lines, which is roughly three medium structure findings (3 × 3 = 9) per KLOC.

Worked example

Project averages: duplication 4, structure 20, hygiene 8 weighted points per KLOC.

Developer Alice wrote 2,000 lines (2 KLOC) and introduced:

  • duplication: 2 medium clones → 6 points
  • structure: 10 medium findings → 30 points
  • hygiene: 1 low + 3 medium → 10 points
text
duplication: adjusted = (6 + 4×1)  / (2 + 1) = 3.33  → 100·e^(−3.33/15) = 80.1
structure:   adjusted = (30 + 20×1)/ (2 + 1) = 16.67 → 100·e^(−16.67/45) = 69.1
hygiene:     adjusted = (10 + 8×1) / (2 + 1) = 6.00  → 100·e^(−6/38)     = 85.4

score = 80.1×0.30 + 69.1×0.45 + 85.4×0.25 = 24.0 + 31.1 + 21.3 = 76.4 → 76

The report shows 76/100 (duplication 80 · structure 69 · hygiene 85). The sub-scores tell Alice to focus on structure.

Score confidence levels

Based only on analyzedLines:

LevelAnalyzed linesHow to read it
insufficient< 300Mostly the project average; do not compare
provisional300 – 999Indicative only
reliable1,000 – 4,999Fair to compare
highly_reliable≥ 5,000Strong signal

These thresholds are in confidenceLevel() and can only be changed in source.

Ranking

  • rankable = the developer has a score and analyzedLines ≥ minLinesForRanking (default 1,000).
  • Non-rankable developers still get a score, but they are listed separately ("Not ranked — insufficient analysed code").
  • A developer with 0 analyzed lines gets score = null ("no analysed code"), which is different from a perfect score.
  • Sorting (CLI, Markdown, dashboard): rankable developers by score, lowest first, then non-rankable developers, then developers with no score.

Other per-developer numbers (display only)

FieldFormulaIn the score?
violationDensitytotal weighted points ÷ KLOC, before shrinkageNo
contributionPercentthe developer's analyzed lines ÷ all analyzed linesNo — shown separately on purpose
netQualityImpactweighted introduced − weighted fixedNo (positive = added more debt than removed)
activeDays, first/last contributionfrom git logNo

Tests

test/scoring.test.ts checks: null for zero lines; a tiny clean sample scores below a large clean one; volume is not punished; one category falls independently; scores stay within 0–100; no division by zero; the weights are 0.30/0.45/0.25; the confidence thresholds; and a configurable ranking threshold.