Scoring
The Clean Code Score formula, Bayesian shrinkage, confidence levels, ranking, worked example
This document explains exactly how the Clean Code Score is computed. The code is in src/scoring/; every constant is in scoringConfig.ts.
Inputs per developer
From the attribution engine (07):
analyzedLines: lines they added or changed in analyzable files, in the analyzed commits;weighted[category]: Σ severity weights of the violations they introduced withhigh/mediumconfidence, forduplication,structureandhygiene.
Only introduced violations count. Existing, fixed, excluded and unattributed violations never lower a score.
The formula
For each category:
kloc = analyzedLines / 1000
adjusted(cat) = ( weighted(cat) + projectDensity(cat) × priorKLOC )
/ ( kloc + priorKLOC )
subScore(cat) = 100 × e^( −adjusted(cat) / scale(cat) )Then:
score = subScore(duplication) × 0.30
+ subScore(structure) × 0.45
+ subScore(hygiene) × 0.25The result is clamped to 0–100 and rounded. The unrounded value is kept as
scoreExact.
Why a density?
Raw violation counts punish productive developers: someone who writes 20,000 lines will have more violations than someone who writes 500. Dividing by KLOC measures how clean the code they write is, not how much they write.
Why an exponential?
e^(−density / scale) maps any density ≥ 0 to 100 … 0:
- density 0 → 100;
- density =
scale→ ~37 (one "e-fold"); - it approaches 0 smoothly and never goes negative.
Each further unit of density costs fewer points than the one before, so one very bad file cannot drive a score to zero.
Why Bayesian shrinkage?
Without it, a developer who wrote 50 clean lines would score a perfect 100, and one who wrote 50 lines with one secret would score near 0. Neither says much.
Shrinkage adds priorKLOC (= 1 KLOC) of "pseudo-code" at the project's
average density to every developer:
- a developer with little code is pulled strongly toward the project average;
- a developer with lots of code is barely affected; their own data dominates.
The project density per category is:
projectDensity(cat) = Σ over all developers of weighted(cat)
/ (total analyzed lines of all commits / 1000)If the whole project is clean, the prior is 0 and there is no division by zero.
Constants
| Constant | Default | Meaning | Configurable via config file? |
|---|---|---|---|
SEVERITY_WEIGHT | low 1, medium 3, high 5, critical 8 | Weight per violation | See note below |
CATEGORY_SCALE | duplication 15, structure 45, hygiene 38 | Density at which a sub-score falls to ~37 | ✅ scoring.categoryScales |
CATEGORY_WEIGHT | duplication 0.30, structure 0.45, hygiene 0.25 | Share of each sub-score in the final score | ✅ scoring.weights |
PRIOR_KLOC | 1 | Strength of the shrinkage | ❌ source only |
DEFAULT_MIN_LINES_FOR_RANKING | 1000 | Lines needed to be ranked | ✅ scoring.minLinesForRanking |
ENGINE_VERSION | 2.0.0 | Cache invalidation key | ❌ source only |
The scales were calibrated on real repositories (CleanLens itself, a Django + React monorepo, two CLI tools) so that, with the default rules, a well-kept codebase scores about 70–85 and one with real debt about 40–55.
Severity weights
What the scale means in practice
Density (weighted points per KLOC) at which a sub-score reaches a given value:
| Sub-score | duplication (scale 15) | structure (scale 45) | hygiene (scale 38) |
|---|---|---|---|
| 90 | 1.6 | 4.7 | 4.0 |
| 80 | 3.3 | 10.0 | 8.5 |
| 70 | 5.4 | 16.1 | 13.6 |
| 50 | 10.4 | 31.2 | 26.3 |
| 37 | 15 | 45 | 38 |
Formula: density = −scale × ln(subScore / 100).
For example, structure 80 allows about 10 weighted points per 1,000 lines,
which is roughly three medium structure findings (3 × 3 = 9) per KLOC.
Worked example
Project averages: duplication 4, structure 20, hygiene 8 weighted points per KLOC.
Developer Alice wrote 2,000 lines (2 KLOC) and introduced:
- duplication: 2 medium clones → 6 points
- structure: 10 medium findings → 30 points
- hygiene: 1 low + 3 medium → 10 points
duplication: adjusted = (6 + 4×1) / (2 + 1) = 3.33 → 100·e^(−3.33/15) = 80.1
structure: adjusted = (30 + 20×1)/ (2 + 1) = 16.67 → 100·e^(−16.67/45) = 69.1
hygiene: adjusted = (10 + 8×1) / (2 + 1) = 6.00 → 100·e^(−6/38) = 85.4
score = 80.1×0.30 + 69.1×0.45 + 85.4×0.25 = 24.0 + 31.1 + 21.3 = 76.4 → 76The report shows 76/100 (duplication 80 · structure 69 · hygiene 85). The
sub-scores tell Alice to focus on structure.
Score confidence levels
Based only on analyzedLines:
| Level | Analyzed lines | How to read it |
|---|---|---|
insufficient | < 300 | Mostly the project average; do not compare |
provisional | 300 – 999 | Indicative only |
reliable | 1,000 – 4,999 | Fair to compare |
highly_reliable | ≥ 5,000 | Strong signal |
These thresholds are in confidenceLevel() and can only be changed in source.
Ranking
rankable= the developer has a score andanalyzedLines ≥ minLinesForRanking(default 1,000).- Non-rankable developers still get a score, but they are listed separately ("Not ranked — insufficient analysed code").
- A developer with 0 analyzed lines gets
score = null("no analysed code"), which is different from a perfect score. - Sorting (CLI, Markdown, dashboard): rankable developers by score, lowest first, then non-rankable developers, then developers with no score.
Other per-developer numbers (display only)
| Field | Formula | In the score? |
|---|---|---|
violationDensity | total weighted points ÷ KLOC, before shrinkage | No |
contributionPercent | the developer's analyzed lines ÷ all analyzed lines | No — shown separately on purpose |
netQualityImpact | weighted introduced − weighted fixed | No (positive = added more debt than removed) |
activeDays, first/last contribution | from git log | No |
Tests
test/scoring.test.ts checks: null for zero lines; a tiny clean sample scores below a large clean one; volume is not punished; one category falls independently; scores stay within 0–100; no division by zero; the weights are 0.30/0.45/0.25; the confidence thresholds; and a configurable ranking threshold.