Attribution engine
Commit replay, fingerprints, attribution status and confidence
The attribution engine decides who introduced each violation. It lives in src/attribution/ and is the core of what makes CleanLens different from blame-based tools.
The idea in one paragraph
For every analyzed commit, CleanLens analyzes each changed file before the commit (parent version) and after it, gives each violation a fingerprint, and compares the two sets. New fingerprints were introduced by the commit; missing fingerprints were fixed by it; shared ones already existed. After the walk, every violation in the current code is looked up by fingerprint to find the commit (and developer) that introduced it.
Step 1 — walking the commits (commitAttribution.ts)
attributeCommits(root, config, opts):
- Lists the commits to replay (newest N non-merge, oldest first; see 04 — Git layer).
- Processes up to 12 commits in parallel (
COMMIT_CONCURRENCY). The work is mostly waiting ongit, so parallelism helps a lot. - For each commit, uses the cached result if there is one; otherwise it computes the commit's contribution (step 2) and stores it in the cache.
- Applies the contributions strictly oldest-first, so the origin registry always reflects the latest state.
- If the run is cancelled, it stops at the first missing contribution and
marks the result
aborted.
Step 2 — one commit (computeCommit)
parent (or empty tree) ──git diff──▶ changed files + added/removed line ranges
│ filter
▼
analyzable, not excluded, not binary, not whitespace-only
│ git cat-file --batch
▼
before blob after blob
│ │
analyzeRevision() analyzeRevision()
│ │
fingerprints (B) fingerprints (A)
\ /
compare the two setsFiles are skipped when they are binary, have no analyzer, match the exclude
list, are a whitespace-only modification (no line ranges under -w), or have a
generated header (when excludeGenerated is on).
For each remaining file:
- Analyzed lines += the number of lines in the commit's
addedranges. This is the "code the developer wrote" that scores are divided by. - The before and after blobs are analyzed. Results are cached by blob id for the whole run, so a file version shared by consecutive commits is only analyzed once.
- Then the sets are compared:
| Fingerprint is in… | Classification | Effect |
|---|---|---|
| after only | introduced | Charged to the author (if confidence is high/medium) |
| before and after | existing | Counted, never charged |
| before only | fixed | Credited to the author (fixedWeighted) |
If the after revision cannot be analyzed (parse error, missing blob), every violation from the before revision is counted as unattributed for that commit.
Step 3 — attribution confidence
Each introduced violation gets a confidence, from classifyConfidence():
| Confidence | When | Scored? |
|---|---|---|
| high | The file is new (added, copied, root commit), or the violation's start line is inside a changed range, or its line span overlaps a changed range | ✅ |
| medium | Not on changed lines, but it belongs to a named symbol (function/class) and the commit changed lines in the file | ✅ |
| low | No evidence the edit caused it (e.g. the parent revision could not be read for an existing file) | ❌ shown only |
Only high and medium count (SCORING_CONFIDENCES in
scoringConfig.ts).
Why this matters, with examples:
- Bob adds 60 lines to an existing function → it now breaks
maxFunctionLines. The function's span overlaps Bob's added lines → high, charged to Bob. - Carol adds one
ifdeep inside a function, and the function crosses the complexity limit. The finding starts at the function header (not changed), but its span overlaps Carol's line → high, charged to Carol. - Dave reformats a file. Under
-wnothing changed → the file is skipped. Nobody inherits anything. - Erin renames
utils.js→helpers.jswithout editing it.-Mdetects the rename, but fingerprints include the path, so every violation looks new under the new path. The rename changed no lines (noaddedranges), so each one gets low confidence and Erin is not charged. Side effects of the current implementation: the old-path fingerprints count asfixedin Erin's display totals (fixed count, net quality impact; not the score), and in the report these violations showintroduced/lowwith Erin's commit. The original author loses the link to them.
Step 4 — fingerprints (fingerprint.ts)
A fingerprint identifies "the same violation" across two revisions without using line numbers (which change whenever someone edits above it):
fingerprint = FNV-1a( ruleId + normalizedPath + symbolName + normalizedContext )- normalizedPath: forward slashes, no leading
./. - symbolName: the enclosing function/class, if known.
- normalizedContext: the source lines from
startLinetomin(endLine, startLine + 2)(at most 3 lines), with//and#comments removed and whitespace collapsed.
Consequences:
- Inserting lines above a violation → same fingerprint (still
existing). - Editing whitespace or comments on those lines → same fingerprint.
- Editing the code on the violation's first lines → a new fingerprint: the old
one is
fixedand the new oneintroducedby the same commit. The author is charged and credited; the net effect is roughly neutral for them. - Renaming the enclosing function → new fingerprint (the symbol changed).
Step 5 — merging with the current code (attributeReport.ts)
After the walk, each violation from the HEAD scan gets its final status:
| Condition | Status | Confidence | developerId |
|---|---|---|---|
| Path matches the exclude list | excluded | — | none |
| File cannot be read from disk | unattributed | low | none |
Fingerprint's latest origin is introduced | introduced | the origin's | the author if high/medium |
| Otherwise (predates the window, or was not seen in the walk) | existing | low | none |
introducedBy, introducedCommit and introducedAt are filled in whenever the
origin is known, even for low confidence, so you can still see the commit.
Step 6 — per-developer totals
For each developer (matched by lowercased email → developer id), the engine sums over their commits:
| Field | Meaning |
|---|---|
analyzedLines | Lines they added/changed in analyzable files |
introducedWeighted[category] | Σ weights of their high/medium introduced violations, per category |
fixedWeighted | Σ weights of violations their commits removed |
counts.introduced | All introduced (any confidence) |
counts.introducedScored | Introduced with high/medium confidence |
counts.fixed, counts.existing, counts.unattributed | As classified per commit |
Note that these are historical counts: they include violations introduced and later fixed. So a developer's "New violations" can be larger than the number of their violations still in the code.
These totals feed the score: see 08 — Scoring.
Guarantees and edge cases
| Case | Behavior |
|---|---|
| Root commit (no parent) | Diffed against the empty tree; every violation is introduced with high confidence |
| Merge commits | Skipped (--no-merges) |
| Whitespace-only change | Ignored |
| Rename | Detected with -M; the renamer is not charged (low confidence), but the rename breaks the link to the original author and adds "fixed" credit for the renamer (display only) |
| Copy | Detected with -C; treated as a new file, so the copier is charged (high confidence) for the violations in the copy |
| Deleted file | Its violations are fixed by the deleter |
| Author not in the developer list | Counted in project totals, not in any developer |
| Violation older than the window | existing; never charged |
| Cancelled run | Contributions after the gap are ignored |
Tests that pin this behavior
test/integration.test.ts builds real temporary Git repositories and checks that: a pre-existing violation is not charged to a later editor; a violation on modified lines is charged; a removed violation is recorded as fixed; a whitespace reformat transfers nothing; generated and excluded files are ignored; a root commit is safe; and a tiny clean contributor does not get 100 and is not ranked.