07

Attribution engine

Commit replay, fingerprints, attribution status and confidence

6 min read9 sections

The attribution engine decides who introduced each violation. It lives in src/attribution/ and is the core of what makes CleanLens different from blame-based tools.

The idea in one paragraph

For every analyzed commit, CleanLens analyzes each changed file before the commit (parent version) and after it, gives each violation a fingerprint, and compares the two sets. New fingerprints were introduced by the commit; missing fingerprints were fixed by it; shared ones already existed. After the walk, every violation in the current code is looked up by fingerprint to find the commit (and developer) that introduced it.

Step 1 — walking the commits (commitAttribution.ts)

attributeCommits(root, config, opts):

  1. Lists the commits to replay (newest N non-merge, oldest first; see 04 — Git layer).
  2. Processes up to 12 commits in parallel (COMMIT_CONCURRENCY). The work is mostly waiting on git, so parallelism helps a lot.
  3. For each commit, uses the cached result if there is one; otherwise it computes the commit's contribution (step 2) and stores it in the cache.
  4. Applies the contributions strictly oldest-first, so the origin registry always reflects the latest state.
  5. If the run is cancelled, it stops at the first missing contribution and marks the result aborted.

Step 2 — one commit (computeCommit)

text
parent (or empty tree) ──git diff──▶ changed files + added/removed line ranges
                                          │ filter
                                          ▼
                 analyzable, not excluded, not binary, not whitespace-only
                                          │ git cat-file --batch
                                          ▼
                            before blob          after blob
                                │                     │
                         analyzeRevision()     analyzeRevision()
                                │                     │
                         fingerprints (B)      fingerprints (A)
                                  \                   /
                                   compare the two sets

Files are skipped when they are binary, have no analyzer, match the exclude list, are a whitespace-only modification (no line ranges under -w), or have a generated header (when excludeGenerated is on).

For each remaining file:

  • Analyzed lines += the number of lines in the commit's added ranges. This is the "code the developer wrote" that scores are divided by.
  • The before and after blobs are analyzed. Results are cached by blob id for the whole run, so a file version shared by consecutive commits is only analyzed once.
  • Then the sets are compared:
Fingerprint is in…ClassificationEffect
after onlyintroducedCharged to the author (if confidence is high/medium)
before and afterexistingCounted, never charged
before onlyfixedCredited to the author (fixedWeighted)

If the after revision cannot be analyzed (parse error, missing blob), every violation from the before revision is counted as unattributed for that commit.

Step 3 — attribution confidence

Each introduced violation gets a confidence, from classifyConfidence():

ConfidenceWhenScored?
highThe file is new (added, copied, root commit), or the violation's start line is inside a changed range, or its line span overlaps a changed range✅
mediumNot on changed lines, but it belongs to a named symbol (function/class) and the commit changed lines in the file✅
lowNo evidence the edit caused it (e.g. the parent revision could not be read for an existing file)❌ shown only

Only high and medium count (SCORING_CONFIDENCES in scoringConfig.ts).

Why this matters, with examples:

  • Bob adds 60 lines to an existing function → it now breaks maxFunctionLines. The function's span overlaps Bob's added lines → high, charged to Bob.
  • Carol adds one if deep inside a function, and the function crosses the complexity limit. The finding starts at the function header (not changed), but its span overlaps Carol's line → high, charged to Carol.
  • Dave reformats a file. Under -w nothing changed → the file is skipped. Nobody inherits anything.
  • Erin renames utils.js → helpers.js without editing it. -M detects the rename, but fingerprints include the path, so every violation looks new under the new path. The rename changed no lines (no added ranges), so each one gets low confidence and Erin is not charged. Side effects of the current implementation: the old-path fingerprints count as fixed in Erin's display totals (fixed count, net quality impact; not the score), and in the report these violations show introduced / low with Erin's commit. The original author loses the link to them.

Step 4 — fingerprints (fingerprint.ts)

A fingerprint identifies "the same violation" across two revisions without using line numbers (which change whenever someone edits above it):

text
fingerprint = FNV-1a( ruleId + normalizedPath + symbolName + normalizedContext )
  • normalizedPath: forward slashes, no leading ./.
  • symbolName: the enclosing function/class, if known.
  • normalizedContext: the source lines from startLine to min(endLine, startLine + 2) (at most 3 lines), with // and # comments removed and whitespace collapsed.

Consequences:

  • Inserting lines above a violation → same fingerprint (still existing).
  • Editing whitespace or comments on those lines → same fingerprint.
  • Editing the code on the violation's first lines → a new fingerprint: the old one is fixed and the new one introduced by the same commit. The author is charged and credited; the net effect is roughly neutral for them.
  • Renaming the enclosing function → new fingerprint (the symbol changed).

Step 5 — merging with the current code (attributeReport.ts)

After the walk, each violation from the HEAD scan gets its final status:

ConditionStatusConfidencedeveloperId
Path matches the exclude listexcluded—none
File cannot be read from diskunattributedlownone
Fingerprint's latest origin is introducedintroducedthe origin'sthe author if high/medium
Otherwise (predates the window, or was not seen in the walk)existinglownone

introducedBy, introducedCommit and introducedAt are filled in whenever the origin is known, even for low confidence, so you can still see the commit.

Step 6 — per-developer totals

For each developer (matched by lowercased email → developer id), the engine sums over their commits:

FieldMeaning
analyzedLinesLines they added/changed in analyzable files
introducedWeighted[category]Σ weights of their high/medium introduced violations, per category
fixedWeightedΣ weights of violations their commits removed
counts.introducedAll introduced (any confidence)
counts.introducedScoredIntroduced with high/medium confidence
counts.fixed, counts.existing, counts.unattributedAs classified per commit

Note that these are historical counts: they include violations introduced and later fixed. So a developer's "New violations" can be larger than the number of their violations still in the code.

These totals feed the score: see 08 — Scoring.

Guarantees and edge cases

CaseBehavior
Root commit (no parent)Diffed against the empty tree; every violation is introduced with high confidence
Merge commitsSkipped (--no-merges)
Whitespace-only changeIgnored
RenameDetected with -M; the renamer is not charged (low confidence), but the rename breaks the link to the original author and adds "fixed" credit for the renamer (display only)
CopyDetected with -C; treated as a new file, so the copier is charged (high confidence) for the violations in the copy
Deleted fileIts violations are fixed by the deleter
Author not in the developer listCounted in project totals, not in any developer
Violation older than the windowexisting; never charged
Cancelled runContributions after the gap are ignored

Tests that pin this behavior

test/integration.test.ts builds real temporary Git repositories and checks that: a pre-existing violation is not charged to a later editor; a violation on modified lines is charged; a removed violation is recorded as fixed; a whitespace reformat transfers nothing; generated and excluded files are ignored; a root commit is safe; and a tiny clean contributor does not get 100 and is not ranked.