01

Overview & core concepts

What CleanLens does, the problem it solves, and the vocabulary used in every other doc

4 min read6 sections

What CleanLens is

CleanLens is a local, Git-aware Clean Code analyzer. It runs in two forms that share one engine:

  • a VS Code extension (sidebar, dashboard webview, inline diagnostics), and
  • a cleanlens command-line tool (text, Markdown or JSON report, CI gate).

It analyzes JavaScript / TypeScript (.js .jsx .mjs .cjs .ts .tsx) and Python (.py) files against 14 configurable rules, then works out which developer introduced each problem and gives each developer a 0–100 Clean Code Score.

Everything runs on your machine. It makes no network requests, collects no telemetry, and needs no sign-in (see PRIVACY.md).

The problem it solves

Most "who wrote bad code" tools use git blame on the current files. That is unfair: if Alice edits one line in a 2,000-line legacy file, blame makes her the latest author of that file, and a file-level tool puts the file's old problems on her.

CleanLens replays the commit history instead. For each commit it analyzes each changed file before and after the commit and compares the two sets of violations:

  • a violation that appears only after the commit, on lines the commit changed → introduced by that commit's author;
  • a violation present before and after → existing, charged to nobody;
  • a violation present before but gone after → fixed, credited to the author.

Scores are then based on how many problems a developer introduced per 1,000 lines they wrote, not on raw counts. A statistical correction (Bayesian shrinkage) stops a small sample from producing an extreme score.

Design principles

  1. Fairness over convenience. A developer is charged only with evidence (high or medium confidence). Anything uncertain is shown but not scored.
  2. Quality is separate from volume. "Contribution %" (how much code you wrote) is always shown next to the score and never mixed into it.
  3. Local and private. Only git subprocesses and file reads. No servers.
  4. Configurable. Rules, severities, limits, exclusions, scoring constants, identities and history depth are all settable (see 10 — Configuration reference and 11 — Customization guide).
  5. One engine, two front ends. The extension and the CLI call the same analyzeRepository() function, so their numbers always match.

Glossary

These terms are used throughout the docs and the UI.

TermMeaning
RuleOne of the 14 checks, e.g. maxFunctionLines. Configured with enabled, severity, and optionally limit / options.
Finding (RawFinding)What an analyzer returns: rule id, file, line range, message, optional symbol name.
ViolationA finding enriched with severity, category, weight and a stable id.
Severitylow, medium, high, critical. Sets the violation's weight (1 / 3 / 5 / 8 by default).
CategoryEvery rule belongs to one of three scoring categories: duplication, structure, hygiene.
Weighted pointsThe sum of severity weights of a set of violations.
KLOC1,000 lines of code. Densities are "weighted points per KLOC".
Analyzed linesLines a developer added or modified, in analyzable files, across the analyzed commits.
FingerprintA hash of rule + path + symbol + normalized code context. It identifies "the same violation" in two file revisions even if its line number moved.
Attribution statusintroduced, existing, fixed, excluded, unattributed. See 07.
Attribution confidencehigh, medium, low: how sure the engine is that the author caused the violation. Only high/medium affect scores.
Score confidence levelinsufficient, provisional, reliable, highly_reliable: how much code the score is based on.
RankableWhether a developer has enough analyzed lines (default ≥ 1,000) to appear in the ranking.
Net quality impactWeighted points introduced minus weighted points fixed. Display only.
HEAD scanThe analysis of the current working-tree files. It produces the violation list.
Commit walkThe replay of the newest N non-merge commits. It produces attribution and scores.
PresetA starter rule configuration: javascript, react, python, django.

What you get from a run

  • Project totals: developers, violations, unattributed violations, analyzed files, analyzed commits, timestamp.
  • Per developer: Clean Code Score and its three sub-scores, analyzed lines, score confidence, contribution %, new / fixed / existing / unattributed counts, weighted points, net quality impact, density per KLOC, active days, first/last contribution, emails.
  • Per violation: severity, rule, message, file and line range, status, confidence, the developer who introduced it, and the commit and date.

Where to go next