Overview & core concepts
What CleanLens does, the problem it solves, and the vocabulary used in every other doc
What CleanLens is
CleanLens is a local, Git-aware Clean Code analyzer. It runs in two forms that share one engine:
- a VS Code extension (sidebar, dashboard webview, inline diagnostics), and
- a
cleanlenscommand-line tool (text, Markdown or JSON report, CI gate).
It analyzes JavaScript / TypeScript (.js .jsx .mjs .cjs .ts .tsx) and
Python (.py) files against 14 configurable rules, then works out which
developer introduced each problem and gives each developer a 0–100
Clean Code Score.
Everything runs on your machine. It makes no network requests, collects no telemetry, and needs no sign-in (see PRIVACY.md).
The problem it solves
Most "who wrote bad code" tools use git blame on the current files. That is
unfair: if Alice edits one line in a 2,000-line legacy file, blame makes her the
latest author of that file, and a file-level tool puts the file's old problems
on her.
CleanLens replays the commit history instead. For each commit it analyzes each changed file before and after the commit and compares the two sets of violations:
- a violation that appears only after the commit, on lines the commit changed → introduced by that commit's author;
- a violation present before and after → existing, charged to nobody;
- a violation present before but gone after → fixed, credited to the author.
Scores are then based on how many problems a developer introduced per 1,000 lines they wrote, not on raw counts. A statistical correction (Bayesian shrinkage) stops a small sample from producing an extreme score.
Design principles
- Fairness over convenience. A developer is charged only with evidence (high or medium confidence). Anything uncertain is shown but not scored.
- Quality is separate from volume. "Contribution %" (how much code you wrote) is always shown next to the score and never mixed into it.
- Local and private. Only
gitsubprocesses and file reads. No servers. - Configurable. Rules, severities, limits, exclusions, scoring constants, identities and history depth are all settable (see 10 — Configuration reference and 11 — Customization guide).
- One engine, two front ends. The extension and the CLI call the same
analyzeRepository()function, so their numbers always match.
Glossary
These terms are used throughout the docs and the UI.
| Term | Meaning |
|---|---|
| Rule | One of the 14 checks, e.g. maxFunctionLines. Configured with enabled, severity, and optionally limit / options. |
Finding (RawFinding) | What an analyzer returns: rule id, file, line range, message, optional symbol name. |
| Violation | A finding enriched with severity, category, weight and a stable id. |
| Severity | low, medium, high, critical. Sets the violation's weight (1 / 3 / 5 / 8 by default). |
| Category | Every rule belongs to one of three scoring categories: duplication, structure, hygiene. |
| Weighted points | The sum of severity weights of a set of violations. |
| KLOC | 1,000 lines of code. Densities are "weighted points per KLOC". |
| Analyzed lines | Lines a developer added or modified, in analyzable files, across the analyzed commits. |
| Fingerprint | A hash of rule + path + symbol + normalized code context. It identifies "the same violation" in two file revisions even if its line number moved. |
| Attribution status | introduced, existing, fixed, excluded, unattributed. See 07. |
| Attribution confidence | high, medium, low: how sure the engine is that the author caused the violation. Only high/medium affect scores. |
| Score confidence level | insufficient, provisional, reliable, highly_reliable: how much code the score is based on. |
| Rankable | Whether a developer has enough analyzed lines (default ≥ 1,000) to appear in the ranking. |
| Net quality impact | Weighted points introduced minus weighted points fixed. Display only. |
| HEAD scan | The analysis of the current working-tree files. It produces the violation list. |
| Commit walk | The replay of the newest N non-merge commits. It produces attribution and scores. |
| Preset | A starter rule configuration: javascript, react, python, django. |
What you get from a run
- Project totals: developers, violations, unattributed violations, analyzed files, analyzed commits, timestamp.
- Per developer: Clean Code Score and its three sub-scores, analyzed lines, score confidence, contribution %, new / fixed / existing / unattributed counts, weighted points, net quality impact, density per KLOC, active days, first/last contribution, emails.
- Per violation: severity, rule, message, file and line range, status, confidence, the developer who introduced it, and the commit and date.
Where to go next
- To run it now: 02 — Getting started.
- To understand the internals: 03 — Architecture.
- To tune it: 11 — Customization guide.