Analyzers
File scanning, glob matching, the JS/TS and Python analyzers, duplicate detection, secret and name heuristics
The analyzers turn source text into findings. They live in src/analyzers/. This document covers how they work; 06 — Rules reference covers what each rule checks.
Overview
┌─▶ jsAnalyzer.analyzeJs() (.js .jsx .mjs .cjs .ts .tsx)
file text ─ analyzerFor┤
└─▶ pyAnalyzer.analyzePy() (.py)
│
▼ RawFinding[]
duplicateCode.detectDuplicates()
│
▼
model.toViolation() → Violation[]Every analyzer has the same signature:
(relPath: string, text: string, config: CleanCodeConfig) => RawFinding[]It is a pure function: no I/O, no Git. That is why the same code runs on disk files (HEAD scan), on historical blobs (commit walk) and on unsaved editor buffers (live diagnostics).
runAnalyzers.ts — the HEAD scan
analyzeProject(root, config, excludeGlobs, onProgress):
scanFiles()collects every file with a supported extension that is not excluded.- For each file:
- read it as UTF-8 (unreadable → skipped);
- if
excludeGeneratedis notfalseand the header marks the file as generated → skipped and counted inexcludedFiles; - run the analyzer. A parse error skips that file, not the whole run.
- Run
detectDuplicates()once across all kept files. - Convert each finding to a
Violationand sort by severity (critical → low), then path, then line.
analyzerFor(relPath) chooses the analyzer from the extension, or returns
null for unsupported files.
fileScanner.ts — which files are read
- Recursively walks the repository root.
- Always skips the directories
.git,node_modules,.hg,.svn. - Skips a directory if the exclude matcher matches
dirordir/, so an excluded folder is never entered. - Keeps a file only if its extension is supported, it is not excluded, and it is
≤ 1,000,000 bytes (
MAX_FILE_BYTES). Bigger files are almost always generated or bundled. - Paths are normalized to forward slashes on every OS.
glob.ts — the exclude pattern language
CleanLens uses a small glob dialect of its own, not a full .gitignore engine.
| Syntax | Meaning |
|---|---|
* | Any characters except / |
** | Any characters including / (**/ also matches zero directories) |
? | One character except / |
| anything else | Literal (regex special characters are escaped) |
ExcludeMatcher adds two conveniences:
- Unrooted patterns match at any depth.
venv/**becomes**/venv/**, so it also excludesbackend/venv/…. Patterns that start with/(rooted) or**/are used as they are, with the leading/removed. - Patterns without
/also match the file's base name.*.min.jsexcludesa/b/c.min.js.
Not supported: negation (!pattern), character classes ([abc]), braces ({a,b}).
excludes.ts — the full exclude list
The final list is built by mergedExcludeGlobs(). Sources are added in this
order and nothing ever removes an earlier entry:
- Built-in defaults (
DEFAULT_EXCLUDE):**/node_modules/**,**/dist/**,**/build/**,**/coverage/**,**/.next/**,**/vendor/**,**/__pycache__/**,**/staticfiles/**,**/generated/**,*.min.js,*.map,*.generated.*. - Migrations
**/migrations/**, unlessexcludeMigrations: false. - Lock files
package-lock.json,yarn.lock,pnpm-lock.yaml, unlessexcludeLockFiles: false. (These are not analyzable anyway; the flag only affects the exclude list and the cache hash.) linguist-generatedpaths from.gitattributes, unlessexcludeGenerated: false. A pattern without/becomes**/<pattern>..gitignoreentries (best effort): comments and!negations are skipped; a trailing/or a name without a dot also adds<name>/**.- The config's own
excludelist (the preset's list plus your additions, plus the local override's and the VS Code setting's entries).
Generated-file headers
hasGeneratedHeader(text) checks the first 5 lines, case-insensitively, for
any of: @generated, generated file, auto-generated, autogenerated,
do not edit, code generated. Matching files are skipped in both the HEAD
scan and the commit walk (unless excludeGenerated: false).
jsAnalyzer.ts — JavaScript and TypeScript
- Parses with the TypeScript compiler API (
ts.createSourceFile), choosing the script kind from the extension (TS, TSX, JSX, JS). No type checking and notsconfigare involved; only syntax. - Walks the AST once with a
depthcounter for nesting. - "Function-like" nodes: function declarations and expressions, arrow functions, methods, constructors, getters and setters.
- Cyclomatic complexity starts at 1 and adds 1 for each
if,?:,for,for…in,for…of,while,do,catch, non-emptycase, and&&,||,??. It does not descend into nested functions; they are counted separately. - Symbol names: findings about a function or class carry its name
(
symbolName). This makes fingerprints stable and enablesmediumattribution confidence. - Usage counts for
detectUnusedCode: every identifier in the file is counted once up front; a declaration that appears only once (itself) is unused.
pyAnalyzer.ts — Python
Python is analyzed without a full parser, using lines and indentation:
- Tabs are expanded to 4 spaces;
#comments are stripped (with string awareness). - A
def/async def/classblock ends at the next non-blank line with indentation ≤ the header's. - Parameters are read from the signature (across lines up to
)), ignoringself,cls,*args,**kwargs. - Complexity is approximate: 1 + lines starting with
if,elif,for,while,except,case+ eachand/orin the body. - Nesting depth is
indent / 4on lines that start withif,elif,else,for,while,with,try,except. It includes the enclosingdef/classlevels, so a method body starts at depth 2. - A docstring is a first body line starting with
"""or'''.
This makes the Python analyzer fast and dependency-free, but less exact than the JS one. See 17 — Limitations.
duplicateCode.ts — duplicate blocks
- Normalize each line: trim and collapse whitespace. Drop lines that are:
shorter than 12 characters, only brackets/punctuation, comments (
//,#,*), orimport/from/exportlines. - Slide a window of
minLines(option, default 6, minimum 3) kept lines over each file. A window is skipped if its lines are spread over more than3 × minLinesoriginal lines (not contiguous enough). - The first time a window's text is seen, it is recorded. When it is seen again (same file or another file), the later location is a hit.
- Overlapping or adjacent hits are merged into one finding per duplicated block. A 40-line clone is one violation, not 35.
Notes:
- The first occurrence is never flagged; only the copies found later in scan order are.
- In the HEAD scan, detection runs across all files. In the commit walk
(
analyzeRevision), it runs within one file only, because only one file is compared before and after.
secrets.ts — hard-coded secret detection
Two checks, shared by both languages:
- Known token patterns anywhere in a string: AWS access key (
AKIA…), PEM private key header, GitHub token (ghp_,gho_,ghu_,ghs_,ghr_), Slack token (xox?-), JWT (eyJ….….…),sk-…API keys. - Secret-looking name + credential-looking value: the variable/property
name matches
password,passwd,secret,api_key,access_key,auth_token,token,credential,private_key,client_secret, …; the value is not a placeholder (changeme,your_…,<…>,${…},example,test,xxx, …); and it looks like a credential (no spaces, ≥ 8 safe characters, and letters+digits or ≥ 20 characters).
names.ts — unclear names
isUnclearName(local variables): flags names of ≤ 2 characters unless allowed (id,db,fs,io,ui,ok,el,fn,cb,x,y,z,i,j,k,n,t, …); vague words followed by a number (data1,tmp2,result3, …); and short-letters-plus-digits (abc1).isVagueName(parameters, function and class names): the same without the short-name check, soe,x,cbare allowed.- Always allowed: names starting with
_, and technical names likeutf8,sha256,base64,ipv4,oauth2,i18n,h1–h6.
model.ts — finding → violation
toViolation(finding, config):
severityfromconfig.rules[ruleId].severity(defaultmedium);categoryfromRULE_CATEGORY;weightfrom the defaultSEVERITY_WEIGHTtable (see the note in 08 — Scoring);id= a djb2 hash of rule + path + line + message (unique within a report, not stable across edits; the attribution engine uses the fingerprint instead);source="internal".
Helper functions used by the analyzers:
ruleEnabled(config, id)— true only ifenabled === true. A rule missing from the config is off.ruleLimit(config, id, fallback),ruleOption(config, id, key, fallback).
analyzeRevision.ts — one historical file version
Used by the commit walk. It runs the file's analyzer plus within-file duplicate detection on a blob, converts findings to violations, and adds a fingerprint to each. It returns:
[]for unsupported file types;nullwhen the text is missing or the analyzer throws, so the caller can count affected violations asunattributedinstead of silently dropping them.