05

Analyzers

File scanning, glob matching, the JS/TS and Python analyzers, duplicate detection, secret and name heuristics

6 min read12 sections

The analyzers turn source text into findings. They live in src/analyzers/. This document covers how they work; 06 — Rules reference covers what each rule checks.

Overview

text
                       ┌─▶ jsAnalyzer.analyzeJs()   (.js .jsx .mjs .cjs .ts .tsx)
file text ─ analyzerFor┤
                       └─▶ pyAnalyzer.analyzePy()   (.py)
                                   │
                                   ▼  RawFinding[]
                     duplicateCode.detectDuplicates()
                                   │
                                   ▼
                         model.toViolation()  →  Violation[]

Every analyzer has the same signature:

ts
(relPath: string, text: string, config: CleanCodeConfig) => RawFinding[]

It is a pure function: no I/O, no Git. That is why the same code runs on disk files (HEAD scan), on historical blobs (commit walk) and on unsaved editor buffers (live diagnostics).

runAnalyzers.ts — the HEAD scan

analyzeProject(root, config, excludeGlobs, onProgress):

  1. scanFiles() collects every file with a supported extension that is not excluded.
  2. For each file:
    • read it as UTF-8 (unreadable → skipped);
    • if excludeGenerated is not false and the header marks the file as generated → skipped and counted in excludedFiles;
    • run the analyzer. A parse error skips that file, not the whole run.
  3. Run detectDuplicates() once across all kept files.
  4. Convert each finding to a Violation and sort by severity (critical → low), then path, then line.

analyzerFor(relPath) chooses the analyzer from the extension, or returns null for unsupported files.

fileScanner.ts — which files are read

  • Recursively walks the repository root.
  • Always skips the directories .git, node_modules, .hg, .svn.
  • Skips a directory if the exclude matcher matches dir or dir/, so an excluded folder is never entered.
  • Keeps a file only if its extension is supported, it is not excluded, and it is ≤ 1,000,000 bytes (MAX_FILE_BYTES). Bigger files are almost always generated or bundled.
  • Paths are normalized to forward slashes on every OS.

glob.ts — the exclude pattern language

CleanLens uses a small glob dialect of its own, not a full .gitignore engine.

SyntaxMeaning
*Any characters except /
**Any characters including / (**/ also matches zero directories)
?One character except /
anything elseLiteral (regex special characters are escaped)

ExcludeMatcher adds two conveniences:

  • Unrooted patterns match at any depth. venv/** becomes **/venv/**, so it also excludes backend/venv/…. Patterns that start with / (rooted) or **/ are used as they are, with the leading / removed.
  • Patterns without / also match the file's base name. *.min.js excludes a/b/c.min.js.

Not supported: negation (!pattern), character classes ([abc]), braces ({a,b}).

excludes.ts — the full exclude list

The final list is built by mergedExcludeGlobs(). Sources are added in this order and nothing ever removes an earlier entry:

  1. Built-in defaults (DEFAULT_EXCLUDE): **/node_modules/**, **/dist/**, **/build/**, **/coverage/**, **/.next/**, **/vendor/**, **/__pycache__/**, **/staticfiles/**, **/generated/**, *.min.js, *.map, *.generated.*.
  2. Migrations **/migrations/**, unless excludeMigrations: false.
  3. Lock files package-lock.json, yarn.lock, pnpm-lock.yaml, unless excludeLockFiles: false. (These are not analyzable anyway; the flag only affects the exclude list and the cache hash.)
  4. linguist-generated paths from .gitattributes, unless excludeGenerated: false. A pattern without / becomes **/<pattern>.
  5. .gitignore entries (best effort): comments and ! negations are skipped; a trailing / or a name without a dot also adds <name>/**.
  6. The config's own exclude list (the preset's list plus your additions, plus the local override's and the VS Code setting's entries).

Generated-file headers

hasGeneratedHeader(text) checks the first 5 lines, case-insensitively, for any of: @generated, generated file, auto-generated, autogenerated, do not edit, code generated. Matching files are skipped in both the HEAD scan and the commit walk (unless excludeGenerated: false).

jsAnalyzer.ts — JavaScript and TypeScript

  • Parses with the TypeScript compiler API (ts.createSourceFile), choosing the script kind from the extension (TS, TSX, JSX, JS). No type checking and no tsconfig are involved; only syntax.
  • Walks the AST once with a depth counter for nesting.
  • "Function-like" nodes: function declarations and expressions, arrow functions, methods, constructors, getters and setters.
  • Cyclomatic complexity starts at 1 and adds 1 for each if, ?:, for, for…in, for…of, while, do, catch, non-empty case, and &&, ||, ??. It does not descend into nested functions; they are counted separately.
  • Symbol names: findings about a function or class carry its name (symbolName). This makes fingerprints stable and enables medium attribution confidence.
  • Usage counts for detectUnusedCode: every identifier in the file is counted once up front; a declaration that appears only once (itself) is unused.

pyAnalyzer.ts — Python

Python is analyzed without a full parser, using lines and indentation:

  • Tabs are expanded to 4 spaces; # comments are stripped (with string awareness).
  • A def / async def / class block ends at the next non-blank line with indentation ≤ the header's.
  • Parameters are read from the signature (across lines up to )), ignoring self, cls, *args, **kwargs.
  • Complexity is approximate: 1 + lines starting with if, elif, for, while, except, case + each and/or in the body.
  • Nesting depth is indent / 4 on lines that start with if, elif, else, for, while, with, try, except. It includes the enclosing def/class levels, so a method body starts at depth 2.
  • A docstring is a first body line starting with """ or '''.

This makes the Python analyzer fast and dependency-free, but less exact than the JS one. See 17 — Limitations.

duplicateCode.ts — duplicate blocks

  1. Normalize each line: trim and collapse whitespace. Drop lines that are: shorter than 12 characters, only brackets/punctuation, comments (//, #, *), or import / from / export lines.
  2. Slide a window of minLines (option, default 6, minimum 3) kept lines over each file. A window is skipped if its lines are spread over more than 3 × minLines original lines (not contiguous enough).
  3. The first time a window's text is seen, it is recorded. When it is seen again (same file or another file), the later location is a hit.
  4. Overlapping or adjacent hits are merged into one finding per duplicated block. A 40-line clone is one violation, not 35.

Notes:

  • The first occurrence is never flagged; only the copies found later in scan order are.
  • In the HEAD scan, detection runs across all files. In the commit walk (analyzeRevision), it runs within one file only, because only one file is compared before and after.

secrets.ts — hard-coded secret detection

Two checks, shared by both languages:

  • Known token patterns anywhere in a string: AWS access key (AKIA…), PEM private key header, GitHub token (ghp_, gho_, ghu_, ghs_, ghr_), Slack token (xox?-), JWT (eyJ….….…), sk-… API keys.
  • Secret-looking name + credential-looking value: the variable/property name matches password, passwd, secret, api_key, access_key, auth_token, token, credential, private_key, client_secret, …; the value is not a placeholder (changeme, your_…, <…>, ${…}, example, test, xxx, …); and it looks like a credential (no spaces, ≥ 8 safe characters, and letters+digits or ≥ 20 characters).

names.ts — unclear names

  • isUnclearName (local variables): flags names of ≤ 2 characters unless allowed (id, db, fs, io, ui, ok, el, fn, cb, x, y, z, i, j, k, n, t, …); vague words followed by a number (data1, tmp2, result3, …); and short-letters-plus-digits (abc1).
  • isVagueName (parameters, function and class names): the same without the short-name check, so e, x, cb are allowed.
  • Always allowed: names starting with _, and technical names like utf8, sha256, base64, ipv4, oauth2, i18n, h1–h6.

model.ts — finding → violation

toViolation(finding, config):

  • severity from config.rules[ruleId].severity (default medium);
  • category from RULE_CATEGORY;
  • weight from the default SEVERITY_WEIGHT table (see the note in 08 — Scoring);
  • id = a djb2 hash of rule + path + line + message (unique within a report, not stable across edits; the attribution engine uses the fingerprint instead);
  • source = "internal".

Helper functions used by the analyzers:

  • ruleEnabled(config, id) — true only if enabled === true. A rule missing from the config is off.
  • ruleLimit(config, id, fallback), ruleOption(config, id, key, fallback).

analyzeRevision.ts — one historical file version

Used by the commit walk. It runs the file's analyzer plus within-file duplicate detection on a blob, converts findings to violations, and adds a fingerprint to each. It returns:

  • [] for unsupported file types;
  • null when the text is missing or the analyzer throws, so the caller can count affected violations as unattributed instead of silently dropping them.