09

Caching & performance

The per-commit cache, invalidation, concurrency, speed tips

3 min read5 sections

Where the time goes

StageCostNotes
Git checks, confignegligible
Developer list (git log)smallOne call over the whole history
HEAD scanproportional to code sizeOne pass per file + one duplicate pass
Commit walkdominantPer commit: 1 git diff + 1 git cat-file --batch + analysis of changed files, before and after
Scoringnegligible

The per-commit cache

Implemented in src/cache/analysisCache.ts.

  • Key: commit hash. Value: that commit's CommitContribution (author, date, analyzed lines, introduced records, fixed fingerprints, counts).
  • A commit's content never changes, so a cached result is valid for as long as the rules and settings are the same. Re-runs only analyze new commits.
  • Storage: one JSON file per repository, commits-<hash of repo path>.json:
Front endDirectory
VS Code extensionthe extension's global storage (context.globalStorageUri)
CLI<os.tmpdir()>/cleanlens-cache (e.g. /tmp/cleanlens-cache)
  • The file is written once at the end of a run, and only if something changed.
  • Failures to read or write the cache are ignored. A missing or damaged cache just means a full re-analysis.

Invalidation

The whole cache for a repository is discarded when either value changes:

  1. ENGINE_VERSION (in scoringConfig.ts) — bumped whenever the engine's logic changes.
  2. Analysis config hash = FNV-1a of { engine, rules, sorted exclude list, scoring }. Changing any rule, severity, limit, option, exclude pattern (including .gitignore / .gitattributes-derived ones) or scoring override invalidates it.

Things that do not invalidate the cache (and do not need to): maxCommits, since, developers, mergeSameNameAuthors. They only choose which cached entries are used and how they are grouped.

Turning it off

  • Extension: "cleanlens.cache.enabled": false.
  • CLI: --no-cache.
  • To clear it by hand, delete the commits-*.json file(s) in the directories above.

Built-in optimizations

  • One git diff per commit for all its files, with --unified=0 (small output).
  • One git cat-file --batch per commit for all its blobs, instead of one git show per file.
  • Blob-level memoization: the analysis of a blob is reused within a run, so the "after" of commit N and the "before" of commit N+1 are parsed once.
  • 12 commits in parallel (COMMIT_CONCURRENCY).
  • Skip early: binary, unsupported, excluded, whitespace-only and generated files are never analyzed.
  • Files over 1 MB are skipped in the HEAD scan.
  • Excluded directories are not entered during the file walk.

Tips for large repositories

  1. Keep the default window (500 commits) for day-to-day use. It covers most active work and is fast.
  2. Use analysis.since (e.g. "12 months ago") for a time-boxed review.
  3. Run --full-history occasionally (e.g. a nightly CI job). The cache makes later runs incremental.
  4. Exclude vendored, generated and fixture code. It saves time and makes results fairer.
  5. Keep the cache enabled. After the first run, only new commits cost time.
  6. Avoid changing rules often during an evaluation period: each change forces a full re-analysis.

Source-level knobs

These are constants, not settings. Change them in source and rebuild if you need to:

ConstantFileDefault
COMMIT_CONCURRENCYsrc/attribution/commitAttribution.ts12
MAX_FILE_BYTESsrc/analyzers/fileScanner.ts1,000,000
DEFAULT_MAX_COMMITSsrc/core/analyzeRepository.ts500
MAX_BUFFER (git output)src/git/diffService.ts / commitHistory.ts512 MB / 128 MB
Live diagnostics debouncesrc/views/diagnostics.ts400 ms

Bump ENGINE_VERSION if a change affects the per-commit results, so that old caches are discarded.