Caching & performance
The per-commit cache, invalidation, concurrency, speed tips
Where the time goes
| Stage | Cost | Notes |
|---|---|---|
| Git checks, config | negligible | |
Developer list (git log) | small | One call over the whole history |
| HEAD scan | proportional to code size | One pass per file + one duplicate pass |
| Commit walk | dominant | Per commit: 1 git diff + 1 git cat-file --batch + analysis of changed files, before and after |
| Scoring | negligible |
The per-commit cache
Implemented in src/cache/analysisCache.ts.
- Key: commit hash. Value: that commit's
CommitContribution(author, date, analyzed lines, introduced records, fixed fingerprints, counts). - A commit's content never changes, so a cached result is valid for as long as the rules and settings are the same. Re-runs only analyze new commits.
- Storage: one JSON file per repository,
commits-<hash of repo path>.json:
| Front end | Directory |
|---|---|
| VS Code extension | the extension's global storage (context.globalStorageUri) |
| CLI | <os.tmpdir()>/cleanlens-cache (e.g. /tmp/cleanlens-cache) |
- The file is written once at the end of a run, and only if something changed.
- Failures to read or write the cache are ignored. A missing or damaged cache just means a full re-analysis.
Invalidation
The whole cache for a repository is discarded when either value changes:
ENGINE_VERSION(inscoringConfig.ts) — bumped whenever the engine's logic changes.- Analysis config hash = FNV-1a of
{ engine, rules, sorted exclude list, scoring }. Changing any rule, severity, limit, option, exclude pattern (including.gitignore/.gitattributes-derived ones) or scoring override invalidates it.
Things that do not invalidate the cache (and do not need to): maxCommits,
since, developers, mergeSameNameAuthors. They only choose which cached
entries are used and how they are grouped.
Turning it off
- Extension:
"cleanlens.cache.enabled": false. - CLI:
--no-cache. - To clear it by hand, delete the
commits-*.jsonfile(s) in the directories above.
Built-in optimizations
- One
git diffper commit for all its files, with--unified=0(small output). - One
git cat-file --batchper commit for all its blobs, instead of onegit showper file. - Blob-level memoization: the analysis of a blob is reused within a run, so the "after" of commit N and the "before" of commit N+1 are parsed once.
- 12 commits in parallel (
COMMIT_CONCURRENCY). - Skip early: binary, unsupported, excluded, whitespace-only and generated files are never analyzed.
- Files over 1 MB are skipped in the HEAD scan.
- Excluded directories are not entered during the file walk.
Tips for large repositories
- Keep the default window (500 commits) for day-to-day use. It covers most active work and is fast.
- Use
analysis.since(e.g."12 months ago") for a time-boxed review. - Run
--full-historyoccasionally (e.g. a nightly CI job). The cache makes later runs incremental. - Exclude vendored, generated and fixture code. It saves time and makes results fairer.
- Keep the cache enabled. After the first run, only new commits cost time.
- Avoid changing rules often during an evaluation period: each change forces a full re-analysis.
Source-level knobs
These are constants, not settings. Change them in source and rebuild if you need to:
| Constant | File | Default |
|---|---|---|
COMMIT_CONCURRENCY | src/attribution/commitAttribution.ts | 12 |
MAX_FILE_BYTES | src/analyzers/fileScanner.ts | 1,000,000 |
DEFAULT_MAX_COMMITS | src/core/analyzeRepository.ts | 500 |
MAX_BUFFER (git output) | src/git/diffService.ts / commitHistory.ts | 512 MB / 128 MB |
| Live diagnostics debounce | src/views/diagnostics.ts | 400 ms |
Bump ENGINE_VERSION if a change affects the per-commit results, so that old
caches are discarded.