Skip to content

perf(explore): keep the stats counts and name list per index state - #2115

Open
professorpalmer wants to merge 1 commit into
colbymchenry:mainfrom
professorpalmer:perf/explore-repeat-work
Open

professorpalmer wants to merge 1 commit into
colbymchenry:mainfrom
professorpalmer:perf/explore-repeat-work

Conversation

@professorpalmer

Copy link
Copy Markdown

Follow-up to #2086, using the same change stamp for two more whole-graph answers that every codegraph_explore recomputed, plus one cheap fix in the ranking predicate.

Reproduced first

I use CodeGraph through a long-lived process (one CodeGraph + ToolHandler, the MCP-server shape) on a ~37k-node Python/TypeScript repo. A CPU profile of 30 warm explores showed three things repeated on every call although their answers depend only on the index:

Work Per explore (before)
getStats(): COUNT(*)s plus GROUP BY kind over nodes and edges. Explore only needs fileCount / nodeCount to pick its output budget ~5.8 ms
searchNodesFuzzy → getAllNodeNames(): SELECT DISTINCT name FROM nodes on every fuzzy search. The comment there says the list "is cached on QueryBuilder by getAllNodeNames()", but only the statement was ~4.7 ms
The deprioritize matcher calls loadDeprioritizePatterns once per candidate path. With no codegraph.json (the zero-config default) each call's statSync threw ENOENT, so every candidate built and caught an Error ~8 ms of statSync self time

Change

  • getDominantFile's memo becomes one private helper, memoizeByChangeStamp(key, compute, keep?), with the same rules: no memo inside a transaction (rollback safety), and cleared on rebind. getDominantFile, getStats and getAllNodeNames use it.
  • getStats returns a fresh copy per call (top level and the three maps), because CodeGraph.getStats writes dbSizeBytes / walSizeBytes into the result. lastUpdated is still Date.now() per call.
  • getAllNodeNames returns a frozen readonly string[]. Both callers (new Set(...) in the resolver, iteration in fuzzy search) only read it. A list above 250k names is recomputed rather than held, so a multi-million-name index (the Linux-kernel case in warmCaches) doesn't pin that memory in a long-running server.
  • loadParsedConfig stats with { throwIfNoEntry: false }. Other stat errors are still caught and treated as "no config", same as before. The live-reload behaviour pinned by deprioritize-config.test.ts ("picks up a config written after the project was opened") is unchanged: the file is still checked on every call, it just no longer throws.

Results

Same repo, same 6 queries × 3 rounds, main and this branch interleaved, three runs each:

main this PR
explore p50 83 / 84 / 87 ms 67 / 67 / 71 ms
explore p90 108 / 114 / 121 ms 87 / 81 / 90 ms

Output is byte-identical. SHA-1 of the codegraph_explore text matches main for 14 calls (7 queries, each twice, including a misspelled one that goes through fuzzy search), and getStats() counts match.

Tests

__tests__/explore-repeat-work-cache.test.ts (real index in a temp dir, real SQLite, no DB mocking):

  • stats: counted once across several explores while unchanged (fails on main); each caller gets its own copy.
  • names: one SELECT DISTINCT name across repeated reads (fails on main).
  • freshness: new counts and names after sync() through the same connection and through another connection (another process).

Transaction and rebind behaviour of the shared helper stays covered by dominant-file-cache.test.ts, which passes unchanged. The full engine suite passes locally on macOS: 310 files, 5288 tests passed, 260 skipped.

CHANGELOG entry under [Unreleased] → New Features, written for users (no internals or numbers).

🤖 Generated with Claude Code

Every codegraph_explore repeated three pieces of work whose answer depends
only on the index: getStats() recounted nodes, edges and files (explore
sizes its output from the counts), fuzzy search re-read every distinct node
name, and the deprioritize ranking predicate stat-ed codegraph.json once per
candidate path, building and throwing an ENOENT error each time when the
project has no config file.

The stats counts and the name list now go through the change-stamp memo
that already holds the dominant file (colbymchenry#1864), generalized into one helper
with the same transaction rule; getStats hands each caller its own copy, and
a name list above 250k entries is recomputed rather than pinned in memory.
The config probe uses statSync's throwIfNoEntry: false.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant