perf(explore): keep the stats counts and name list per index state - #2115
Open
professorpalmer wants to merge 1 commit into
Open
professorpalmer wants to merge 1 commit into
professorpalmer wants to merge 1 commit into
Conversation
Every codegraph_explore repeated three pieces of work whose answer depends only on the index: getStats() recounted nodes, edges and files (explore sizes its output from the counts), fuzzy search re-read every distinct node name, and the deprioritize ranking predicate stat-ed codegraph.json once per candidate path, building and throwing an ENOENT error each time when the project has no config file. The stats counts and the name list now go through the change-stamp memo that already holds the dominant file (colbymchenry#1864), generalized into one helper with the same transaction rule; getStats hands each caller its own copy, and a name list above 250k entries is recomputed rather than pinned in memory. The config probe uses statSync's throwIfNoEntry: false. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up to #2086, using the same change stamp for two more whole-graph answers that every
codegraph_explorerecomputed, plus one cheap fix in the ranking predicate.Reproduced first
I use CodeGraph through a long-lived process (one
CodeGraph+ToolHandler, the MCP-server shape) on a ~37k-node Python/TypeScript repo. A CPU profile of 30 warm explores showed three things repeated on every call although their answers depend only on the index:getStats():COUNT(*)s plusGROUP BY kindover nodes and edges. Explore only needsfileCount/nodeCountto pick its output budgetsearchNodesFuzzy→getAllNodeNames():SELECT DISTINCT name FROM nodeson every fuzzy search. The comment there says the list "is cached on QueryBuilder by getAllNodeNames()", but only the statement wasloadDeprioritizePatternsonce per candidate path. With nocodegraph.json(the zero-config default) each call'sstatSyncthrew ENOENT, so every candidate built and caught anErrorstatSyncself timeChange
getDominantFile's memo becomes one private helper,memoizeByChangeStamp(key, compute, keep?), with the same rules: no memo inside a transaction (rollback safety), and cleared onrebind.getDominantFile,getStatsandgetAllNodeNamesuse it.getStatsreturns a fresh copy per call (top level and the three maps), becauseCodeGraph.getStatswritesdbSizeBytes/walSizeBytesinto the result.lastUpdatedis stillDate.now()per call.getAllNodeNamesreturns a frozenreadonly string[]. Both callers (new Set(...)in the resolver, iteration in fuzzy search) only read it. A list above 250k names is recomputed rather than held, so a multi-million-name index (the Linux-kernel case inwarmCaches) doesn't pin that memory in a long-running server.loadParsedConfigstats with{ throwIfNoEntry: false }. Other stat errors are still caught and treated as "no config", same as before. The live-reload behaviour pinned bydeprioritize-config.test.ts("picks up a config written after the project was opened") is unchanged: the file is still checked on every call, it just no longer throws.Results
Same repo, same 6 queries × 3 rounds,
mainand this branch interleaved, three runs each:mainOutput is byte-identical. SHA-1 of the
codegraph_exploretext matchesmainfor 14 calls (7 queries, each twice, including a misspelled one that goes through fuzzy search), andgetStats()counts match.Tests
__tests__/explore-repeat-work-cache.test.ts(real index in a temp dir, real SQLite, no DB mocking):main); each caller gets its own copy.SELECT DISTINCT nameacross repeated reads (fails onmain).sync()through the same connection and through another connection (another process).Transaction and
rebindbehaviour of the shared helper stays covered bydominant-file-cache.test.ts, which passes unchanged. The full engine suite passes locally on macOS: 310 files, 5288 tests passed, 260 skipped.CHANGELOG entry under
[Unreleased]→ New Features, written for users (no internals or numbers).🤖 Generated with Claude Code