perf(context): stop redoing per-query work per candidate on long prompts - #2260
Open
danusha2345 wants to merge 2 commits into
Open
danusha2345 wants to merge 2 commits into
danusha2345 wants to merge 2 commits into
Conversation
On a long prompt, explore (and so the Claude Code prompt hook) spent most of its time repeating work whose answer cannot change within one retrieval (colbymchenry#2184): - scorePathRelevance re-split the whole query into words and re-ran extractSearchTerms for every candidate path. The query half is now prepared once (preparePathRelevanceQuery) and scored per path (scorePreparedPathRelevance); scorePathRelevance keeps its signature. - The deprioritize predicate stat'ed codegraph.json for every candidate. The matcher set on QueryBuilder is now a source that returns a predicate for the config in force; searchNodes and ContextBuilder take one per pass. Config edits are still picked up on the next search. - getAllNodeNames re-read SELECT DISTINCT name FROM nodes on every fuzzy fallback, once per query term. It is now memoized per database change stamp, like getDominantFile (colbymchenry#1864), held through a WeakRef so a large index's name list is not pinned between retrievals. Ranking is unchanged: explore, buildContext and findRelevantContext outputs are byte-identical to main on 18 prompts (short symbol queries, 500 B to 8 KB prose and subagent reports), with and without a deprioritize config. Prompt-hook wall time on a 377-file index: 4 KB report 4.6 s -> 2.2 s, 8 KB report 8.6 s -> 3.0 s. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The sync's own resolution may read the name list again (any write after that read moves the change stamp), so the test now checks what the memo promises: the new name is seen after a sync, and a repeat call in the same state does not re-read. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
danusha2345
pushed a commit
to danusha2345/codegraph
that referenced
this pull request
Oct 1, 2026
…bymchenry#2259, colbymchenry#2260) into fork main Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 5, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Refs #2184, follow-up to #2190.
#2190 measured
codegraph prompt-hookgrowing faster than prompt length and noted three hot spots that redo work whose answer cannot change within one retrieval. This PR removes those three and nothing else: ranking is untouched and outputs are byte-identical.Change
scorePathRelevancere-split the whole query and re-ranextractSearchTermsfor every candidate path (explore's CamelCase/compound steps score up to ~400 LIKE hits per query symbol against the full prompt). The query half is now prepared once (preparePathRelevanceQuery) and each path scored against it (scorePreparedPathRelevance);scorePathRelevancekeeps its signature and delegates.searchNodesprepares once per call instead of once per result.deprioritizeconfig. The predicate stat'edcodegraph.jsonper candidate.QueryBuilder.setDeprioritizedPathMatchernow takes a source that returns the predicate for the config in force;searchNodesandContextBuilder.findRelevantContexttake one per pass. Edits tocodegraph.jsonare still picked up by the next search (the explore/query relevance: a generic token's exact name-match overboosts in peripheral dirs — follow-up to #746 #982 "config written after open" test passes unchanged).getAllNodeNames. The fuzzy fallback re-readSELECT DISTINCT name FROM nodesonce per query term that FTS/LIKE miss. It is now memoized per database change stamp, exactly likegetDominantFile(Performance: getDominantFile adds ~5s to every generic explore on large repos #1864): not inside a transaction, dropped onrebind(). The list is held through aWeakRef, so it survives one retrieval but a large index's name list is not pinned in a long-lived server.Note:
QueryBuilderis exported, so two signatures change for library users:setDeprioritizedPathMatchertakes a factory, andgetAllNodeNames()returnsreadonly string[](the list is shared). If you'd rather keep the old setter, I can add a separate one instead.Not in this PR
The remaining cost is linear in the number of query terms/symbols (per-term
searchNodes, per-symbol substring scans). Capping or deduplicating them changes ranking, so it is left for a separate decision.Verification
src/+ui/src),codegraph_exploreoutput,buildContextmarkdown and thefindRelevantContextsubgraph are byte-identical tomainfor 18 prompts (short symbol queries; 500 B–8 KB prose; 500 B–8 KB subagent reports), with and without adeprioritizeconfig and a project name.searchNodesid+score lists match on 16 terms including fuzzy typos. The prompt-hook's own stdout matched on all 16 timed prompts.__tests__/node-names-cache.test.ts: a 14-term query that falls through to fuzzy reads the name list once (14 times onmain), and async()is seen.Short symbol queries are unchanged (~0.4 s).
🤖 Generated with Claude Code