Skip to content

perf(context): stop redoing per-query work per candidate on long prompts - #2260

Open
danusha2345 wants to merge 2 commits into
colbymchenry:mainfrom
danusha2345:perf/context-long-query
Open

danusha2345 wants to merge 2 commits into
colbymchenry:mainfrom
danusha2345:perf/context-long-query

Conversation

@danusha2345

Copy link
Copy Markdown
Contributor

Refs #2184, follow-up to #2190.

#2190 measured codegraph prompt-hook growing faster than prompt length and noted three hot spots that redo work whose answer cannot change within one retrieval. This PR removes those three and nothing else: ranking is untouched and outputs are byte-identical.

Change

  • Path relevance. scorePathRelevance re-split the whole query and re-ran extractSearchTerms for every candidate path (explore's CamelCase/compound steps score up to ~400 LIKE hits per query symbol against the full prompt). The query half is now prepared once (preparePathRelevanceQuery) and each path scored against it (scorePreparedPathRelevance); scorePathRelevance keeps its signature and delegates. searchNodes prepares once per call instead of once per result.
  • deprioritize config. The predicate stat'ed codegraph.json per candidate. QueryBuilder.setDeprioritizedPathMatcher now takes a source that returns the predicate for the config in force; searchNodes and ContextBuilder.findRelevantContext take one per pass. Edits to codegraph.json are still picked up by the next search (the explore/query relevance: a generic token's exact name-match overboosts in peripheral dirs — follow-up to #746 #982 "config written after open" test passes unchanged).
  • getAllNodeNames. The fuzzy fallback re-read SELECT DISTINCT name FROM nodes once per query term that FTS/LIKE miss. It is now memoized per database change stamp, exactly like getDominantFile (Performance: getDominantFile adds ~5s to every generic explore on large repos #1864): not inside a transaction, dropped on rebind(). The list is held through a WeakRef, so it survives one retrieval but a large index's name list is not pinned in a long-lived server.

Note: QueryBuilder is exported, so two signatures change for library users: setDeprioritizedPathMatcher takes a factory, and getAllNodeNames() returns readonly string[] (the list is shared). If you'd rather keep the old setter, I can add a separate one instead.

Not in this PR

The remaining cost is linear in the number of query terms/symbols (per-term searchNodes, per-symbol substring scans). Capping or deduplicating them changes ranking, so it is left for a separate decision.

Verification

  • Identity: on a 377-file index (codegraph's own src/ + ui/src), codegraph_explore output, buildContext markdown and the findRelevantContext subgraph are byte-identical to main for 18 prompts (short symbol queries; 500 B–8 KB prose; 500 B–8 KB subagent reports), with and without a deprioritize config and a project name. searchNodes id+score lists match on 16 terms including fuzzy typos. The prompt-hook's own stdout matched on all 16 timed prompts.
  • New test __tests__/node-names-cache.test.ts: a 14-term query that falls through to fuzzy reads the name list once (14 times on main), and a sync() is seen.
  • Full suite on Linux passes.
  • Prompt-hook wall time (median of 3, includes ~0.4 s process start):
prompt main this PR
2 KB prose 1.8 s 1.1 s
4 KB prose 2.9 s 1.5 s
8 KB prose 5.7 s 2.4 s
2 KB subagent report 2.2 s 1.2 s
4 KB subagent report 4.6 s 2.2 s
8 KB subagent report 8.6 s 3.0 s

Short symbol queries are unchanged (~0.4 s).

🤖 Generated with Claude Code

danusha2345 and others added 2 commits October 1, 2026 19:34
On a long prompt, explore (and so the Claude Code prompt hook) spent most
of its time repeating work whose answer cannot change within one
retrieval (colbymchenry#2184):

- scorePathRelevance re-split the whole query into words and re-ran
  extractSearchTerms for every candidate path. The query half is now
  prepared once (preparePathRelevanceQuery) and scored per path
  (scorePreparedPathRelevance); scorePathRelevance keeps its signature.
- The deprioritize predicate stat'ed codegraph.json for every candidate.
  The matcher set on QueryBuilder is now a source that returns a
  predicate for the config in force; searchNodes and ContextBuilder take
  one per pass. Config edits are still picked up on the next search.
- getAllNodeNames re-read SELECT DISTINCT name FROM nodes on every fuzzy
  fallback, once per query term. It is now memoized per database change
  stamp, like getDominantFile (colbymchenry#1864), held through a WeakRef so a large
  index's name list is not pinned between retrievals.

Ranking is unchanged: explore, buildContext and findRelevantContext
outputs are byte-identical to main on 18 prompts (short symbol queries,
500 B to 8 KB prose and subagent reports), with and without a
deprioritize config. Prompt-hook wall time on a 377-file index:
4 KB report 4.6 s -> 2.2 s, 8 KB report 8.6 s -> 3.0 s.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The sync's own resolution may read the name list again (any write after that
read moves the change stamp), so the test now checks what the memo promises:
the new name is seen after a sync, and a repeat call in the same state does
not re-read.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
danusha2345 pushed a commit to danusha2345/codegraph that referenced this pull request Oct 1, 2026
…bymchenry#2259, colbymchenry#2260) into fork main

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant