Skip to content

perf(context): compute the dominant file once per index state (#1864) - #2086

Merged
colbymchenry merged 4 commits into
mainfrom
carry/1864-dominant-file-memo
Sep 29, 2026
Merged

colbymchenry merged 4 commits into
mainfrom
carry/1864-dominant-file-memo

Conversation

@colbymchenry

Copy link
Copy Markdown
Owner

Carries @danusha2345's #1916, rebased onto current main with their commits and authorship intact, plus one maintainer fix on top. Supersedes #1916.

Reproduced first

getDominantFile() finds the file with the most in-file edges; context ranking then boosts results in that file's directory. The answer doesn't depend on the query, yet on main every generic codegraph_explore recomputes it with a whole-graph join and group-by.

  • Raw SQL: a synthetic 10,000-file TypeScript index (610k nodes, 600k edges) takes a steady ~1 s per call. This repo's own index (73k edges) takes ~40 ms warm.
  • End to end: 4 explores in one process through ToolHandler, the MCP-server shape, with time inside the aggregation measured separately. Machine under load, main and this branch interleaved:
main this PR
generic symbol query the aggregation runs on every call; 765–2,600 ms per call runs once (~670 ms, first call); later calls 90–104 ms
natural-language query runs 4 more times runs 0 times (reused)
exact-file query never reaches it never reaches it

The reporter's A/B also reproduces: with getDominantFile() forced to null on main, the same generic query takes ~92 ms. After its first call, this PR lands at that bypass number.

The contributor's fix

The answer is memoized per QueryBuilder and tagged with an O(1) change stamp: total_changes() for this connection's writes, PRAGMA data_version for any other connection's commits. It isn't kept inside a transaction, because total_changes() counts rows a ROLLBACK undoes. The isTransaction getter it relies on exists in Node 22.20 and in the bundled Node v24.16.0.

Maintainer fix: drop the memo on rebind()

QueryBuilder.rebind() swaps a query-pool worker onto a new connection, for example after the index is rebuilt. The stamp is per connection, and fresh read-only connections to two different databases both report 0:2. So a rebound worker kept serving the old database's dominant file. rebind() now clears the memo. A new test rebinds onto a second project and fails without the reset.

Scope, honestly

This helps every long-lived process: the MCP server and daemon, where agents call codegraph_explore, and codegraph ui. A one-shot process still computes it once. That covers a CLI codegraph explore, which is what the issue's timings measured, and the prompt hook, which opens the index in a fresh process for a structural prompt. Removing that one-per-process cost needs the per-file counts kept up to date by the index itself, a schema and write-path change left for a follow-up.

Verification

  • dominant-file-cache.test.ts: 6 tests covering reuse across explores, a sync on this connection, a sync on another connection, a rolled-back transaction, a rebuilt-and-reopened database, and rebinding onto another database.
  • Full suite on this branch: 320 files, 5,480 passed, 0 failed.

Fixes #1864

Co-authored-by: danusha2345 danusha2345@users.noreply.github.com

🤖 Generated with Claude Code

danusha2345 and others added 4 commits September 28, 2026 23:26
getDominantFile() backs context ranking's core-directory boost. Its answer
depends only on the graph, but the whole-graph aggregation behind it
(edges joined to both endpoints' nodes, grouped by file) ran on every
generic codegraph_explore - seconds per call on a large index.

Memoize it in QueryBuilder against a database change stamp:
total_changes() (rows this connection wrote) plus PRAGMA data_version
(commits by any other connection or process). Both are O(1), so no write
path has to remember to invalidate, and a sync made by another process is
seen on the next call. Any write forces a recompute; the heuristic's
result is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…1864)

total_changes() counts writes that are later rolled back, so a result
computed inside a transaction could survive the ROLLBACK under an unchanged
stamp and describe edges that no longer exist. getDominantFile now skips the
memo while a transaction is open (and on a runtime without isTransaction).
Ported from the maintainer's hardening of this fix in #1919.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…1864)

The memo's change stamp (total_changes + data_version) is per connection,
and fresh read-only connections to two different databases report the same
one ("0:2"). A pool worker rebound onto a rebuilt index kept serving the old
database's dominant file until something wrote through the new connection.
A test rebinds onto a second project and fails without the reset.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Performance: getDominantFile adds ~5s to every generic explore on large repos

2 participants