Skip to content

perf(context): share project-wide aggregates across processes - #2349

Draft
gtarpenning wants to merge 1 commit into
colbymchenry:mainfrom
gtarpenning:griffin/cross-process-aggregate-memo
Draft

gtarpenning wants to merge 1 commit into
colbymchenry:mainfrom
gtarpenning:griffin/cross-process-aggregate-memo

Conversation

@gtarpenning

Copy link
Copy Markdown

Summary

getDominantFile() and getStats() aggregate the whole graph, and neither depends on the query. The prompt hook is a fresh process per prompt, so the in-process memo from #1864 never helps it: every prompt recomputes both (about 1.3 s warm, more cold). This keeps the results where a new process can read them, and refreshes them when the graph changes.

Independent of the trigram PR; either can merge first. #2115 (an in-process getStats cache) and #2260 touch nearby code in queries.ts; this stays minimal around them.

What changed

  • graph_epoch is a project_metadata row replaced whenever the graph changes: indexAll, sync (unless it changed nothing), indexFiles, resolveReferences, resolveReferencesBatched and clear. No schema change.
  • Sidecar memo (<db>.memo.json) holds both results, tagged with the epoch. A tag mismatch, a missing file or unreadable JSON recomputes and rewrites it. It is a file, not a metadata row, because the hook's connection would otherwise need the database write lock the daemon's writer holds.
  • Warm after writes: indexAll recomputes inline; other writers do it after 15 s of quiet on an unref'd timer, so a burst of edits pays once.
  • indexFiles now awaits orchestrator.indexFiles inside its try, so the file lock is held until indexing finishes instead of being released as the promise is created.
  • Reads are skipped inside a transaction and on an index with no epoch yet; any sync seeds one.

Measured

Prompt hook on a large TS/Go/Python monorepo index (about 1.5 GB, 350k nodes, 1.2M edges), same JSON on stdin:

Wall time
Before 3.2-3.3 s
With this PR 1.6 s
With this PR and the trigram PR about 1.1 s

The first run after a real change pays the recompute once. Hook stdout is byte-identical to the unmodified build on every run.

Compatibility

  • A writer without this change (an older install on a shared index) does not bump the epoch, so the memo can be stale until a newer build writes.
  • One-shot codegraph sync exits before the deferred warm; the next hook recomputes once.

Testing

  • npx tsc --noEmit
  • npx vitest run __tests__/persisted-memo.test.ts __tests__/dominant-file-cache.test.ts __tests__/sync.test.ts: a stale tag and a corrupt sidecar both recompute and repair; a fresh process sees the new graph after indexFiles and clear; the deferred warm populates the sidecar.
  • npx vitest run: 5747 passed, 23 failed. 20 are in daemon-older-version, daemon-registry and index-command and fail the same way on unmodified main; 3 are git-index-currency timeouts under load that pass in isolation on both.

🤖 Generated with Claude Code

getDominantFile and getStats scan the whole graph and were recomputed by every
fresh process, such as the prompt hook. They are now kept in a sidecar file
tagged with a graph_epoch bumped by every graph writer, and warmed after an
index or sync.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant