Conversation
ee7f586 to
05165b8
Compare
|
Hi @QingNagi This proposal looks really interesting. I like the idea of treating Markdown documents as part of the same graph context, especially because many projects keep important architecture notes, runbooks, ADRs, setup guides, and agent instructions in Markdown files. One question/suggestion: would it make sense to also include a small document or section-level summary as part of the indexed metadata? For example, besides indexing headings, table rows, links, and references, CodeGraph could optionally store a lightweight summary derived from the document structure, such as: a short summary of what the Markdown file is about; The main benefit would be helping agents quickly decide whether a Markdown file is relevant before loading more context. It could also improve search results by making documentation easier to discover when the exact heading or filename is not known. I understand that full LLM-generated summaries may be outside the scope or introduce extra complexity, but even a deterministic summary based on headings, links, tables, and references could be useful. |
|
@QingNagi what is the status of this? |
Add Markdown extraction for headings, table rows, command templates, and file-symbol references. Resolve Markdown anchors and file-symbol references, plus code string references back to Markdown files/headings. Cover Markdown extraction/resolution and full-pipeline md/code graph edges with tests.
05165b8 to
f69718a
Compare
Give the Markdown file node a deterministic, LLM-free digest in its docstring: a one-line intro (what the doc is about) plus the key files/symbols it references (basename + ::symbol/#anchor), derived purely from the structure already extracted. Because docstrings are in the FTS index and surface in node details, this lets an agent judge a doc's relevance from search results before loading it, and makes a doc discoverable by the symbols it documents even when the query matches no heading or filename. Responds to the PR colbymchenry#361 review thread. Also make heading extraction fenced-code- and YAML-frontmatter-aware and add Setext (===/---) underline headings. This fixes #-in-code-fence false positives, adds more anchor targets, and gives the digest a cleaner outline. Covered by new extraction tests for both behaviours. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
@Alpha018 Great suggestion — and I agree it lands right on the main value of indexing docs: helping an agent decide whether a Markdown file is relevant before loading it. I've implemented a deterministic version of this (no LLM) in the latest commit (f296bb5). Mapping to your four points: Short summary of what the file is about — the file node's docstring is now a digest: a one-line intro (first real prose line, skipping frontmatter/badges/headings) instead of the raw first lines. I deliberately kept LLM-generated summaries out of scope for the reasons you noted (cost, non-determinism, indexing-time complexity) — everything here is derived from the document structure and is byte-stable. While in there I also made heading extraction fenced-code/frontmatter-aware and added Setext (===/---) headings, which gives the outline (and the digest) better coverage. Happy to tune the digest format (field order, length budget, code-vs-doc ref prioritisation) if you have a preference. |
Thanks for checking in! Honest status: this is running in my own personal setup, and for my day-to-day it does exactly what I wanted — Markdown docs (runbooks, ADRs, agent instructions) become searchable and link into the code graph, so retrieval over my .md files works well. Local speed/efficiency test — I benchmarked it against a real Markdown corpus I use (~25 files, ~367 KB / ~4k lines):
That ~99% context cut (plus turning “which doc covers X / what does this doc point at” into a single graph hop instead of grep-then-open) is where the real speed-up comes from for doc retrieval. Next directions I’d like input on:
Suggestions and different priorities very welcome — especially on the digest format and how deep to index doc structure. Happy to split any of these into separate PRs. |
|
@colbymchenry can this get reviewed? i think this is an important capability |
Resolve conflict in src/extraction/tree-sitter.ts field fallback: combine upstream's fieldKind (constant vs field kind for Java/C# const fields) with this branch's fieldNode capture + extractMarkdownPathReferencesFromSubtree call, matching the sibling declarator branch. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
+1 on this |
|
Where are we with this? @colbymchenry @QingNagi |
|
We will never get this into codegraph |
Resolve PR colbymchenry#361 conflicts against upstream main f1ca991 while preserving the Markdown extractor and the latest extraction-kernel changes.
|
Resolved the merge conflicts. @colbymchenry |
|
Let's get this merged in, this would be a great feature! |
Ports QingNagi/codegraph#361 (markdown extractor, heading nodes, name-matcher and resolution hooks) onto experimental and adds the section-first doc tier from feature/md-section-first: a doc-shaped query renders the best headed sections of the markdown file it names, ranked by idf-weighted line hits with path tokens weighted zero, capped at DOC_FILE_CAP per file. Markdown reaches an answer only through that tier, generated-file detection ignores markdown bodies, and the budget tiers count code files only so a README-heavy repo keeps its code answers.
|
Built this branch (3a73fed, 1.4.1 base) and measured it under headless Claude Code (Opus, As merged, the index would not get used. 24 cells over the doc bank: 0
Section-first doc answers fix the second. On top of this branch I added a doc tier in
Every explore call in all 36 cells chose the right file and section. The misses are the model's habit, not retrieval: one prompt about "one-shot reminders" is read as a question about the session's own scheduling tools in 0 of 6 cells under this build, and the other misses are a Grep on a file the prompt already names. Cost per cell was $0.30 against $0.54 on the shipped build. Porting to 1.6.0 needed three more changes, which apply to this PR as well:
The port is on my fork's |
Ports QingNagi/codegraph#361 (markdown extractor, heading nodes, name-matcher and resolution hooks) onto experimental and adds the section-first doc tier from feature/md-section-first: a doc-shaped query renders the best headed sections of the markdown file it names, ranked by idf-weighted line hits with path tokens weighted zero, capped at DOC_FILE_CAP per file. Markdown reaches an answer only through that tier, generated-file detection ignores markdown bodies, and the budget tiers count code files only so a README-heavy repo keeps its code answers.
Ports QingNagi/codegraph#361 (markdown extractor, heading nodes, name-matcher and resolution hooks) onto experimental and adds the section-first doc tier from feature/md-section-first: a doc-shaped query renders the best headed sections of the markdown file it names, ranked by idf-weighted line hits with path tokens weighted zero, capped at DOC_FILE_CAP per file. Markdown reaches an answer only through that tier, generated-file detection ignores markdown bodies, and the budget tiers count code files only so a README-heavy repo keeps its code answers.
Ports QingNagi/codegraph#361 (markdown extractor, heading nodes, name-matcher and resolution hooks) onto experimental and adds the section-first doc tier from feature/md-section-first: a doc-shaped query renders the best headed sections of the markdown file it names, ranked by idf-weighted line hits with path tokens weighted zero, capped at DOC_FILE_CAP per file. Markdown reaches an answer only through that tier, generated-file detection ignores markdown bodies, and the budget tiers count code files only so a README-heavy repo keeps its code answers.
|
Great work bompus! there is something we can do in order to push further this pr? |
Ports QingNagi/codegraph#361 (markdown extractor, heading nodes, name-matcher and resolution hooks) onto experimental and adds the section-first doc tier from feature/md-section-first: a doc-shaped query renders the best headed sections of the markdown file it names, ranked by idf-weighted line hits with path tokens weighted zero, capped at DOC_FILE_CAP per file. Markdown reaches an answer only through that tier, generated-file detection ignores markdown bodies, and the budget tiers count code files only so a README-heavy repo keeps its code answers.
Why this feature is needed
Agent and skill workflows often mix Markdown instructions, phase documents, script templates, and implementation scripts. Important logic is not always in code; it is frequently defined in
SKILL.md, runbooks, checklists, tables, and workflow docs.This change extends CodeGraph’s Markdown indexing so Markdown files can participate in the same graph-based lookup flow as source code.
Markdown files are now indexed structurally: headings become searchable section nodes, stable table rows such as
API-AUTHbecome searchable nodes, and Markdown references to files, scripts, command templates, andfile::symboltargets are extracted as graph edges.It also adds reverse mapping from code back to Markdown. For example, a Python function that opens
docs/setup.md#database-setupnow creates a graph edge from that function to the Markdown heading node. This lets agents move both ways: from docs to implementation, and from scripts back to the documentation rules or templates they depend on.Main benefits:
Example graph relationships:
Validation covered Markdown extraction, file-symbol resolution, Markdown anchor resolution, and full-pipeline doc/code graph edges.
Infrastructure
src/types.ts— added markdown to the Language unionsrc/extraction/grammars.ts— added .md, .mdx, and .markdown extension mapping; marked Markdown as a supported custom extractor languagesrc/extraction/markdown-extractor.ts— new lightweight Markdown extractor for headings, table rows, links, command templates, and file::symbol referencessrc/extraction/tree-sitter.ts— added Markdown extractor dispatch; added code-string scanning for Markdown path references such as docs/setup.md#database-setupsrc/resolution/name-matcher.ts— added resolution for Markdown anchors and file::symbol references into target filessrc/resolution/index.ts— updated fast pre-filtering so path-like references with #anchor can still resolve correctlyTest updates
__tests__/extraction.test.ts— added Markdown language detection, heading/link extraction, table row extraction, file-symbol extraction, and code-to-Markdown reference tests__tests__/resolution.test.ts— added Markdown file, anchor, and file::symbol resolution tests__tests__/integration/full-pipeline.test.ts— added full-pipeline tests for Markdown-to-script/function edges and code-to-Markdown heading edges