Skip to content

feat(markdown): index documentation graph links - #361

Open
QingNagi wants to merge 5 commits into
colbymchenry:mainfrom
QingNagi:feature/md-index-graph-links
Open

QingNagi wants to merge 5 commits into
colbymchenry:mainfrom
QingNagi:feature/md-index-graph-links

Conversation

@QingNagi

@QingNagi QingNagi commented May 23, 2026 •

Copy link
Copy Markdown

Why this feature is needed

Agent and skill workflows often mix Markdown instructions, phase documents, script templates, and implementation scripts. Important logic is not always in code; it is frequently defined in SKILL.md, runbooks, checklists, tables, and workflow docs.

This change extends CodeGraph’s Markdown indexing so Markdown files can participate in the same graph-based lookup flow as source code.

Markdown files are now indexed structurally: headings become searchable section nodes, stable table rows such as API-AUTH become searchable nodes, and Markdown references to files, scripts, command templates, and file::symbol targets are extracted as graph edges.

It also adds reverse mapping from code back to Markdown. For example, a Python function that opens docs/setup.md#database-setup now creates a graph edge from that function to the Markdown heading node. This lets agents move both ways: from docs to implementation, and from scripts back to the documentation rules or templates they depend on.

Main benefits:

  • Direct lookup of Markdown headings and table rows.
  • Fewer grep and full-file reads when locating documentation rules.
  • Faster discovery of scripts/functions referenced by Markdown.
  • Easier reverse tracing from code to related Markdown files or sections.
  • Better validation that documentation and implementation stay aligned.

Example graph relationships:

API-AUTH          -> src/auth.ts::loginUser
API-AUTH          -> src/auth.ts::refreshToken
initializeDatabase -> docs/setup.md#database-setup
SETUP_DOC         -> docs/setup.md

Validation covered Markdown extraction, file-symbol resolution, Markdown anchor resolution, and full-pipeline doc/code graph edges.

Infrastructure
src/types.ts — added markdown to the Language union
src/extraction/grammars.ts — added .md, .mdx, and .markdown extension mapping; marked Markdown as a supported custom extractor language
src/extraction/markdown-extractor.ts — new lightweight Markdown extractor for headings, table rows, links, command templates, and file::symbol references
src/extraction/tree-sitter.ts — added Markdown extractor dispatch; added code-string scanning for Markdown path references such as docs/setup.md#database-setup
src/resolution/name-matcher.ts — added resolution for Markdown anchors and file::symbol references into target files
src/resolution/index.ts — updated fast pre-filtering so path-like references with #anchor can still resolve correctly

Test updates
__tests__/extraction.test.ts — added Markdown language detection, heading/link extraction, table row extraction, file-symbol extraction, and code-to-Markdown reference tests
__tests__/resolution.test.ts — added Markdown file, anchor, and file::symbol resolution tests
__tests__/integration/full-pipeline.test.ts — added full-pipeline tests for Markdown-to-script/function edges and code-to-Markdown heading edges

@QingNagi
QingNagi force-pushed the feature/md-index-graph-links branch from ee7f586 to 05165b8 Compare May 25, 2026 16:03
@Alpha018

Copy link
Copy Markdown

Hi @QingNagi

This proposal looks really interesting. I like the idea of treating Markdown documents as part of the same graph context, especially because many projects keep important architecture notes, runbooks, ADRs, setup guides, and agent instructions in Markdown files.

One question/suggestion: would it make sense to also include a small document or section-level summary as part of the indexed metadata?

For example, besides indexing headings, table rows, links, and references, CodeGraph could optionally store a lightweight summary derived from the document structure, such as:

a short summary of what the Markdown file is about;
a short summary per major heading;
key referenced files/symbols found in the document;
a quick “reference map” showing which docs point to which code symbols or other docs.

The main benefit would be helping agents quickly decide whether a Markdown file is relevant before loading more context. It could also improve search results by making documentation easier to discover when the exact heading or filename is not known.

I understand that full LLM-generated summaries may be outside the scope or introduce extra complexity, but even a deterministic summary based on headings, links, tables, and references could be useful.

@elpapi42

Copy link
Copy Markdown

@QingNagi what is the status of this?

Add Markdown extraction for headings, table rows, command templates, and file-symbol references.

Resolve Markdown anchors and file-symbol references, plus code string references back to Markdown files/headings.

Cover Markdown extraction/resolution and full-pipeline md/code graph edges with tests.
@QingNagi
QingNagi force-pushed the feature/md-index-graph-links branch from 05165b8 to f69718a Compare June 14, 2026 04:31
Give the Markdown file node a deterministic, LLM-free digest in its
docstring: a one-line intro (what the doc is about) plus the key
files/symbols it references (basename + ::symbol/#anchor), derived
purely from the structure already extracted. Because docstrings are in
the FTS index and surface in node details, this lets an agent judge a
doc's relevance from search results before loading it, and makes a doc
discoverable by the symbols it documents even when the query matches no
heading or filename. Responds to the PR colbymchenry#361 review thread.

Also make heading extraction fenced-code- and YAML-frontmatter-aware
and add Setext (===/---) underline headings. This fixes #-in-code-fence
false positives, adds more anchor targets, and gives the digest a
cleaner outline.

Covered by new extraction tests for both behaviours.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@QingNagi

Copy link
Copy Markdown
Author

@Alpha018 Great suggestion — and I agree it lands right on the main value of indexing docs: helping an agent decide whether a Markdown file is relevant before loading it. I've implemented a deterministic version of this (no LLM) in the latest commit (f296bb5).

Mapping to your four points:

Short summary of what the file is about — the file node's docstring is now a digest: a one-line intro (first real prose line, skipping frontmatter/badges/headings) instead of the raw first lines.
Key referenced files/symbols — the digest appends refs: ::symbol / #anchor, …, de-duplicated and derived from the references already extracted.
Per-heading summary — this already exists: each heading node carries a docstring built from its section body.
Reference map (which docs point to which symbols/docs) — that's exactly the resolved imports/references/calls edges in the graph; the per-file refs: line is a compact, query-time-free slice of it.
Two reasons this is more than cosmetic: docstring is in the FTS index, so a doc now becomes discoverable by the symbols it documents even when the query matches no heading or filename (your "exact heading unknown" case); and it's kept under the ~200-char threshold so it actually surfaces in node details for fast triage.

I deliberately kept LLM-generated summaries out of scope for the reasons you noted (cost, non-determinism, indexing-time complexity) — everything here is derived from the document structure and is byte-stable.

While in there I also made heading extraction fenced-code/frontmatter-aware and added Setext (===/---) headings, which gives the outline (and the digest) better coverage. Happy to tune the digest format (field order, length budget, code-vs-doc ref prioritisation) if you have a preference.

@QingNagi

Copy link
Copy Markdown
Author

@QingNagi what is the status of this?

Thanks for checking in! Honest status: this is running in my own personal setup, and for my day-to-day it does exactly what I wanted — Markdown docs (runbooks, ADRs, agent instructions) become searchable and link into the code graph, so retrieval over my .md files works well.

Local speed/efficiency test — I benchmarked it against a real Markdown corpus I use (~25 files, ~367 KB / ~4k lines):

  • The whole corpus extracts/indexes in ~35 ms (~10 MB/s, ~1.4 ms per file) — indexing cost is negligible.
  • It yields ~1,140 graph nodes (headings, table/list items, commands, links) and ~250 cross-references into code.
  • Every file gets a deterministic digest, avg ~186 chars. So to decide whether a doc is relevant, an agent reads the ~186-char digest instead of opening the ~15 KB file — about 99% less content per relevance decision, before any deeper read.

That ~99% context cut (plus turning “which doc covers X / what does this doc point at” into a single graph hop instead of grep-then-open) is where the real speed-up comes from for doc retrieval.

Next directions I’d like input on:

  1. Cross-dir / multi-repo links — ../foo/bar.md is currently dropped; resolve it against the workspace root.
  2. Single-pass extraction — collapse the current multi-pass scan + duplicated fence handling into one tokenizer (faster, removes a correctness footgun).
  3. A first-class “reference map” query, building on the per-file refs: digest.

Suggestions and different priorities very welcome — especially on the digest format and how deep to index doc structure. Happy to split any of these into separate PRs.

@stack-stitch-admin

Copy link
Copy Markdown

@colbymchenry can this get reviewed? i think this is an important capability

Resolve conflict in src/extraction/tree-sitter.ts field fallback: combine
upstream's fieldKind (constant vs field kind for Java/C# const fields) with
this branch's fieldNode capture + extractMarkdownPathReferencesFromSubtree
call, matching the sibling declarator branch.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@fleps

fleps commented Jul 2, 2026

Copy link
Copy Markdown

+1 on this

@elpapi42

elpapi42 commented Jul 2, 2026 •

Copy link
Copy Markdown

Where are we with this? @colbymchenry @QingNagi

@elpapi42

Copy link
Copy Markdown

We will never get this into codegraph

Resolve PR colbymchenry#361 conflicts against upstream main f1ca991 while preserving the Markdown extractor and the latest extraction-kernel changes.
@QingNagi

Copy link
Copy Markdown
Author

Resolved the merge conflicts. @colbymchenry

@AlexG-UltraCamp

Copy link
Copy Markdown

Let's get this merged in, this would be a great feature!

bompus referenced this pull request in bompus/codegraph Sep 5, 2026
Ports QingNagi/codegraph#361 (markdown extractor, heading nodes, name-matcher and
resolution hooks) onto experimental and adds the section-first doc tier from
feature/md-section-first: a doc-shaped query renders the best headed sections of
the markdown file it names, ranked by idf-weighted line hits with path tokens
weighted zero, capped at DOC_FILE_CAP per file. Markdown reaches an answer only
through that tier, generated-file detection ignores markdown bodies, and the
budget tiers count code files only so a README-heavy repo keeps its code answers.
@bompus

bompus commented Sep 5, 2026

Copy link
Copy Markdown
Contributor

Built this branch (3a73fed, 1.4.1 base) and measured it under headless Claude Code (Opus, --effort medium) on a repo with 109 markdown files and a doc-question bank, against the shipped build with the same repo rules. Two findings, then a fix I ended up carrying and can send as a follow-up.

As merged, the index would not get used. 24 cells over the doc bank: 0 codegraph_explore calls in every cell, calls and correctness identical to the shipped build (3.33 vs 3.00 calls, all correct). Two causes, both in the branch:

  1. src/mcp/server-instructions.ts still lists docs under what codegraph does not index, so the model is told not to ask. Correcting that line (and the client-side rule that mirrored it) took adoption from 0 of 6 to 6 of 6 cells.
  2. With adoption, every cell still ran the same Grep after explore, because the doc answer leads with ~23k characters of code-symbol blast radius and relationships before the section body, and the section it carries is a single matched line. The call was prepended to the chain, not substituted for it.

Section-first doc answers fix the second. On top of this branch I added a doc tier in tools.ts: when a query is doc-shaped and names a markdown file, the top three sections of that file render first, whole, ranked by idf-weighted line hits (a term that appears in the file's own path weighs zero, a heading the query covers word for word counts as named), capped at 8k characters per file; graph sections and the "additional files" pointer list stay off unless a code file rendered. Code queries are unaffected. Measured three times on the six-task doc bank, 12 fresh cells each, pre-registered criteria, $11 total:

Criterion Round 1 Round 2 Round 3
Correct 12 / 12 12 / 12 12 / 12
Tool calls, median (shipped build: 4) 1 1 1
Explore adopted 10 / 12 11 / 12 10 / 12
No Grep or Read after explore 8 / 12 9 / 12 9 / 12

Every explore call in all 36 cells chose the right file and section. The misses are the model's habit, not retrieval: one prompt about "one-shot reminders" is read as a question about the session's own scheduling tools in 0 of 6 cells under this build, and the other misses are a Grep on a file the prompt already names. Cost per cell was $0.30 against $0.54 on the shipped build.

Porting to 1.6.0 needed three more changes, which apply to this PR as well:

  • detectGeneratedFile reads the file body, so a README that quotes "generated by" is dropped from the index. Markdown should skip the header check.
  • Markdown section bodies land in the same FTS table as code, so an English word in a README matches into a code query's subgraph. I drop markdown nodes from the subgraph unless the doc tier seeded them.
  • The explore budget tiers key on fileCount; markdown took this repo from 466 to 575 indexed files, across the 500-file breakpoint, and every code answer grew a Relationships block. Counting code files only keeps code answers as they were (11 of 25 byte-identical, 3 reorder, 10 swap a fourth- or fifth-ranked padding file from FTS rank shifts).

The port is on my fork's experimental branch, commit bompus/codegraph@26c0190 (this PR's 13 files plus the tier). Happy to open it as a follow-up PR against this branch or against main, whichever the maintainer prefers, and to share the harness cells.

bompus referenced this pull request in bompus/codegraph Sep 10, 2026
Ports QingNagi/codegraph#361 (markdown extractor, heading nodes, name-matcher and
resolution hooks) onto experimental and adds the section-first doc tier from
feature/md-section-first: a doc-shaped query renders the best headed sections of
the markdown file it names, ranked by idf-weighted line hits with path tokens
weighted zero, capped at DOC_FILE_CAP per file. Markdown reaches an answer only
through that tier, generated-file detection ignores markdown bodies, and the
budget tiers count code files only so a README-heavy repo keeps its code answers.
bompus referenced this pull request in bompus/codegraph Sep 28, 2026
Ports QingNagi/codegraph#361 (markdown extractor, heading nodes, name-matcher and
resolution hooks) onto experimental and adds the section-first doc tier from
feature/md-section-first: a doc-shaped query renders the best headed sections of
the markdown file it names, ranked by idf-weighted line hits with path tokens
weighted zero, capped at DOC_FILE_CAP per file. Markdown reaches an answer only
through that tier, generated-file detection ignores markdown bodies, and the
budget tiers count code files only so a README-heavy repo keeps its code answers.
bompus referenced this pull request in bompus/codegraph Sep 28, 2026
Ports QingNagi/codegraph#361 (markdown extractor, heading nodes, name-matcher and
resolution hooks) onto experimental and adds the section-first doc tier from
feature/md-section-first: a doc-shaped query renders the best headed sections of
the markdown file it names, ranked by idf-weighted line hits with path tokens
weighted zero, capped at DOC_FILE_CAP per file. Markdown reaches an answer only
through that tier, generated-file detection ignores markdown bodies, and the
budget tiers count code files only so a README-heavy repo keeps its code answers.
@friedinando

Copy link
Copy Markdown

Great work bompus! there is something we can do in order to push further this pr?

bompus referenced this pull request in bompus/codegraph Oct 2, 2026
Ports QingNagi/codegraph#361 (markdown extractor, heading nodes, name-matcher and
resolution hooks) onto experimental and adds the section-first doc tier from
feature/md-section-first: a doc-shaped query renders the best headed sections of
the markdown file it names, ranked by idf-weighted line hits with path tokens
weighted zero, capped at DOC_FILE_CAP per file. Markdown reaches an answer only
through that tier, generated-file detection ignores markdown bodies, and the
budget tiers count code files only so a README-heavy repo keeps its code answers.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants