On this page
Claude Code doesn't build a codebase index; it greps and reads files each session, which costs tokens on large repos. For a prebuilt index, add an MCP server: CodeGraph or code-review-graph for a fast call graph, Serena for language-server symbols, cocoindex-code for semantic search, Graphify for a knowledge graph, or repowise for graph plus history, docs and health.
| Tool | What it indexes | Output tokens vs bare agent (Codex) | Tool calls per answer (Codex) | Local only | License |
|---|---|---|---|---|---|
| repowise | Call graph, git history, generated docs, decisions, health, embeddings | -31.6% | 3.8 | Yes (hosted optional) | AGPL-3.0 |
| CodeGraph | Symbol and call graph in SQLite with full-text search | -24.4% | 4.0 | Yes | MIT |
| Serena | Live symbols through language servers (no stored graph) | -14.8% | 10.1 | Yes | GPL-3.0 (core), MIT (SolidLSP) |
| Graphify | Knowledge graph from AST, plus docs and PDFs through your model | -8.9% | 7.4 | Code yes; docs go to your model | Apache-2.0 |
| code-review-graph | Tree-sitter graph of calls, imports and tests in SQLite | -6.0% | 7.2 | Yes | MIT |
| cocoindex-code | AST-chunked embeddings for semantic search | not run in the agent loop | not run | Yes with local embeddings | Apache-2.0 |
Scroll the table sideways to see every column.
Bare agent: 1,828 output tokens and 7.2 tool calls per answer on the same questions. Numbers are from our August 2026 benchmark on django/django, explained under "How we measured" below.
We build repowise, so weigh our entry accordingly. The measurements we publish about it, with their methods, are on the benchmarks page.
Cursor-style indexing and what Claude Code does instead
When people ask for Cursor-style indexing in Claude Code, they usually mean one thing: the editor has already read the repo before you ask, so a question like "where do we validate webhooks?" doesn't start with ten rounds of searching.
Cursor's own approach has changed. Its current search docs describe a local "Instant Grep" index built on your machine, and say Cursor "does not store embeddings of your codebase for search." So Cursor itself now relies on a fast index for exact matches plus an agent that searches well, with no large vector store behind it.
Claude Code takes the plain route. It comes with tools to grep, glob and read files, and nothing persists between sessions. On a small repo that works well and costs little. On a large one, every new session repeats the same exploration: list the folders, grep a name, open five files, discover four were wrong, open three more. That exploration is most of the token bill.
An MCP server fixes this by doing the expensive reading once, offline, and answering questions from the result. The six tools below differ in what they read and what they can answer.
The six options
repowise
repowise parses the repo into a call and import graph, mines the git history (hotspots, owners, files that change together), writes a documentation page per module, extracts architectural decisions, scores code health and builds search embeddings. Agents get ten tools such as get_overview, get_context, get_risk, search_codebase and get_answer.
That breadth is why it led the agent-loop run, and also why it is slow to build. On django it took 366.8 seconds without generated docs and 1,058 seconds with them, against CodeGraph's 16.4 seconds. Updates after the first index are incremental, so you pay that once. Disclosure: we build it.
uv tool install repowise
cd your-repo
repowise init --yes --no-prose # no API key, no spend; writes .mcp.json for Claude Code
CodeGraph
CodeGraph builds a symbol and call graph into a local SQLite database with full-text search, and keeps it current with a file watcher (a 2-second debounce by default). It advertises a single tool to the agent, which keeps its schema small. It was the clear second place on Codex and the fastest indexer we measured.
npm i -g @colbymchenry/codegraph
codegraph install
Serena
Serena works differently from the other five because it doesn't store a graph. It drives real language servers (the same engines behind go-to-definition in your editor) and exposes symbol-level tools for reading and editing. That makes it strong for precise refactors across 40+ languages. On questions about how a codebase works, it made more tool calls than the bare agent (10.1 against 7.2) while still cutting output tokens by 14.8%.
uv tool install -p 3.13 serena-agent
claude mcp add serena -- serena start-mcp-server --context claude-code --project "$(pwd)"
Graphify
Graphify builds a knowledge graph from tree-sitter parsing with no API calls for code, plus an interactive HTML view and a markdown report. Docs, PDFs and images are sent to whatever model your assistant runs, unless you pass --code-only. It can serve the graph over MCP, and a git hook rebuilds it after commits. It is the most-starred project here and the most visual.
uv tool install graphifyy
graphify install # then type /graphify . in Claude Code
code-review-graph
code-review-graph maps calls, imports, inheritance and test coverage with tree-sitter into one SQLite file, aimed at reviews: given a change, which code is affected. It covers 40+ languages and lets you add parsing rules for others without forking.
pip install code-review-graph
code-review-graph install --platform claude-code
cocoindex-code
cocoindex-code is the closest thing here to classic "semantic search": it chunks code along the syntax tree, embeds the chunks (locally with a small default model, or through a cloud provider), and re-indexes only changed files. It was in our retrieval benchmark, not in the agent loop, so the table above has no token figure for it.
pipx install 'cocoindex-code[full]' # [full] includes local embeddings
claude mcp add cocoindex-code -- ccc mcp
One more you'll see in lists: Claude Context from Zilliz. It does semantic search well, but its documented setup needs a Zilliz Cloud or Milvus vector database and an OpenAI key for embeddings, so it isn't local-only by default. We didn't benchmark it.
The numbers, and how far to trust them
We ran every tool against the same django questions in an agent loop, each with a fresh index on the same pinned commit and byte-identical prompts.
On Codex, every tool was called on every question, so the comparison is like for like. repowise cut output tokens by 31.6% and needed 3.8 tool calls instead of 7.2. CodeGraph cut 24.4% with 4.0 calls. That is a real second place, and the honest reading is that more than one tool in this field works.
On Claude Code (Sonnet), the picture is messier, and you should know that before installing anything. The agent decides whether to call an MCP tool at all, and on 15 questions it called repowise 15 times, CodeGraph 13, Serena 4, Graphify 3 and code-review-graph 0. When we reran the same setup on later days with nothing changed on our side, our own adoption dropped to 4 of 15 and then 3 of 15. Adoption is a property of the tool, the harness and the day together.
Whichever tool you pick, tell Claude Code to use it. A line in CLAUDE.md such as "Use the repowise tools before grepping for architecture questions" moves adoption more than any tool choice. We wrote more about this in your MCP server is probably not being called.
The table leaves out two results that matter for the choice:
- On a separate retrieval test (ContextBench, 42 held-out bug reports in Python and Go), repowise's
get_answerfound 0.876 of the files each real fix touched, CodeGraph 0.610, Graphify 0.546, code-review-graph 0.445 and cocoindex 0.361. We came last on this same test, at 0.228, the first time we ran it. The cause was a bug in how our tool discarded candidates, the fix ships to every user, and the story is in we benchmarked ourselves and came last. - On index time, repowise is 22x slower than CodeGraph like for like, and 135x with generated docs on. If a call graph is all you want, that cost buys you nothing.
Which one to pick
- For the cheapest improvement and a fast build, use CodeGraph.
- For precise symbol-level edits, use Serena, ideally alongside one of the graph tools.
- For a picture of the codebase that you can read as well as the agent, use Graphify.
- If your main use is PR review, use code-review-graph.
- To find code by meaning ("find code that does X"), use cocoindex-code.
- If the agent should also know history, owners, risk, decisions and health, or you want the same index in Claude.ai, ChatGPT, Cursor and Claude Code through a hosted MCP address, use repowise. The hosted version adds one command for Claude Code:
claude mcp add --transport http repowise https://api.repowise.dev/mcp/OWNER/REPO --header "Authorization: Bearer YOUR_API_KEY". See the setup guide for Cursor and other editors.
How we measured
The agent-loop numbers come from August 2026 runs on django/django at one pinned commit: Codex (gpt-5.6-sol) on 43 questions that all six arms completed, and Claude Code (claude-sonnet-5) on 15 questions. Versions: CodeGraph 1.5.0, Graphify 0.9.31, Serena 1.6.2.dev0, code-review-graph 2.3.7. Each tool had its full advertised tool list and a freshly built index. The figures are output tokens per answer, because cached input tokens make dollar costs depend on which arm ran first. Raw data, including failed runs, is in the public benchmark repository linked from our benchmarks page.
These numbers have clear limits: the agent loop used one repository in one language, with questions of four to nine turns. Your repo, your model and your prompt will move these numbers. Tool versions have moved since August too. cocoindex-code and Claude Context were not in the agent loop. Setup commands and licenses were checked on the projects' own pages on 6 October 2026.
FAQ
Does Claude Code index my codebase like Cursor?
No. Claude Code has no persistent index. It uses grep, glob and file reads in each session. To get an index built ahead of time, you add an MCP server that indexes the repo and answers the agent's questions from it.
What is the best codebase MCP server for Claude Code?
It depends on the job. In our django agent-loop run on Codex, repowise cut output tokens the most (31.6%) and CodeGraph was second (24.4%) with a much faster build. Serena is strongest for symbol-level edits. Under Claude Code, adoption varied so much between runs that you should instruct the agent to use whichever tool you install.
Is there a semantic code search MCP that runs locally?
Yes. cocoindex-code embeds code locally with a small default model when you install the `[full]` variant. Claude Context does semantic search too but, as documented, needs a Milvus or Zilliz Cloud database and an OpenAI key.
How do I make Claude Code index a repo?
Install an indexing MCP server and register it. For repowise: `uv tool install repowise`, then `repowise init --yes --no-prose` inside the repo, which builds the index without an API key and writes `.mcp.json` so Claude Code picks it up. Other tools have their own `install` commands, listed in this post.
Why does my agent ignore the MCP server I installed?
The model decides when to call a tool, and many skip MCP tools in favour of built-in grep. In our Claude Code runs, adoption varied from 0 to 15 of 15 questions between tools, and even our own dropped to 3 of 15 on a rerun with nothing changed. Add a line to `CLAUDE.md` telling the agent when to use the server.