Serena vs Graphify vs CodeGraph vs repowise: Benchmarked on the Same Repos

Raghav Chamadiya11 min read

serena vs graphify · graphify alternatives · codegraph alternatives · serena alternative · graphify vs codegraph · code knowledge graph mcp

On this page

Serena, Graphify, CodeGraph and repowise all give coding agents a better view of a codebase, but they do different jobs. Serena drives language servers for symbol-level reading and editing. Graphify builds a knowledge graph people can open and share. CodeGraph builds the fastest local call graph. repowise adds git history, docs, decisions and health on top of the graph.

SerenaGraphifyCodeGraphrepowise
What it buildsNo stored index; queries live language serversKnowledge graph from AST, plus docs and PDFs via your modelSymbol and call graph in SQLite with full-text searchCall graph, git history, generated docs, decisions, health, embeddings
Best atPrecise symbol edits and refactorsA browsable graph and report for humans and agentsCheap, fast structural answers"How does this work, who owns it, what breaks" questions
Output tokens vs bare agent (Codex, django)-14.8%-8.9%-24.4%-31.6%
Tool calls per answer (bare agent 7.2)10.17.44.03.8
File retrieval, 42 held-out instancesnot run0.5460.6100.876 (0.228 on first run)
Index time, djangonone needed141.5 s16.4 s366.8 s (1,058 s with docs)
Tools advertised to the agent2910110
LicenseGPL-3.0 (MIT for SolidLSP)Apache-2.0 (relicensed from MIT)MITAGPL-3.0

Scroll the table sideways to see every column.

Disclosure: we build repowise. Every number here comes from the same public benchmark, which also includes the rows where we lose. Method and caveats are at the end.

We build repowise, so weigh our entry accordingly. The measurements we publish about it, with their methods, are on the benchmarks page.

Four tools, four different ideas

All four show up in the same "give your agent codebase context" lists, but each is built on a different design choice, and that choice decides what the tool is good at. The benchmark numbers below make more sense once you know which choice each tool made.

Serena doesn't build an index at all. It starts real language servers, the same engines your editor uses for go-to-definition and rename, and gives the agent tools like "find symbol", "find references" and "replace symbol body". Because nothing is stored, nothing goes stale. The cost is that a broad question ("how does request routing work here?") has to be answered by many small lookups, which is why Serena made more tool calls than an agent with no tools at all (10.1 against 7.2) while still using fewer output tokens. It covers 40+ languages through language servers, and a paid JetBrains plugin can stand in for them.

Graphify parses code with tree-sitter (no model calls for code), and can also pull in Markdown, PDFs and images by sending them to the model your assistant already uses. The output is three files: an interactive graph.html, a GRAPH_REPORT.md of key concepts and surprising connections, and a graph.json the agent can query, including over MCP. Graphify treats the graph as something a team opens and shares as well as something the agent queries, and it is the most-starred of the four by a wide margin (about 124,000 GitHub stars against Serena's 30,000).

CodeGraph builds a symbol and call graph into a local SQLite database with full-text search, keeps it current with a file watcher, and exposes a single tool to the agent, which keeps its schema small (1,567 characters). Under Claude Code it was the second most-called tool in our runs.

repowise builds a graph too, then mines git history (hotspots, owners, files that change together), writes a page per module, extracts architectural decisions, scores code health, and builds search embeddings. Agents reach it through ten tools such as get_overview, get_context, get_risk and get_answer. That's more to build, which is why it's the slowest indexer here.

Agent loop: tokens and tool calls

We ran each tool against the same django questions with a freshly built index on the same pinned commit and byte-identical prompts, plus a bare agent as the control. On Codex, every tool was called on every question, so this is a like-for-like comparison.

Tool (Codex, 43 questions)Output tokensvs bare agentTool callsLeaner than bare on
repowise1,250-31.6%3.837 of 44
CodeGraph1,383-24.4%4.037 of 44
Serena1,550-14.8%10.135 of 43
Graphify1,658-8.9%7.431 of 43
bare agent1,828baseline7.2n/a

Scroll the table sideways to see every column.

repowise and CodeGraph completed 44; the deltas use the 43 every tool completed.

CodeGraph's result is a close second place from a tool that indexes in seconds, and the fair reading is that more than one tool in this group works well.

On Claude Code (Sonnet, 15 questions), the result depends on whether the agent decides to call the tool at all. It called repowise on 15 questions, CodeGraph on 13, Serena on 4 and Graphify on 3. Output tokens fell 15.9% for repowise, 11.7% for CodeGraph, 11.3% for Serena and 0.0% for Graphify. Then we reran the identical setup on later days and our own adoption dropped to 4 of 15 and then 3 of 15, with nothing changed on anyone's side. Treat any adoption number, ours included, as a fact about one harness on one day. If you install any of these, add a line to CLAUDE.md telling the agent when to use it.

Retrieval: finding the right files, and the run where we came last

Graphify, CodeGraph and repowise were also tested on ContextBench: a real GitHub issue goes in, and the score is how many of the files the actual fix touched come back. Grading is deterministic, with no model acting as judge. Serena wasn't run here, since it has no index to retrieve from.

The first time we ran this benchmark, we came last, scoring 0.228 against CodeGraph's 0.609. On the 20 Go instances we scored 0.025 while every other tool cleared 0.50, and we published that result before fixing anything.

The cause was a set of query-time shortcuts that returned early, before the ranked candidates were used. One of them fired on ordinary words like field and schema that appear in most bug reports on database-backed code. In 18 of the 19 cases where CodeGraph beat us, our own search_codebase had already found the file on the same index. We fixed the shortcuts in ordinary releases that every user gets.

We had split the 112 instances into 70 to work on and 42 sealed ones before the first run, and only touched the sealed 42 once, after the fix:

ToolFile coverage (42 sealed)Files returned per question
repowise get_answer0.87619.2
repowise search_codebase0.7428.2
CodeGraph0.61014.0
Graphify0.54634.5

Scroll the table sideways to see every column.

Coverage alone rewards returning more files, and Graphify reaches its 0.546 by returning 34.5 files per question, about four times as many as search_codebase returns. get_answer also returns about 19 files to reach 0.876, so search_codebase, at 8.2 files for 0.742, is the better tool if you count coverage per file. The full write-up is in we benchmarked ourselves and came last.

Call-graph accuracy

A call graph is only useful if its edges are real. We graded each tool's call edges against compiler-produced graphs on five Go and TypeScript repositories (37,853 reference edges), as precision (how many edges are correct) and recall (how many of the real edges were found). Two examples:

ReporepowiseCodeGraphGraphify
zod (TypeScript)0.992 / 0.7030.729 / 0.3730.825 / 0.248
cobra with tests (Go)0.972 / 0.6840.929 / 0.7630.971 / 0.433

Scroll the table sideways to see every column.

Precision / recall. On zod we lead both. On cobra, CodeGraph finds more of the real graph, and across all five Go cells we don't lead recall. What holds in all seven repo and test-set cells is narrower: no tool that finds as much of the graph as we do gets more of it right. Serena isn't in this table because it asks language servers at query time and stores no graph to grade.

Build cost and the row we lose

SerenaGraphifyCodeGraphrepowise
Full index, djangonone141.5 s16.4 s366.8 s, or 1,058 s with docs
Graph only, median over 35 reposn/a12.23 s, 860 MB3.65 s, 757 MB2.77 s, 75 MB

Scroll the table sideways to see every column.

When it builds only the call graph, repowise is the fastest and lightest of the three (2.77 seconds and 75 MB at the median). When it builds everything a default install builds, it is 22x slower than CodeGraph, and 135x slower with generated docs on. In that django run the extra time went into 36,485 graph nodes, git history across 2,630 files, 3,392 written and embedded doc pages and 5,317 health findings. If a call graph is all you need, that extra work gives you nothing, and CodeGraph is the better choice. All three indexers can update after edits: CodeGraph with a file watcher, Graphify with a git hook or --watch, and repowise incrementally.

Which to pick

  • Pick Serena if your agent edits a lot of code and you want IDE-grade symbol operations with nothing to rebuild. It pairs well with one of the graph tools for broad questions.
  • Pick Graphify if you want a graph you can open, explore and share with your team, and you want docs and PDFs in the same graph.
  • Pick CodeGraph if you want the cheapest reliable token saving on code questions and a build measured in seconds.
  • Pick repowise if your agent's questions are about history and risk as much as structure (who owns this, what changes with it, why is it built this way, what is safe to delete), or if you want the same index in Claude.ai, ChatGPT, Cursor and Claude Code through a hosted MCP address.

Our compare pages for CodeGraph, Serena and Graphify go through each pairing one at a time.

How we measured

All numbers were measured in August 2026 with CodeGraph 1.5.0, Graphify 0.9.31, Serena 1.6.2.dev0 and repowise at the versions recorded in the benchmark repository. Agent loop: django/django at one pinned commit, Codex (gpt-5.6-sol) on the 43 questions all arms completed, Claude Code (claude-sonnet-5) on 15. Retrieval: ContextBench, 42 sealed Python and Go instances. Call graph: five repositories graded against compiler output. Index time: django on one machine. Each competitor was set up following its own docs, and Serena needs an explicit project activation before it answers anything, which we did.

What this can't tell you: the agent loop is one repository and short questions. Retrieval says we find the right files, not that an agent writes a better fix. Tool versions have moved since August, and all four projects ship often. Raw data, including failed and invalidated runs, is in the public benchmark repository linked from the benchmarks page. Licenses and setup details were checked on each project's own page on 6 October 2026.

FAQ

What is the difference between Serena and Graphify?

Serena has no stored index. It queries language servers live and gives the agent symbol-level read and edit tools. Graphify builds a knowledge graph ahead of time from tree-sitter parsing, optionally including docs and PDFs, and produces files people can browse. In our Codex run Serena cut output tokens by 14.8% and Graphify by 8.9% against a bare agent.

What are the best Graphify alternatives?

CodeGraph if you want a fast local call graph with one simple tool, code-review-graph if your focus is PR review, and repowise if you also want git history, ownership, generated docs and health in the same index. In our agent-loop run on django, CodeGraph and repowise both reduced output tokens more than Graphify did.

What are the best CodeGraph alternatives?

Serena suits symbol-level editing with no index to build. Graphify suits teams that want a shareable graph view. repowise had the larger token reduction (31.6% against 24.4%) and higher retrieval coverage (0.876 against 0.610), but indexes 22x slower. If speed of indexing matters most, CodeGraph is hard to beat.

Is there a Serena alternative that doesn't need language servers?

Yes. CodeGraph, Graphify and repowise parse code with tree-sitter and store the result, so they don't need a language server per language. The trade-off is staleness: a stored index must be updated after edits, which all three do incrementally or with a file watcher or git hook.

Can I use more than one of these together?

Yes. A common pairing is Serena for edits plus one graph tool for broad questions. Every tool's schema is loaded into the prompt, so check how often the agent actually calls each server, and add a line to `CLAUDE.md` saying which one to use for which kind of question.