How 7 AI agents find context: we read the code

Raghav Chamadiya12 min read

how coding agents find context · coding agent context window · agent context compaction · agent harness architecture · grep vs index coding agent · openclaw architecture

On this page

I read the context code of seven open-source agent repos to answer one question: how does each one decide what the model gets to see? The short answer is that none of them builds an index of your code by default. They all find code the same way a person in a terminal would, with grep, find and read. Where they really differ is in what they throw away, and when.

That surprised me a little, because a lot of the public talk about agents is about retrieval: embeddings, code graphs, repo maps. In these seven codebases, the code that decides what the model reads is small. The code that decides what the model stops reading is large, and it is where the engineering effort has gone.

Each section below names the files involved, so you can check my reading yourself.

The seven repos

These are the agent repos we already had indexed at repowise. I did not pick them to make a point, and they fall into three different kinds of project, which matters for how you read the rest.

Full agent harnesses. A harness is the program that runs the loop: send messages to the model, run the tools it asks for, feed the results back, repeat.

  • pi (earendil-works/pi): a coding agent CLI plus the libraries under it (an LLM API layer, an agent loop, a terminal UI).
  • DeepSeek Harness (deepseek-ai/deepseek-harness): a harness where every feature is a plugin, built on the Cordis plugin framework.

General assistants that also code. These are personal assistant agents. They connect to chat apps, run scheduled jobs and browse the web, and coding is one of many jobs.

Layers on top of someone else's agent.

  • OmO (code-yeongyu/oh-my-openagent): plugins and agents for OpenCode and Codex, and since version 5.0 its own CLI on "senpi", which its README describes as a fork of pi.
  • free-claude-code (alishahryar1/free-claude-code): a proxy that sits between a coding agent (Claude Code, Codex, pi and eight others) and the model provider, and routes requests to other models.
  • Claude Code's public repo (anthropics/claude-code): the agent's core is not in this repo. What is there are 13 plugins, examples, and the source of four built-in "mods", which are plugins that hook the engine's events. So for Claude Code I can only describe the parts it exposes, not how its own search or compaction works.

Codex CLI, Gemini CLI, OpenCode, Aider, Cline and OpenHands are not in this post because we have not indexed them. I did not want to write about code I had not read.

The numbers side by side

RepoWhat it isFilesLines of codeHealthFinds code withShrinks context by
piCoding harness2,010379k6.9read, bash; grep, find, ls optionalSummary when under 16,384 tokens left; keeps last 20,000
DeepSeek HarnessPlugin harness12,578899k6.6read, glob, grep, bash; LSP tool optionalSummary at 80% of window; keeps 16%
OpenClawAssistant19,3614.36M6.7read, ls, shell (ripgrep, fd)Caps tool output, clears old results, summary
Hermes AgentAssistant8,7582.12M5.3read_file, search_files (ripgrep), terminalSummary by a cheaper model at 50% of window on large-window models, 75% under 512K
OmOLayer5,591591k8.1Host agent's tools plus LSP and ast-grep serversHost agent's compaction; frozen memory block
free-claude-codeProxy744168k5.6Not its jobAnswers some side requests locally; optional output filter
Claude Code (public repo)Plugins and mods1,12426k7.6Not in this repoNot in this repo

Scroll the table sideways to see every column.

Health: repowise's health score (1 to 10) rates how likely a file is to cause bugs and how hard it is to change, using static checks plus git history; the repo score is the average across files, weighted by lines of code. Shown to one decimal, as on each repo page. Files and lines are from our latest index of each repo, which is between 26 Jun and 30 Sep depending on the repo (dates in "How we measured"). Lines of code means non-blank, non-comment lines.

The size range is the first thing to notice: OpenClaw is about 167 times the size of Claude Code's public repo. Claude Code's engine isn't in its repo, so the ratio says nothing about the two agents themselves. It does show how much code an open agent grows once it has to talk to chat apps, phones, browsers, schedulers and dozens of model providers, and the coding part of OpenClaw is a small corner of it.

Finding code: grep, read, repeat

Every repo here that runs its own loop, plus OmO, gives the model the same basic kit: read a file, list a directory, search file contents, run a shell command.

  • pi ships four tools by default: read, bash, edit and write. grep, find and ls exist in packages/coding-agent/src/core/tools/ and can be switched on. The default system prompt in system-prompt.ts lists [read, bash, edit, write]. The model is expected to search with the shell.
  • DeepSeek Harness has read, write and edit in its tool-fs package and glob and grep in tool-fs-search. There is also a tool-lsp package that lets the model jump to definitions and find references through a language server (the same machinery your editor uses for "go to definition"). Its README is direct about the order: use it "when textual search is ambiguous", and "ordinary navigation should continue to use search and read". It is not in the default bundle.
  • OpenClaw builds its file tools in src/agents/core-coding-tools.ts: read, ls, edit, write, apply_patch, and a shell. I did not find a dedicated grep tool in the main runner. src/agents/utils/tools-manager.ts downloads pinned copies of ripgrep and fd, so search goes through the shell with those binaries.
  • Hermes Agent has search_files, backed by ripgrep in tools/file_operations_search.py, next to read_file and terminal.
  • OmO adds structure on top of whatever agent it runs in: a shared language server daemon (packages/lsp-daemon), LSP and ast-grep tools exposed as MCP servers, and "hashline" edits, where the model points at a line by a short hash of its content instead of by line number. When we indexed OmO in July, its Codex plugin also started a background CodeGraph index at session start (packages/omo-codex/plugin/components/codegraph). That component is gone now, and the current config migration drops the old codegraph setting as retired.

So the one index-like feature in the set was removed, and the language-server tools are optional. I read this as a deliberate bet by the teams: models have become good at grep, a grep result is always fresh, and an index is one more thing that can be out of date or broken on a user's machine. I think that bet is right for finding a function. I am less sure it holds for questions like "what breaks if I change this", where the answer is spread across files that never mention each other by name.

Building the system prompt

The second source of context is what the agent loads before the user types anything.

All of them read the project's instruction file. pi's resource-loader.ts looks for AGENTS.override.md, AGENTS.md and CLAUDE.md. DeepSeek Harness has an agent-instructions plugin that loads AGENTS.md or CLAUDE.md and reloads them after the agent edits a file. OmO has agents-md-core and a rules-engine package for nested instruction files. Claude Code's public repo includes the source of its agents-md mod, which hands AGENTS.md files to the engine as project instruction files, placed and framed like CLAUDE.md. It also attaches nested ones when the model reads a file in that folder. So every agent here that loads instruction files reads AGENTS.md, Claude Code included: its built-in agents-md mod loads AGENTS.md when a project has no CLAUDE.md (or alongside it, as an option).

Two of the repos handle instructions and memory in ways I think are well designed.

DeepSeek Harness puts instructions in the session history

Its context package README says injected instructions and references "enter session history as user-role messages, so they persist, replay, and compact like other conversation content". The upside is that a change to AGENTS.md mid-session does not rewrite the system prompt.

Hermes Agent and OmO freeze memory at session start

Hermes keeps two memory files, MEMORY.md (agent notes) and USER.md (a profile of the user). The docstring in tools/memory_tool.py says both enter the system prompt "as a FROZEN snapshot at session start; mid-session writes hit disk but never change the prompt". OmO's changelog describes the same fix: a memory commit during a session no longer rewrites that session's system prompt.

Both teams are protecting the prompt cache. Providers charge less for a prompt prefix they have seen before and answer it faster, and changing one byte near the top of the prompt loses the whole cache. A surprising amount of context code in these repos exists to keep the top of the prompt byte-for-byte stable.

Trimming context, where most of the code is

The seven repos differ most in how they trim context, and this is also where most of their context code lives. They trim in two ways: they cap each tool result as it comes in, and they compact, which means that when the conversation gets close to the model's limit, older turns are replaced by a summary written by a model.

RepoPer-result capCompaction triggerWhat it keeps
pi2,000 lines or 50 KB; grep lines cut at 500 charsFewer than 16,384 tokens leftThe last ~20,000 tokens
OpenClaw16,000 chars by default, 32,000 above 100k-token windows, 64,000 above 200k; tool output capped at 30% of the windowReserve capped at 25% of the window; up to 3 overflow retriesAgent core defaults: 16,384 reserve, last 20,000 tokens
Hermes Agentread_file 100,000 chars50% of the window on large-window models, 75% under 512KFirst 3 and last 20 messages; middle summarised
DeepSeek HarnessOversized results pruned first; optional "spill" to disk80% of the windowAbout 16% of the window

Scroll the table sideways to see every column.

Hermes Agent compacts earlier than most

Its default trigger is 50% of the window on models with 512K tokens or more, and 75% on smaller windows. The code comment explains the floor: on a small window, compacting at 50% frees so little room that compaction would fire again every turn or two. On a large window that is much sooner than DeepSeek Harness (80%) or pi (16,384 tokens left, about 97% of a 512K window). It also does the summary with a separate, cheaper "auxiliary" model, and it prunes big tool outputs before summarising (agent/context_compressor.py). Compacting sooner costs a summary call more often. In return, the model rarely works with a nearly full window, where answers tend to get worse. For a personal assistant that runs long sessions over days, I think that is a sensible trade.

DeepSeek Harness waits, then spills

It compacts at 80% and keeps about 16%. Its spill packages can also move a large tool result out of the context entirely: the full text goes to a file, and the model gets back a short pointer with instructions on how to read more. That is the same idea as a person saving a long log to disk instead of pasting it into a chat.

OpenClaw clears old tool results

Besides capping each result, src/agents/embedded-agent-runner/tool-result-truncation.ts replaces tool output that has outlived the prompt cache window with [Old tool result content cleared], and old images with [image removed during context pruning]. The model keeps the record that it ran the tool but loses the bytes.

pi keeps compaction short

pi's whole compaction logic is in packages/coding-agent/src/core/compaction/compaction.ts: compact when the context is within 16,384 tokens of the limit, keep the most recent ~20,000 tokens, summarise everything before that. It is the easiest of the four to read in one sitting.

free-claude-code cuts down requests

As a proxy, free-claude-code never sees a repo directly, but src/free_claude_code/api/optimization_handlers.py answers five kinds of side request from Claude Code locally without calling a model: quota checks, command-prefix detection, title generation, suggestions, and file-path extraction. Its README also offers an optional filter that shortens terminal output before the model sees it.

pi is underneath more of this than its size suggests

This was the finding I did not expect: pi is one of the smaller repos in the set (379k lines), but it shows up inside three others.

  • OpenClaw depends on @earendil-works/pi-tui in its package.json. Its packages/agent-core has the same compaction file names as pi (compaction.ts, branch-summarization.ts) and the same defaults: 16,384 reserve tokens and 20,000 recent tokens. Its truncate.ts has pi's exact limits: 2,000 lines, 50 KB, 500 characters per grep line. OpenClaw also reserves the tool names bash, edit, find, grep, ls, read, write, which is pi's tool list. I don't know the history behind this and won't guess at it. The code shows that OpenClaw's agent core shares pi's design and numbers.
  • OmO says in its README that version 5 "runs on senpi, our fork of pi".
  • DeepSeek Harness has a dsh-llm-pi-ai package that adapts pi's model API library. Its own description calls it a "design-verification twin" of the main DeepSeek adapter, so it is not the default path.

If you want to learn how a coding agent's loop works from source, I would start with pi for this reason. It is small enough to read, and many larger agents copied its choices.

Giant functions, and what happened to them

When we indexed these repos, two of them held the two longest functions in the whole set. Hermes Agent's run_conversation in agent/conversation_loop.py was 6,591 lines (index of 19 Aug). OpenClaw's runEmbeddedAttempt in src/agents/embedded-agent-runner/run/attempt.ts was 4,972 lines (index of 26 Jun). A function like that is usually the agent loop itself: every retry rule, provider quirk and recovery path ends up as another branch in the one place that sees everything.

Both teams have since split them up. On 6 Oct, Hermes's conversation_loop.py is 1,835 lines, down from 8,298 at our index, and run_conversation is a short wrapper around _run_conversation_turn. OpenClaw's attempt.ts is 567 lines, down from 5,808. The run/ folder next to it went from 65 non-test files to 173, with names like attempt-prompt-build.ts, attempt-history-prepare.ts and attempt-tool-search-executor.ts.

I like this as a story about how agent code grows. The loop is the first thing you write and the place every fix lands, so it gets big fast. Our indexes of both repos predate the split, so the health scores in the table above describe the code before it.

What these numbers can't tell you

  • Health measures the code. The health score, defined under the table above, says nothing about how well the agent solves tasks. The lowest score in this set belongs to Hermes Agent, one of the two assistants, which carry chat integrations, desktop apps and test suites in the same repo. The second lowest is free-claude-code, a small proxy, so size and scope don't explain every low score here.
  • Claude Code is mostly invisible here. Its score describes 26k lines of plugins and mods, not the agent. Treat its row as "the public extension surface".
  • Snapshots are dated. Our OpenClaw index is from 26 Jun and OmO's from 17 Jul. The code I quote is from each repo's main branch on 6 Oct, which is why some numbers in the text are newer than the table.
  • Defaults can change. I quoted default settings. Users change them, and some repos pick different values per model or per surface.
  • I read code and didn't run benchmarks. Whether early compaction or late compaction gives better results is a measurable question, and this post does not measure it.

Where repowise fits

Disclosure: I build repowise, which is a context layer for agents. It indexes a repo's dependency graph, git history and docs and serves them over MCP, which is a standard way for agents to call outside tools. Given what I found, I'd describe it as something you add next to grep. Grep answers "where is this name". An index helps with "what depends on this" and "why is it shaped like this", which grep can't answer in one call. Every repo page linked in this post was produced that way, and you can connect any of them to your own agent.

How we measured

  • Repos: the seven agent repos already indexed at repowise with a code health score. No repos were indexed for this post.
  • Numbers: files, lines of code, health and longest functions come from each repo's latest ready snapshot in our database, read on 6 Oct 2026. Snapshot dates: pi 30 Sep, DeepSeek Harness 28 Sep, Claude Code 27 Sep, free-claude-code 19 Sep, Hermes Agent 19 Aug, OmO 17 Jul, OpenClaw 26 Jun. Health is shown to one decimal, the way the repo pages show it.
  • Code reading: shallow clones of each repo's default branch on 6 Oct 2026, plus files fetched at the indexed commits for the before-and-after comparisons. Every file path in the post is from those clones.
  • Health score: defined under the first table. The method is on each repo's code health page.

If you maintain one of these repos and something here is wrong, tell me and I will fix it.

Try it: connect any of these repos to Claude or ChatGPT and ask your own agent how it handles context. It's a good way to check my reading.