On this page
- The numbers first
- LangChain: about seven steps across three packages
- Google ADK: six steps inside one function
- Composio: the call leaves your machine
- mem0: one LLM call and an anti-hallucination trick
- GPT Researcher: the model never picks a tool
- Abstraction versus overhead
- What the numbers can't tell you
- How we measured
Most of the "abstraction or overhead" argument about agent frameworks is about one thing: what happens between the model saying "call this tool" and your function running. So I traced that path, in the actual code, through five popular repos that each sit at a different layer of an agent stack.
The short version: every step I found exists for a reason, and most of those reasons are things you would end up writing yourself. The cost is that the steps are spread across packages, so when something breaks you are reading three codebases instead of one.
I only used repos we have indexed, because I did not want to quote numbers you can't check. As a result, the five repos do different jobs, and you could stack all of them in one product. I think that makes a more useful comparison than five frameworks competing for the same job.
- langchain-ai/langchain: the general framework. Model and tool abstractions, plus an agent loop.
- google/adk-python: a framework from a platform vendor. Agents, flows, sessions, plugins.
- composiohq/composio: the tools layer. Toolkits for outside services, plus the auth to call them.
- mem0ai/mem0: the memory layer.
- assafelovic/gpt-researcher: an application built on top.
The numbers first
| Repo | Layer | Code health | Lines of code | Files | Hotspots | Recorded decisions |
|---|---|---|---|---|---|---|
| composio | Tools | 7.7 | 260,291 | 2,969 | 286 | 69 |
| gpt-researcher | Application | 7.2 | 63,716 | 614 | 7 | 16 |
| mem0 | Memory | 7.1 | 228,481 | 1,491 | 94 | 2 |
| adk-python | Framework (vendor) | 6.7 | 418,576 | 2,369 | 238 | 14 |
| langchain | Framework (general) | 6.6 | 320,283 | 2,859 | 57 | 38 |
Scroll the table sideways to see every column.
repowise's health score (1 to 10) rates how likely a file is to cause bugs and how hard it is to change, using static checks plus git history; the repo score is the average across files, weighted by lines of code. A hotspot is a file that changes often and is also complex. "Recorded decisions" are design decisions we could find written down in PR descriptions, ADRs and code comments.
The scores don't sort by size. Composio, the third largest, scores highest; GPT Researcher, by far the smallest, comes second; and the two frameworks, ADK and LangChain, score lowest. Across our whole index median health drops as a repo grows, but five repos are too few to show that pattern. LangChain's score has a specific cause that is visible in its hotspots, which I come back to below.
The diagram below shows how the five fit together. Each arrow is an integration that exists in these repos' code.
LangChain: about seven steps across three packages
The new agent entry point is create_agent in libs/langchain_v1/langchain/agents/factory.py. It builds a graph with a model node and a tools node. When the model replies with a tool call, this is the path, in order:
- The model node returns an
AIMessagewithtool_calls. - The edge built by
_make_model_to_tools_edgedecides where to go next. It skips calls that already have results, and sends each pending call to the"tools"node as its ownSend, so calls can run in parallel. - LangGraph's
ToolNodepicks the tool by name.ToolNodelives in the separatelanggraphpackage, outside this repo. - Any middleware you registered with
wrap_tool_callruns, chained by_chain_tool_call_wrappers. BaseTool.runinlibs/core/langchain_core/tools/base.pyfireson_tool_startcallbacks and validates the arguments with_parse_inputand_to_args_and_kwargs.StructuredTool._runcalls your function._format_outputturns the return value into aToolMessage, and_make_tools_to_model_edgesends it back to the model.
Steps 2, 4 and 5 are the ones people call overhead, and they are also the ones I would rebuild by hand on day three of a real project: parallel calls, a place to add retries or approval, and argument validation before your code sees bad input. One thing I liked in step 2: a comment explains they stopped copying the full state into every Send because it was "previously O(N^2)".
The real cost is that the path crosses three packages: langchain, langchain_core and langgraph. A stack trace from a failed tool call will jump between all three.
LangChain's lower score has a lot to do with how the repo is organised. It is a monorepo with 16 partner packages under libs/partners/ (anthropic, openai, ollama, groq and so on). Several of its top hotspots are model-profile data files such as libs/partners/openrouter/langchain_openrouter/data/_profiles.py, which is 8,718 lines of generated model metadata with a "DO NOT EDIT THIS FILE MANUALLY" header. Others are the provider chat model files, like langchain_openai/chat_models/base.py, which change every time a provider ships a model. That is the cost of supporting every provider in one repo, and it pulls the average down without saying much about the agent loop itself.
Google ADK: six steps inside one function
In ADK, the model response comes back as an event, and BaseLlmFlow in src/google/adk/flows/llm_flows/base_llm_flow.py hands its function calls to handle_function_calls_async. From there _batch_executor.py prepares the calls and runs them together, and each one goes through _execute_single_prepared_call in flows/llm_flows/tools/_caller.py.
That function is the most honest piece of code in this whole comparison, because it numbers its own steps in comments:
- Plugin
before_tool_callbackgets a chance to answer instead of the tool. - If no plugin answered, the agent's own before-tool callbacks get a chance.
- If nobody answered, call the tool:
_call_tool_asyncrunsFunctionTool.run_async, which preprocesses arguments and then calls your function. - Plugin
after_tool_callbackcan replace the result. - The agent's after-tool callbacks can replace it.
- If any of them did, use that result.
Errors get their own path through on_tool_error callbacks at both the before and the call stage.
So ADK has two layers of hooks (plugins, then agent) on both sides of every call. That is more hooks than LangChain, but they all live in one function you can read top to bottom. I find that easier to debug than middleware spread across packages, even if the total amount of machinery is similar.
One more thing you notice in the code: flows/llm_flows/functions.py is now a 135-line "backward-compatibility shim" that only re-exports names from the new flows/llm_flows/tools/ folder. The tool-calling code was recently split into smaller files. Our snapshot is from 11 August, so some of its hotspots, like tests/unittests/flows/llm_flows/test_request_confirmation.py, are probably that same work in progress.
Composio: the call leaves your machine
Composio gives your agent tools for outside services (GitHub, Gmail, Slack and many more) and handles the logins for them, so it sits underneath a framework. Its repo has adapters for most frameworks: 13 Python providers under python/providers/ (including langchain and google_adk) and 11 TypeScript providers under ts/packages/providers/.
If you use it from LangChain, python/providers/langchain/composio_langchain/provider.py wraps each Composio tool as a LangChain StructuredTool. So you get all seven LangChain steps above, and then step 6 calls into Composio:
Tools.executeinpython/composio/core/models/tools.pylooks up the tool's schema, caching it.- It swaps local file paths for uploads if you turned that on.
- It runs your
before_executemodifiers, which are hooks that can change the arguments. _execute_toolpicks the toolkit version and refuses to run against"latest"unless you passdangerously_skip_version_check.- It sends one HTTP request to the Composio API.
- It runs your
after_executemodifiers on the result.
The auth and the call to GitHub or Gmail happen on Composio's servers, outside this repo, so the SDK here is a client, and a careful one. Step 5 has my favourite comment of the five repos: retries are disabled because "tool execution is a non-idempotent write, and a silent retry after a read timeout can duplicate the side effect." Non-idempotent means running it twice does something twice, like sending an email twice. Most HTTP clients retry by default. Turning that off for tool calls is the right call.
Most of the repo is TypeScript (1,390 of the files, against 234 Python), because the TypeScript SDK and the composio CLI live here too. Its hotspots are the CLI's tools.execute command and the tool router session code, which is where the product is moving.
mem0: one LLM call and an anti-hallucination trick
mem0 sits off the tool-call path: an agent calls it to remember things, so I traced Memory.add instead, in mem0/memory/main.py. The work happens in _add_to_vector_store, which labels its own phases:
- Phase 0: load the last 10 messages for this user or session.
- Phase 1: embed the new messages and search the vector store for the 10 closest existing memories.
- Phase 2: one LLM call, asking for JSON, to extract new facts given the existing ones.
- Phase 3: embed the extracted facts in one batch, insert them, and write history.
So every add costs one LLM call, two rounds of embedding and a vector search. That is the price of a memory layer that decides which facts are worth keeping, instead of storing every message as it arrives.
The detail I liked is right before Phase 2. Existing memories have UUIDs, long random ids. Before showing them to the model, mem0 replaces each with a small number ("0", "1", "2") and keeps a map back, with the comment "Map UUIDs to integers (anti-hallucination)". Models copy short numbers reliably and mangle long random strings. If you build anything where a model must refer back to records by id, steal this.
At our index, the file is 3,856 lines, with a sync Memory class and an async AsyncMemory class that mirror each other. main.py is one of mem0's top hotspots, which makes sense: every change to memory logic is made twice. Only 2 design decisions were recorded for mem0, both from code comments.
GPT Researcher: the model never picks a tool
GPT Researcher is the application layer, and it answers the "abstraction or overhead" question differently from the rest: in its default path, the model does not choose tools at all.
ResearchConductor.conduct_research in gpt_researcher/skills/researcher.py works like this. The model plans sub-queries (plan_research). The code runs them in parallel with asyncio.gather. For each one it calls the search retrievers you configured, scrapes the pages, and keeps the parts most similar to the question. The model then writes the report. Which search engines to call is a config setting, resolved by get_retriever in gpt_researcher/actions/retriever.py, a match statement with 21 cases from google and duckduckgo to arxiv, pubmed_central and mcp.
MCP is the one place the model gets a say, through _get_mcp_strategy, which can be "disabled", "fast" (run once with the original query, the default) or "deep" (run for every sub-query).
This is the smallest repo of the five, with the second-highest score and only 7 hotspots. I think the design helps, because when the code decides the control flow, there are fewer paths to test. The multi-agent version in multi_agents/ goes the other way and uses a LangGraph StateGraph with seven nodes (browser, planner, researcher, writer, fact_checker, visualizer, publisher), which is a nice example of an application leaning on the framework layer when it actually needs a graph.
Abstraction versus overhead
Counting steps, a tool call costs about seven in LangChain and six inside one function in ADK. Put Composio under LangChain and you add six more, plus a network hop. Every one of those steps has a reason I could find in the code: parallel calls, hooks for approval and retries, argument validation, version pinning, not retrying writes.
After reading all five, the overhead I'd worry about is how many repos you have to read when something goes wrong. ADK keeps its steps in one function. LangChain spreads them across three packages. Composio moves the most important ones to a server you can't read. GPT Researcher avoids most of them by not letting the model choose.
If your agent's job has a fixed shape, like research or a pipeline, GPT Researcher's approach is worth copying: let the model plan and write, and let code decide what to call. If the model really must choose tools, pick the framework whose tool path you can read in one sitting, because you will be reading it.
If you want to trace a path like this in your own stack, a quick way is to ask repowise's MCP server from Claude ("where does a tool call go after the model returns it?"), get the function names back, and then read each one on GitHub. That beats opening five browser tabs and guessing at folder names.
What the numbers can't tell you
- Health scores compare files inside each repo against the same rules. They don't say whether a framework is good to build on. LangChain's lower average is partly generated data files and provider adapters that churn because providers churn.
- The repos were indexed on different days, from 11 August (adk-python) to 6 October (langchain). I read the code on each repo's main branch today, so a few file paths may have moved since the snapshot.
- Step counts are my reading of the code. I did not run a benchmark or measure latency. A step that is a dictionary lookup costs nothing next to a model call.
- "Recorded decisions" counts what we could find written down. mem0's 2 means that few of its design decisions are written down in places we read.
How we measured
All five repos were indexed by repowise from their public GitHub repos. Code health, line counts, file counts, hotspots and decisions come from each repo's latest ready snapshot, shown on the repo pages linked above. Health is shown to one decimal and cut off without rounding (7.28 shows as 7.2), exactly as the repo pages show it. The tool-call paths come from reading each repo's source on its default branch on 6 October 2026; every function named above exists at the path given.
If you want to ask your own questions of any of these repos, you can connect repowise to Claude or ChatGPT and point it at the repo. Browsing any public repo page is free and needs no signup.