Agent memory, three ways: how mem0, cognee and hindsight are built

Swati Ahuja10 min read

mem0 vs cognee · agent memory · mem0 architecture · cognee knowledge graph · hindsight agent memory · llm memory layer

On this page

We all reach for a memory layer the first time an agent forgets something the user said two messages ago. I opened three popular ones expecting three versions of the same feature, and found three different ideas of what "memory" even means.

mem0 treats memory as a list of short facts about a user. cognee treats it as a knowledge graph built from your documents. hindsight treats it as something that should change its mind over time: it keeps the raw facts, forms higher-level observations from them, and can reflect on them when asked.

So I read the code that runs when an agent saves something (the write path) and when it asks for something back (the read path), in all three. Below are both paths for each project, with the files and functions, and then what the code health numbers do and don't say.

The three at a glance

mem0cogneehindsight
GitHub stars66,65331,44646,150
What "a memory" isA short fact with an embeddingNodes and edges in a knowledge graph, plus chunks and summariesA fact, later folded into observations and mental models
LLM calls on writeOne extraction call per add()Several per chunk (graph extraction and summary)Fact extraction on retain, more in background consolidation
StorageAny of 24 vector stores, plus a SQLite history tableA graph database (Kuzu, Neo4j, Neptune, Postgres and others) plus a vector store (LanceDB or pgvector)Postgres (or Oracle) for everything
Read pathVector + keyword (BM25) + entity boost, one ranked listVector search to find graph triplets, then an LLM writes the answerFour retrieval arms fused, reranked, optional reflect
Lines of code (non-blank, non-comment)228,481254,682605,089
Code health (out of 10)7.16.66.0

Scroll the table sideways to see every column.

The line counts include everything in each repo: SDKs in other languages, integrations, frontends, docs code and tests. The memory logic itself is a small part of each, which matters later when we get to the scores.

mem0: a list of facts, written by one LLM call

Everything in mem0's Python library goes through one class, Memory, in mem0/memory/main.py. The file is 3,856 lines and holds both the sync Memory and the async AsyncMemory.

Write path

When you call add() with some messages, the work happens in _add_to_vector_store. The code comments call it a "phased batch pipeline", and it runs in this order:

  1. Load the last 10 messages of the session for context.
  2. Embed the new messages and search the vector store for the 10 most similar memories that already exist.
  3. Make one LLM call. The system prompt (ADDITIVE_EXTRACTION_PROMPT) tells the model its "sole operation is ADD": read the new messages, look at what is already stored, and return only facts that are new.
  4. Embed the new facts in one batch.
  5. Drop exact duplicates by hashing each fact's text (MD5) and comparing it against the hashes of what is already stored.
  6. Insert the rest into the vector store and write an "ADD" row to a history table.
  7. Pull entities (names, places, things) out of each fact and link them in a separate entity collection.

One detail made me happy in a very specific engineering way. Before the LLM sees the existing memories, their UUIDs are swapped for short numbers ("0", "1", "2"), and the code comment calls this "anti-hallucination", because models copy short numbers back far more reliably than 36-character IDs.

The bigger surprise was what's missing. In mem0 v1.0.11, add() asked the LLM to choose ADD, UPDATE or DELETE for each memory on every write (main.py, lines 591 to 665), but in the commit we indexed the automatic path only adds. update() and delete() still exist, and you call them yourself. There is also no mem0/graphs folder any more, because entity linking inside the vector store does that job now, so if your picture of mem0 is a year old, it's worth re-reading.

Read path

search() calls _search_vector_store, which runs these lookups and merges them:

  • a semantic search on the embedding (it asks for 4 times the results you wanted, at least 60),
  • a keyword search (BM25, a classic ranking formula that scores word overlap) on a lemmatized copy of the text, where "lemmatized" means words reduced to their base form,
  • an entity boost, so memories linked to the same entities as your query score higher.

score_and_rank combines them into one list, and an optional reranker can reorder it at the end.

The design is easy to state: keep the unit of memory small and flat, spend one LLM call per write, and make reads cheap. You can see it on the mem0 architecture page.

cognee: build a knowledge graph, then search the graph

cognee splits memory into two verbs. cognee.add() ingests raw data, cognee.cognify() turns it into a knowledge graph, and then cognee.search() reads from it.

Write path

cognify() in cognee/api/v1/cognify/cognify.py delegates the work. It builds a list of tasks in get_default_tasks and hands them to a pipeline runner:

  1. classify_documents: turn raw items into typed documents.
  2. extract_chunks_from_documents: split them into chunks sized to the LLM's limit.
  3. extract_graph_and_summarize: ask the LLM for entities and relationships in each chunk, and for a summary.
  4. add_data_points: write nodes and edges to the graph database, and their embeddings to the vector database.
  5. extract_dlt_fk_edges: add edges from foreign keys when the data came from a database.

You can swap the graph schema (graph_model), pass an ontology file, or run a different task list for time-aware data (get_temporal_tasks). Storage sits behind two interfaces, graph_db_interface.py and vector_db_interface.py. At the commit we indexed, the graph side has adapters for Kuzu, Ladybug, Neo4j, Neptune and Postgres, and the vector side has LanceDB and pgvector.

Read path

search() takes a query_type, and SearchType has 17 of them, ranging from plain CHUNKS through GRAPH_COMPLETION_COT (chain of thought over the graph) to CYPHER (run a graph query directly). The default is GRAPH_COMPLETION. Its retriever, graph_completion_retriever.py, embeds the query, runs a vector search over node and edge collections, and scores graph triplets (subject, relation, object) by distance in brute_force_triplet_search. It then turns the best triplets into text and asks an LLM to answer from them.

cognee's bet is that memory is structure, and that the structure is worth paying for at write time. That costs more LLM calls per document than mem0, and in return you can ask "how are these two things connected", which a flat list of facts answers badly. The cognee architecture page shows how much of the repo is these adapters: the file that changed most in our window is the LanceDB adapter.

hindsight: keep the facts, form observations, reflect

hindsight's README says it wants agents to "learn, not just remember". The code backs that up with three verbs, retain, recall and reflect, and all of them live in one class, MemoryEngine, in hindsight_api/engine/memory_engine.py.

Write path

retain_async and retain_batch_async hand off to the retain/ package. There, fact_extraction.py pulls facts from the input with an LLM, entity_processing.py resolves entities, link_creation.py links facts to each other, and fact_storage.py writes them to Postgres. Memories live in "banks", one per agent or user, each with its own settings.

Then a second write happens later, and that second write interested me most. consolidation/consolidator.py runs as a background job after retain, and its docstring explains it well: for each new fact it either creates a new observation or updates an existing one "when new evidence supports/contradicts/refines" it. Each observation keeps a proof_count, the IDs of the facts behind it, and a history of how it changed.

On top of that sit "mental models". You define a question once ("what are this user's preferences?"), and hindsight writes the answer, stores it and rewrites it as the bank learns more, so reading one is a plain database read.

Read path

recall_async runs four retrieval arms: semantic, BM25, graph (following links between facts), and temporal (for questions about time). Each arm is capped so one can't crowd out the rest, and then they are merged with reciprocal rank fusion in search/fusion.py. The method is simple: each result scores 1 / (60 + its rank) in every list it appears in, and the scores add up. A cross-encoder reranker then orders the final list, and reflect_async goes a step further, with an agent in reflect/agent.py that uses tools to search memory and reason about it before answering.

hindsight wants memory to hold beliefs with evidence behind them, and to revise those beliefs. It's the most ambitious design of the three, and the code size shows it: memory_engine.py is 23,543 lines at the commit we indexed. The HTTP layer next to it, api/http.py, has one function, _register_routes, that is 5,209 lines long because it declares every endpoint inline.

The health numbers

repowise's health score (1 to 10) rates how likely a file is to cause bugs and how hard it is to change, using static checks plus git history; the repo score is the average across files, weighted by lines of code. The per-repo detail below comes from the pages linked above.

mem0cogneehindsight
Code health7.16.66.0
Hotspots (files that change often and are complex)94298500
File that changed most in our windowmem0/memory/main.py (41 commits)vector/lancedb/LanceDBAdapter.py (61)engine/memory_engine.py (353)
Longest functionregisterCliCommands, 1,473 linesDatasetDetailPage, 1,160 lines_register_routes, 5,209 lines

Scroll the table sideways to see every column.

The longest functions mostly sit outside the memory logic. mem0's is in integrations/openclaw/cli/commands.ts, a CLI for an integration that lives outside the core library, and cognee's is a React page in its frontend. Only hindsight's sits in the main service, and it is route wiring.

Each project's busiest file is the one its design runs through. In hindsight that is memory_engine.py: in the window we looked at (the last 2,000 commits, starting 4 May 2026), 353 commits touched it, and our classifier labelled 180 of them as bug fixes and 119 as features. mem0's main.py shows the same pattern at a smaller scale, with 25 of its 41 commits labelled as fixes. When one file holds the main class, every fix to retain, recall or reflect lands there, and if I maintained hindsight I would read that as a map of where to split the file.

The scores follow scope more than craft. mem0 does the least per write and has the smallest core, while hindsight does the most, runs the most moving parts, and carries a 605,089-line repo with clients in several languages. Across the 3,126 public repos we have scored, health drops as repos grow, so a bigger, more ambitious project starting lower is what you'd expect.

Picking one

My take, for what it's worth:

  • If you want user preferences and facts recalled cheaply in a chat product, mem0's shape fits, with one LLM call per write, simple reads, and almost any vector store you like.
  • If your agent needs to answer questions about how things in your documents relate, cognee's graph is the reason to pay the extra write cost.
  • If you want memory that revises itself as evidence comes in, and you're happy running Postgres and background jobs, hindsight is the only one of the three built around that idea.

Limits of these numbers

  • Code health doesn't measure retrieval quality, so none of these numbers say which project remembers better. For that you need a benchmark on your own data.
  • The snapshots are from different dates: cognee on 19 July, mem0 on 18 August, hindsight on 27 September 2026. All three move fast, so details may already have changed.
  • "Bug fix" means the commit message contains a word such as fix, bug, patch, revert, regression or crash. Some commits with those words aren't fixes, and some fixes use none of them.
  • Lines of code count everything in the repo, including tests and other-language SDKs, so they overstate the size of the memory logic itself.

How we measured

We indexed each repo with repowise at the commits linked above and read the latest snapshot's numbers from our database on 6 October 2026. Health scores are shown with one decimal, the way the repo pages show them, and the code descriptions come from reading the source at the same commits on GitHub. If you want to ask your own questions about any of these codebases, you can connect the repowise MCP server to Claude or ChatGPT and point it at the repo.

That's three architectures for a feature request most of us would write in one line: remember what the user said.