Vector Database vs Knowledge Graph vs Files for Agent Memory
The substrate decides how a memory comes back, not whether it is still true. Compare vector stores, knowledge graphs, and version-controlled files on the write decision, retrieval, and what happens the day a stored fact stops being true.
Someone on your team has been asked to give the agents long-term memory, and the first argument is about where to put it. A vector database, because that is the reflex the RAG years left behind. A knowledge graph, because the facts have relationships and relationships change. Or plain files in the repository, because the agent reads those anyway and nobody has to operate a new service.
The three get compared on retrieval quality, which is the smaller half of the decision. What separates them once something is in production is narrower and less discussed. Two questions decide it: what makes a memory get written, and what happens to a stored fact on the day it stops being true.
What the substrate settles, and what it leaves you
A memory system has four moving parts. What a record looks like, what decides to write one, how a record gets back into the context window, and what happens when the world moves on. Our engineering team used that same decomposition reading through the agent-native memory literature in the building blocks of agent memory.
Picking a substrate settles the first and the third. It leaves you the second and the fourth, and those are where memory systems rot.
A vector store holds an embedding of a passage next to the passage and some metadata, and returns the records whose embeddings sit nearest the query. A knowledge graph holds entities and the facts that join them, so a lookup follows relationships rather than resemblance. Files hold text in paths the agent loads at startup or opens when it needs them. The Claude Code documentation is blunt that each session "begins with a fresh context window" and that the files it recognizes are what carry anything across.
How the three compare
| Substrate | What a record is | How it comes back | When a fact stops being true |
|---|---|---|---|
| Vector store | An embedding, the source passage, and metadata | Nearest neighbours to the embedded query, cut at a limit you set | Nothing in the index marks the old record superseded. The layer above has to delete, filter, or outrank it |
| Knowledge graph | Entity nodes joined by a fact edge, often with the source text retained | Graph traversal, usually fused with semantic and keyword search | A temporal graph can end the edge's validity and keep it when the dates allow. Retrieval still has to filter by period to act on it |
| Files in version control | Readable text in a path the agent loads | A bounded set loaded at session start, with subdirectory, path-scoped and topic files read on demand | Nothing in the file notices. It keeps asserting the claim until someone or something revises it |
Vector stores recall well and hold no opinion about truth
The vector database is not the part that decides anything. Pinecone, Qdrant, Chroma, and Weaviate store vectors and search them. The write policy lives in the memory layer above, and that layer's choices are what you are buying.
Those choices differ more than the category name suggests. Mem0's README describes its current extraction path as "Single-pass ADD-only extraction -- one LLM call, no UPDATE/DELETE," and states that memories accumulate and nothing is overwritten. Append-only is a reasonable design, and also a decision to handle supersession at read time rather than write time, which you inherit along with it.
Retrieval embeds the question and returns the nearest records. Under a tight context budget the consequence is mechanical: several near-duplicate memories of the same fact can fill the slots you allowed, because they are all close to the query and nothing in the index knows they say the same thing.
Contradiction is the sharper limit. Suppose a record from January says the service runs on MySQL and a record from March says it runs on PostgreSQL. Both sit near a question about the database, and nearest-neighbour ranking is a statement about wording, not about which one is current. Recency filters and metadata predicates can fix the specific case, but the application has to supply them. Vector proximity carries no notion of a fact having been replaced.
One more operational cost is easy to miss at design time. Vectors from two different embedding models are not comparable, so changing the embedding model means re-embedding the corpus or running two indexes side by side through the migration.
Vector stores fit an agent that needs fuzzy recall over a large body of loosely structured text, where the application above the store is prepared to own the truth question.
A knowledge graph can let a fact expire
A graph record is two entity nodes and the fact edge between them. Graphiti, the open-source engine behind Zep, also keeps the original episode that produced each extracted fact, so an edge can be traced back to the text it came from.
Writes cost more here, and the cost is structural rather than incidental. Each episode goes through entity and relationship extraction, embedding and search against what is stored, and a resolution step that decides whether this is a new entity or one you have seen before. Several model calls run before anything lands. The Zep team's paper, Zep: A Temporal Knowledge Graph Architecture for Agent Memory, sets out that pipeline. You are also running a graph database underneath. Graphiti's README lists Neo4j, FalkorDB, and Amazon Neptune among its supported backends.
What a temporal graph can add is a superseded fact kept rather than overwritten. Graphiti uses a bitemporal model, recording when a fact held in the world and when the system learned or invalidated it. Its resolution pipeline identifies contradictory facts, then checks their dates. When the validity windows overlap and both start dates exist in the required order, it can end the older edge's validity and retain the edge as history. Missing or equal start dates can leave a contradiction unresolved. The January MySQL edge can therefore acquire an end date when the March PostgreSQL fact arrives, but that outcome depends on the extracted dates as well as the detected contradiction. Graphiti's validity filters default to unset, so retrieval must still filter by period or pass the dates to the agent. The write path and the read path both need checking.
Retrieval is hybrid inside the engine. Graphiti's README describes combining semantic embeddings, keyword search via BM25, and graph traversal. The mature graph implementation is using vectors too, which takes some of the air out of the versus.
The failure mode moves to the write path. Temporal fields preserve history only when extraction identified the entity and the relationship correctly in the first place, and a wrong edge is now a wrong edge with a confident validity window. Schema or ontology upkeep is ongoing work. Benchmark results in this category are mostly vendor-published, so treat them as claims to reproduce on your own traces rather than as settled findings.
Temporal graphs fit agents reasoning over relationships that change, especially when someone will eventually ask what the system believed last quarter.
Files combine startup context with on-demand recall
Files can provide both a predictable startup set and notes the agent reads when needed. Committed instruction files are diffable, reviewable in a pull request, and travel through version control with the code they describe. Auto memory loads a bounded index at startup and reads topic files on demand.
Two write paths run here with different authors, and the Claude Code documentation keeps them separate. You write CLAUDE.md or AGENTS.md, holding instructions and rules. Claude writes auto memory, holding learnings and patterns drawn from your corrections. Both are loaded at the start of every conversation.
Determinism buys a hard ceiling. For auto memory, "the first 200 lines of MEMORY.md, or the first 25KB, whichever comes first, are loaded at the start of every conversation," and content past that threshold is not loaded at session start. For instruction files the guidance is to target under 200 lines, Claude Code loads a file up to 4 MiB in full and skips a larger one, and splitting into imports does not help, because imported files also load at launch. Two mechanisms do load conditionally. A CLAUDE.md in a subdirectory is read when Claude opens a file in that directory, and a rule under .claude/rules/ carrying a paths pattern loads only when Claude works on a matching file. So the ceiling binds on the always-on set rather than on everything you have written, and scoping is the lever that keeps the always-on set small.
The reliability limit is stated by the vendor rather than inferred. Claude treats these files "as context, not enforced configuration," there is "no guarantee of strict compliance," and when two instructions contradict each other Claude "may pick one arbitrarily." Claude can revise or delete its own memory files, so the write path is not limited to people. What file storage does not do is connect a source change to the notes it makes stale. A maintenance process still has to notice the change, check the affected guidance, and update it, whether that process is a person or an agent.
The limit that surprises teams is locality. Auto memory "is machine-local" and its files "are not shared across machines or cloud environments." A committed CLAUDE.md or AGENTS.md reaches the whole team through git. What one developer's agent worked out for itself does not reach anybody else's agent.
Files fit compact, stable conventions that a named person is willing to keep current.
Pick the failure you can live with
The three substrates fail in ways that are not interchangeable. Which failure your team can absorb is the decision.
A vector store returns something plausible when the truth has changed, and it does so without any signal that the record is stale. Detection is on you.
A temporal graph can record supersession and keep the history, and it charges you a multi-call write pipeline plus schema maintenance for that. Whether an older edge expires depends on the extracted dates, and the retrieval filter is still yours to apply, so a superseded fact can reach the agent from either end. An extraction error becomes a confidently dated wrong fact.
Files are the easiest to audit and the easiest to let rot. Content past the startup budget is not loaded at session start, so it reaches the agent only if the agent goes looking, and auto memory stays on one machine.
The substrates can work together
Graphiti combines embeddings, keyword search, and traversal. Mem0 combines semantic, BM25, and entity signals. Those implementations show that retrieval techniques can coexist inside one memory layer.
A team can also combine a recall store with version-controlled instructions. The store can hold material retrieved for a particular task, while instruction files carry conventions the team wants to review and share. Whether that split is useful depends on what needs to load consistently and what needs selective recall.
The part the substrate does not decide
Run the test on your own system. Take the last ten things it stored and ask, for each one, what would have to happen in the world for this to become false, and what in your stack would notice.
The second half is the one to sit with. A vector store has no opinion about whether a stored record is still true. A temporal graph can end a fact's validity when something tells it the fact changed, and something still has to tell it. A file holds what was written until someone or something revises it.
For repository knowledge, pull-request changes provide a concrete signal to check. A Monitor on a connected Source compares those changes with relevant published Documents and proposes targeted edits. Reviewers can Accept, Edit, or Decline proposals in the pull request or Review page. Library settings control whether pending Document updates are accepted automatically on merge or wait for manual approval. Agents can retrieve published knowledge through the Dosu MCP Server's read_knowledge tool in Claude Code, Cursor, or VS Code Copilot.
Dosu is not a memory API for an agent product you are shipping, and it does not replace one. If what your agents keep losing is user state inside your application, a memory layer is the right purchase and the substrate comparison above is the right way to choose between them. If what they keep losing is the repository, the missing piece is something that notices the merge, not a different store.
So pick the substrate on two questions rather than on the category argument. Who or what is allowed to decide a memory gets written, and how much staleness the agent can carry before the answers stop being worth having. For a vendor-by-vendor view of the memory layers themselves, we compared them in the best AI agent memory API providers, and for the adjacent question of whether you need memory at all we compared agent memory against RAG and bigger context windows.
FAQ
Should I use a vector database or a knowledge graph for agent memory?
Use a vector store when recall is about resemblance across a lot of loosely structured text and the application above it can own deletion and filtering. Use a knowledge graph when the queries follow relationships, or when facts change and you need to tell the current one from the one it replaced. Mature implementations of both end up fusing semantic and keyword retrieval, so the gap is less about search quality than about whether supersession is handled for you.
What is a temporal knowledge graph?
A knowledge graph that attaches time information to facts or relationships. Graphiti uses a bitemporal model, recording both when a fact held in the world and when the system learned or invalidated it. On a detected contradiction it can end the earlier edge's validity and keep the edge, though that depends on the extracted start dates. The end date makes a point-in-time question answerable rather than automatic, because retrieval still has to filter by period.
Can plain files replace a vector database for agent memory?
Files can work for compact conventions and for notes read on demand. Claude Code loads a bounded MEMORY.md index at startup, then reads topic files when needed. The startup limit does not cap all file-backed memory. As the collection grows, check whether the agent can find relevant files reliably. Auto memory is machine-local, while committed instruction files can reach the team through git.
What happens when a stored fact becomes false?
In a vector store, nothing, until the layer above deletes, filters, or outranks the old record. In a temporal graph, the edge can be given an end date and kept when the extracted dates allow, though retrieval still has to filter by period for the agent to see the current fact. In a file, the claim stands until a person or an agent revises it.
Audit what your agents are storing against what the code says
Pick one repository and one claim your agents rely on, such as which queue a service publishes to or which auth path an API uses. Find where the claim lives, then compare it with the current implementation and the pull request that last changed the relevant behavior. If the claim is now wrong and nothing flagged it, you have found a gap in the maintenance process. A claim that predates a merge may still be correct.
Dosu closes that loop for repository knowledge: it captures context from pull requests and connected Sources, proposes an update when the code changes, and serves the published result to your coding agents over MCP. Setup runs from one command:
curl -fsSL https://cli.dosu.dev/install | sh
Then connect your first repo to Dosu and let the next session start from something that is still true.