Memory OS for AI Agents: What It Is and Which Platforms Offer One
Of six agent memory projects, only MemOS calls itself an operating system. The label is not a capability, so compare these platforms on what each one does when a stored fact turns out to be wrong.
Six agent memory projects go on the shortlist, and they describe themselves with six different nouns. One is a memory operating system. The rest are a memory layer, a memory and context engine, an AI memory platform, a framework for temporal context graphs, and a harness for stateful agents. Reading all six projects' own READMEs on 9 October 2026, exactly one of them uses the operating system word.
The label is doing less work than it looks. Used strictly it can mean a policy for what occupies the context window, and used broadly it can mean a whole resource manager for several kinds of memory. Both are real designs, and the word alone does not tell you which one you are buying. One question separates these products much faster. What happens when a fact you stored last month turns out to be wrong today?
Only one of these projects calls itself an operating system
Each of these descriptions is the project's own wording, from its own repository, checked 9 October 2026.
| Project | How it describes itself |
|---|---|
| MemOS | "a Memory Operating System for LLMs and AI agents" |
| Mem0 | "The Memory Layer for Personalized AI" |
| Supermemory | "the memory and context layer for AI" |
| Cognee | "The Free Open-Source AI Memory Platform for Agents" |
| Graphiti (the open core of Zep) | "a framework for building and querying temporal context graphs for AI agents" |
| Letta | "Build stateful agents with memory that can learn and improve over time" |
"Memory OS" is a category label that readers and analysts apply from outside. Only MemTensor's MemOS claims it. Treating the term as a product tier, where the memory-OS tier does more than the memory-layer tier, will mislead you, because the vendors are not competing on that word.
The OS comparison borrows two different things
Two analogies travel under the same word, and they claim very different amounts.
The narrow one is paging. A model has a fixed context window, the material an agent might need is larger than that window, and a piece of software sits between the two deciding which facts are resident in the prompt, which get written out, and which get pulled back in. Paging between a fast tier and a slow one is recognizably virtual memory. That is the MemGPT line of thinking, and its creators now build Letta.
The broad one is resource management, and MemOS states it directly. Its documentation makes the comparison explicit, "just as an operating system schedules CPU, RAM, and files," and describes the system as one that "schedules, transforms, and governs multiple memory types." It names three. Parametric memory is "knowledge distilled into model weights," activation memory is "KV-caches and hidden states for inference reuse," and plaintext memory is "text, docs, graphs, vector chunks, user-visible facts." It documents a lifecycle in which a memory unit moves through states such as generated, activated, merged, archived, and frozen, plus a governance module covering access control, redaction, compliance, and audit trails.
The operating-system analogy does not establish how a product isolates agents or controls inference execution. Check access control, agent isolation, and the underlying inference scheduler separately from its memory-management policies. Preemption does exist lower down the stack, in the GPU and in inference servers, but that is a different layer from an agent memory policy and no memory product's label tells you about it.
So the label stretches from an eviction policy at one end to a resource manager with its own lifecycle and audit trail at the other. Which one a product means is the first thing to pin down, and the name will not tell you.
A vector store retrieves, a memory system decides what to keep
Three adjacent things get conflated with agent memory, and the distinction is mechanical rather than philosophical.
A vector store answers a query with similar text. It has no opinion about what should have been saved in the first place, when a saved item stopped being true, or what deserves space in a crowded prompt. You can build a memory system on top of one. The vector store is not the memory system.
A RAG pipeline retrieves material to ground a model's response. That material can include documents, conversation history, or observations generated during earlier work. A memory system adds policies for what persists, how it is updated, and when it is retrieved, and can use RAG as its retrieval path. We covered the relationship in agent memory vs RAG vs the context window.
A longer context window gives the agent more room in a single session. It does not carry anything into the next one. Every session that starts from an empty window re-derives context your team already paid to establish.
Six questions that separate a memory system from a key-value store
Use these when you read a platform's documentation. Each one asks whether the platform does the work or hands it back to you, and either answer can be the right one for your team.
- What triggers a write. Does the service decide which facts are worth keeping from an interaction, or does your code call an endpoint for each one? Automatic extraction removes work and removes control.
- What comes back, and in what order. Retrieval has to choose what fits in limited prompt space. Ranking by similarity alone will surface a chatty recent message over the architectural decision that matters.
- What happens on contradiction. The most damaging entry in an agent's memory is rarely an invented one. The damaging entry is a fact that was right in March. Ask which layer does the correcting, and whether a default read hands back the superseded entry alongside the current one.
- What gets forgotten. Is there a decay, expiry, or pruning policy, do you set it, and does anything remove a memory that nothing has read in six months?
- How memory is scoped. Per user, per session, per project, or per organization. Scoping decides what retrieval mixes together. Scoping is not access control, and treating it as access control is how one team's context ends up in another team's answer.
- Whether several agents can share one store safely. Concurrent writes, and some rule about who may read and write a shared scope.
The real question is when correction happens and what a default read returns
Mem0's managed platform and Graphiti both retain superseded facts as history. Compare when each system identifies a contradiction and whether a default query returns that history alongside the current fact. Do not assume the same retention policy across every product or deployment.
Mem0 separates the two layers, and reading only one of them is misleading. Its README describes an extraction algorithm introduced in April 2026 built on "Single-pass ADD-only extraction," one LLM call with no update or delete step, whose own wording is that "Memories accumulate; nothing is overwritten." That describes extraction. On Mem0's managed platform, Dream adds a lifecycle layer above it, where Supersede and Merge run as memories are added. An older fact is marked superseded and linked to the memory that replaced it, and a default read still returns that labelled history. Pass latest_only=true and you get active memories only. Those are Platform capabilities, so do not assume the open-source library behaves the same way.
Graphiti, the open-source framework at the core of Zep's context infrastructure, puts the same job in the graph. Its README describes "Explicit bi-temporal tracking with automatic fact invalidation." A superseded fact is invalidated rather than deleted, the temporal history is preserved, and you can ask what the graph believed at an earlier date. The cost is operational. Graphiti runs against a graph database you choose and operate, with Neo4j, FalkorDB, and Amazon Neptune listed as supported, and it needs an OpenAI API key by default for inference and embeddings.
So the question to carry into a trial is not whether a platform reconciles contradictions. It is whether correction happens at write time or at query time, whether your default read returns superseded entries alongside current ones, and what flag you need to get only what is true now. Get that wrong and your agent cites a fact the store already knows is stale.
What each platform is, in its own terms
A caveat on sourcing before the list. Everything below comes from each project's own repository, published package metadata, or official documentation. Benchmark figures published by a vendor about its own product are reported as exactly that, with the baseline named.
MemOS owns the framing most completely, and it is Apache 2.0 licensed. It exposes a REST API, available as a managed cloud service or self-hosted on Docker or uvicorn, plus plugins for several agent harnesses. It uses memory cubes as composable units for isolating and sharing knowledge bases. In MemOS's open-source tree-memory reorganizer, when enabled, an unresolved contradiction can trigger a timestamp-based fallback that keeps the newer memory and deletes the older node. This runs automatically following an add, but reorganization is disabled by default. MemOS publishes a table of benchmark results evaluated via OmniMemEval, including 88.83 on LoCoMo and 89.20 on LongMemEval. Those are vendor-reported, and no independent replication was available. Its README also claims production stability under high concurrency.
Letta carries the MemGPT lineage directly, and its repository still lists it as "Letta (f.k.a. MemGPT)." It has also moved. The letta package on PyPI is now Letta Code, a terminal coding agent, and states plainly that it "is the Letta Code CLI, not a Python SDK or API server." Python applications calling the API use letta-client instead. If you read an older comparison of Letta, check which product it meant. Letta Code describes agents that improve by rewriting their own context through memory blocks, and tracks all context including those blocks in git, syncable to a repository you choose. Keeping memory in git is an unusually legible design, because you can read and diff what the agent believes. Letta Cloud is the default and stores agent memory and conversations even when execution runs on your computer. For local agent state, select the local backend. The App Server quickstart uses letta server --backend local --listen ws://127.0.0.1:4500.
Mem0 is Apache 2.0, ships as a library, a self-hosted server, and a managed cloud platform, and fuses semantic, BM25 keyword, and entity matching at retrieval. Its published benchmark gains, including LoCoMo moving from 71.4 to 92.5, are measured against Mem0's own previous algorithm rather than a third-party system, and its README states the numbers reflect the managed platform and that open-source users should expect results that are "directionally similar" but not identical. Read those two qualifications together before using the figures to compare vendors.
Zep is the managed product, and Graphiti is its self-hosted open core. The split matters for evaluation, because Graphiti's temporal guarantees are the part you can read and run yourself, while the hosted engine, built-in users and threads, and enterprise support are the part you buy.
Supermemory extracts facts from conversations and states that it handles temporal changes, contradictions, and automatic forgetting, so the question to ask it is the same one as the others: when the correction lands, and what a default read returns. It scopes memory with projects, passed as a containerTag. It ships connectors for Google Drive, Gmail, Notion, OneDrive, and GitHub. A local single-binary option exists, though its README notes that URL ingestion still depends on a hosted reader service, so local does not mean fully disconnected.
Cognee is graph-centric and exposes remember, recall, improve, and forget as its operations, with forget being a first-class verb that most of this list lacks. It supports a single-Postgres memory stack, but the bundled Postgres graph store is a demo feature. Cognee also offers a licensed version it describes as production-ready, so check which implementation and license your deployment requires.
Where this fits for a team running coding agents
If the agents in question are Claude Code, Cursor, or Codex and the gap is repository context rather than conversation history, a conversational memory layer may not be the missing piece. Repository decisions often sit in pull requests, review threads, and the Slack conversation where an approach was settled, without ever reaching a document the agent retrieves. That is one failure mode among several, and retrieval and correction are worth ruling out too.
Writing that knowledge down and keeping it current is the job Dosu does. It connects to GitHub, Slack, and Notion, with Slack and Notion on a paid plan, drafts Documents from what happens in those Sources, and serves them to coding agents through the Dosu MCP Server. Monitors watch a Source, check relevant code changes against published Documents, decide whether an update is needed, and send a proposed edit to Review for the ones they identify as stale. Whether a person approves that edit is configurable, and Auto Accept Review is on by default, so turn it off in the Library's settings if you want every proposed Document update held for approval.
Dosu is not a general-purpose agent memory API and does not compete with these platforms on per-user personalization. If you are building memory into a product you ship to users, pick from the list above. If your engineering team's agents keep rediscovering decisions your team already made, the gap is documentation that maintains itself. We compared the API-shaped options in agent memory API providers, and sorted the MCP-based options by where the store lives in best MCP memory servers.
FAQ
Is a memory OS just RAG with better branding?
They are not alternatives. RAG is a retrieval path, and it can run over documents, conversation history, or observations generated during earlier work. A memory system adds the policies around that path: what gets written down, what is evicted from a full context window, what happens when a stored fact is contradicted, and what a default read returns. Many memory products use RAG underneath. A product that only retrieves, with none of those policies, is a retrieval system whatever it is called.
Does a longer context window remove the need for one?
Not for anything that has to survive the session. A larger window helps within one conversation and carries nothing into the next. If your agents start every run with an empty window, the cost is paid again each time, and the fix is persistence rather than a bigger window.
Does memory scoping give us access control?
No, and conflating them is the common mistake. A scope tells retrieval which memories belong together. Access control decides who is allowed to read them. If two agents can both request the same scope, something still has to check whether each one is entitled to it. Verify that separately, and test it by revoking access and then querying through the real agent path.
Which of these should we self-host?
Graphiti and Cognee are built to be run by you, MemOS and Mem0 publish both self-hosted and managed paths, and Letta defaults to its cloud. Supermemory's local single-binary server is built for individual developers, is free within its lite license limit, and supports one organization with one API key on one machine. Its server source is not public. Enterprise adds organization-wide access controls and managed scaling. Local URL ingestion still uses a hosted reader service. Two things decide it. The first is who operates the graph or vector database underneath and who is on call when it needs an upgrade. The second is licensing, which is not a formality here: Cognee's single-Postgres graph store is a demo feature, with the production-ready implementation offered as a licensed product. Check which implementation your deployment needs. Also separate running the process from storing the state. With Letta, running letta server on your machine still keeps agent memory in Letta Cloud unless you select the local backend.
Before you pick one
Run one test against your own data before you compare feature tables. Store a fact that is true today, such as who owns a service or which queue library you use. Change it. Then ask the agent the same question and see what comes back. Whether you get the new answer, the old one, or both tells you more about a platform than its category label does.
If a repository decision is missing from the context your agents retrieve, start by documenting it. Create a Library and connect that repository, publish a Document for the service your agents ask about most, and enable its Monitor. Dosu checks relevant code changes and proposes updates to Documents it identifies as stale. Turn Auto Accept Review off in the Library's settings if you want each proposed Document update held for approval. It is on by default. Private repositories need a paid plan.
NOTE
Product descriptions here are quoted from each project's own repository, published package metadata, or official documentation, checked 9 October 2026. Benchmark figures are vendor-reported, measured by the vendor against the baseline named in each case. Vendors change these systems often, so confirm any behavior you intend to rely on against current documentation before you integrate.