Episodic vs Semantic vs Procedural Memory in AI Agents
Three memory types, borrowed from two different cognitive-psychology taxonomies. What each one holds, what writes it, how it goes stale, and how to decide which store a new observation belongs in.
Give an agent a memory layer and one early decision is what kind of thing it just learned. A session transcript, a fact about the codebase, and a deployment runbook need different metadata and different update rules. They can share a storage backend, but retrieval has to preserve their type, provenance, and validity, or a settled convention looks the same as something one developer tried once on a Tuesday.
The vocabulary the field uses for the three kinds is episodic, semantic, and procedural memory. The useful version of that taxonomy is not a definition exercise. It is a rule about when you classify: at write time, while you still know where the observation came from, rather than at read time, when all you have is a chunk of text that resembles the query.
Three kinds, because undifferentiated records answer the wrong questions
A bigger context window does not solve this. A window is a buffer for the current run. It holds more conversation, and it still has no opinion about which parts of that conversation should outlive the session, which should harden into a fact, and which describe a method worth repeating.
The cost of getting the classification wrong shows up later and is hard to trace. Write raw transcripts in as general knowledge and retrieval starts returning four near-duplicate passages where you wanted one current fact, each one burning context. Promote a single successful decision into a procedure and the agent repeats it in conditions that never justified it.
One thing worth knowing about the three-way split: it is stitched together from two different taxonomies. Endel Tulving's 1972 proposal divided long-term declarative memory into episodic, for events experienced at a particular time and place, and semantic, for general facts and concepts. Procedural memory comes from the other branch. In Larry Squire's taxonomy, declarative memory covers Tulving's two, and procedural sits under non-declarative memory alongside priming and conditioning. So episodic and semantic are siblings in a way that procedural is not.
That matters more than a footnote, because the AI usage has drifted further still. Procedural memory in neuroscience means learned skills acquired without conscious recall. In agent architectures it covers several representations at once: instructions in a prompt, an executable skill, the agent's own code, and the model weights themselves. This article stays with the representations a team can inspect and revise. Treat the three names as engineering labels borrowed from cognitive science, not as a claim that the system models human cognition.
Episodic memory holds what happened, with its time attached
An episodic entry records a specific event the agent took part in or observed: what happened, when, in which session, and against which artifacts. "The reviewer rejected the migration on Tuesday because it required downtime" is an episode. The time and the circumstances are part of the meaning, not metadata you can drop.
Write an episode when the event might later explain a decision. Keep the temporal context on the record and resist generalizing at write time, because the generalization is a separate step with its own evidence bar. Read episodes when the question depends on sequence, recency, or provenance, such as who approved the last deploy or what the agent tried before the current attempt.
Similarity search alone handles this badly. "What broke yesterday" and "what broke last month" embed almost identically while asking for different events, so retrieval that ranks only on meaning cannot reliably separate them. Episodic retrieval wants filters on time and session, plus explicit links between related events, rather than ranking on resemblance.
Episodes do not usually become false. They become superseded. The migration rejected on Tuesday may be approved in March once the downtime requirement goes away, and both events stay true as history. What changes is which one describes the current state. Keep both, record the order, and link the correction back to the original rather than overwriting it. Then archive aggressively, because an agent that carries every episode it has ever recorded into context is paying tokens for history it will not use.
Semantic memory holds what is true now
A semantic entry is a claim the agent treats as true independent of the occasion it was learned on. "Migrations live in supabase/migrations." "Pull requests need two approvals." The entry says what was concluded, not when the conclusion was reached.
What separates a semantic entry from an episode is the evidence bar, not a required route. A convention can be extracted from repeated episodes, read directly from the repository or an authoritative document, or supplied by an explicit correction. One session showing a developer putting a migration in a particular directory is an episode, and it is thin evidence on its own. Keep the source, scope, and validity information with the claim so it can be checked when disputed. A single observed action should not become a general convention without supporting evidence.
The payoff at read time is compression. An agent editing a migration needs the convention, not the six sessions that established it, and a compact claim costs a fraction of the context that the underlying transcripts would.
The failure mode here is the one most worth designing against. When a system writes whole sessions into semantic memory as though they were facts, abandoned ideas and half-finished reasoning end up competing with settled knowledge at the same rank. Retrieval returns several overlapping passages, none of them marked as current, and the agent has no basis for choosing. Semantic memory also goes stale on its own schedule whenever the code or the policy underneath it changes, which is why it needs supersession rules and revalidation rather than only an append path.
Procedural memory holds how the work gets done here
Procedural memory holds the method. Across agent systems that spans prompt instructions, executable skills, agent code, and model weights, which is why it is the one kind every agent already has whether or not anyone designed it. The part a team can read and revise is an explicit instruction set or runbook: the steps, their sequence, the preconditions that make them apply, and what failure looks like. Deploying a service, recovering from a specific build failure, cutting a release.
Order and preconditions are what a procedure has to carry with it. Similarity search can return a whole runbook or skill intact, so the retrieval method is not the problem by itself. The failure is returning disconnected fragments: chunk a runbook for embedding and you can get passages that each resemble the query while the dependency that the migration runs before the application deploy sits in none of them. Retrieve a procedure as a unit, selected by goal and precondition, rather than as whichever chunks rank highest.
The write bar should be higher than for the other two kinds. One successful run is evidence about one occasion. Promotion into a procedure should wait for repeated success under the intended conditions, or for a human approving a runbook. The failure this prevents is specific and common: an agent watches a test get skipped during an urgent production patch, records skipping the test as the deployment method, and then skips it during routine work. The exceptional circumstance that justified the shortcut is exactly the part that does not survive into the stored sequence unless preconditions are stored with it.
The three at a glance
| Type | What it holds | What writes it | How it goes stale |
|---|---|---|---|
| Episodic | Events tied to a time and session | Direct observation during a run | Later events supersede it, or it stops being relevant |
| Semantic | Claims treated as currently true | Evidence fit to the claim: repeated episodes, an authoritative source, or a correction | The code, policy, or environment underneath it changes |
| Procedural | A method, from prompt instructions and skills to code, with its preconditions | Designer-authored to begin with, then repeated success or an approved runbook | Tools, interfaces, or accepted practice change |
Routing an observation at write time
Classify on the way in. Three questions, in order:
Is the meaning tied to a specific moment? If the timestamp, participants, and circumstances change what the record means, it is an episode. Keep all of them.
Does it support a claim that holds outside that moment? If so, extract the claim into semantic memory, bounded to its real scope, with the source episode linked and a validity signal attached. "Driver v4 conflicts with the current connection pool" is a fact. "The deploy failed Tuesday" is the episode it came from.
Is it a repeatable method with conditions? Then it is a procedure, once something beyond a single run supports it. Store prerequisites and expected results alongside the steps.
One observation often produces more than one record. A failed deployment can be kept as an episode, yield a compatibility constraint as a semantic fact, and much later contribute to an upgrade procedure once the corrected sequence has worked several times. Linking those records preserves the evidence trail without treating every detail as equally reusable.
When you cannot tell, keep it episodic. An episode can be promoted later when evidence accumulates. A premature fact or procedure has to be found and corrected, usually after it has already produced a wrong answer.
Where this is not the same as RAG
RAG retrieves from an index at query time, usually by similarity. The line between the two is not who wrote the documents, because you can index an agent's own session notes as easily as a handbook. It is that RAG does not own the write path. Something else decides what enters the index and when an entry stops being true. Memory owns that decision, which is what lets it record that an event happened, mark one fact as superseding another, and keep a procedure's steps in their order. A memory layer will use retrieval internally, and retrieval on its own settles none of that.
That is why retrieval without memory-specific structure produces the same handful of symptoms: temporal questions that similarity cannot separate, procedures fragmented across chunks, relationship chains the ranking cannot follow, and contradictions returned side by side with nothing marking which one won.
Both still earn their place. RAG is the right tool for large reference collections that change slowly. Memory is the right tool for events, for the current facts distilled from them, and for methods learned by doing the work. We compared those mechanisms directly in agent memory vs RAG vs bigger context windows.
What the taxonomy tells you about a memory tool
Read the three types as evaluation lenses rather than product categories. No tool maps one-to-one onto them, and the architecture a tool picks constrains the classification problem without solving it.
Memory APIs such as Mem0 and Zep implement part of the write path themselves. Mem0 extracts facts from messages by default, with raw storage available when inference is turned off. Zep stores what you send as an episode and derives entities, facts, and temporal bounds from it. Your application still chooses what to send and what to do with the results, so evaluate a service's extraction, provenance, and update behavior alongside your own policy rather than treating it as a passive record store. Ordered procedures still need a representation that preserves their steps and preconditions.
Knowledge-graph approaches, including Cognee and temporal graph designs, organize around entities and the relationships between them. A temporal graph can hold when a relationship applied, which is what makes the episodic-to-semantic link expressible and supersession representable. Execution order is still something you have to model deliberately. We went through those tradeoffs in vector database vs knowledge graph vs files for agent memory.
Tiered context paging solves a different problem again. MemGPT, and the Letta line of work that followed it, treats the context window like RAM and external storage like disk, moving material in and out under the model's own control. That governs what a memory costs you at inference time. It says nothing about what the memory means or when it expires, so classification stays yours either way.
Codebase memory belongs to the repository, not the laptop
For a team running coding agents, the semantic and procedural stores have a property worth noticing: almost nothing in them is personal. Which queue a service publishes to, how releases get cut, which auth path the API uses. These are facts about a shared system, and they do not vary by who opened the editor.
When every developer's agent keeps its own private store, that shared truth fragments. One agent learns the API constraint was lifted, another is still working from the old one, and a third has quietly recorded a different release sequence. Each store is internally consistent and they disagree with each other, which is harder to debug than a single store that is out of date. Episodic memory can reasonably stay local, since it is a record of one session. The other two want to be shared, versioned, and revised in one place when the system changes.
The other half is where the evidence comes from. Codebase facts and procedures are produced continuously by normal engineering work, in pull requests, review threads, and discussions, and asking people to restate them in a wiki is how they go stale. Capture from the work itself, then keep a human in the loop on promotion, so a temporary workaround does not become documented team policy.
FAQ
What is agent memory architecture?
The set of decisions about what an agent persists, where each kind of record goes, and what causes a record to be written, updated, or retired. The storage backend is the smaller half of it. The part that determines behavior is the write policy: what gets classified as an event, what gets promoted to a durable fact, and what evidence is required before something becomes a procedure.
Is agent memory the same thing as context engineering?
They overlap but answer different questions. Context engineering is about what goes into the window for a given run and at what cost. Memory is about what persists between runs and whether it is still true. You can do one well and the other badly. An agent with a well-packed context and no memory starts cold every session, and an agent with a large memory and no context discipline pays for records it never needed.
Should short-term and long-term memory map onto these three types?
They are different axes. Short-term and long-term describe how long something survives and whether it sits in the window or in a store. Episodic, semantic, and procedural describe what the information is. A semantic fact can be in the window right now, and an episode from this morning can already be archived. Classify by kind on the write path and handle duration separately with paging and retention.
How do you stop agent memory from going stale?
Decide, per type, what invalidates a record. Episodes are superseded rather than falsified, so order them and keep the history. Semantic facts need supersession rules and periodic revalidation against the source they describe. Procedures need versioning and preconditions, so an old method stays readable as history without being selected for new work. A store with only an append path accumulates contradictions by design.
Check what your agents would write down
Pick one repository and one claim your agents rely on, such as which service owns a queue or which directory migrations belong in. Find where that claim currently lives for each engineer: a CLAUDE.md, a private memory store, a doc nobody has opened in six months. Then compare it against the pull request that last changed the behavior. If the claim is wrong and three agents are wrong in three different ways, the problem is not retrieval quality.
Dosu keeps that layer in one place for a team. It reads the work you already do in connected Sources such as GitHub, Slack, and Notion, drafts the resulting knowledge as Documents in a shared Library, uses Monitors to flag a Document when the code moves underneath it, routes changes through Review before they are posted or published unless you turn on auto-accept, and serves the result to coding agents over the Dosu MCP server. Connect your first repo and let the next session start from something the whole team can correct in one place.