BACK TO ARTICLES

How to Prune Stale Agent Memory

A fact your agent wrote down six months ago outlives the decision that produced it, and nothing in a normal session rechecks it. Here is how to tell which memories to retire, and how to fire that retirement off the change that made them false.

Dosu
/
Oct 1, 2026/11 min read

Agent memory needs maintenance as the codebase changes. A dependency gets replaced, an API is retired, or a team revises a convention. The code moves on, while instructions and saved notes can continue pointing agents toward the old approach. That guidance may have been useful when it was written. Keeping it around without revisiting it is what makes it a problem.

Updating code should include checking the knowledge that describes it. Which instructions still apply? Which need revision? Which should be removed? Making those questions part of the change gives teams a practical way to keep outdated guidance from shaping future work.

A forgotten fact costs you a lookup, a stale one costs you a wrong change

Two things go wrong with what an agent knows, and they pull in opposite directions.

When an agent loses a fact that is true, you pay for a lookup. The agent re-reads the file, searches the repository, or asks you which package manager the project uses for the fourth time this week. The cost is irritating and recoverable, and the session-boundary version of it has its own diagnosis in Why Coding Agents Forget Between Sessions.

When an agent holds on to a fact that is false, you pay in output. The agent writes the import, follows the convention you retired, or re-proposes the design the team ruled out last spring, with the same confidence it brings to a correct fact. From inside the session a stale instruction looks exactly like a current one.

Keeping useful context available and keeping it current are separate maintenance jobs. More retained context does not establish that a stored claim is still true. When a source changes, review the old claim and decide whether to correct it, retire it, or keep it as dated history.

What an agent can read again is what you have to retire

Not every wrong statement is equally durable. The line worth drawing runs between what an agent loads again on its own and what it only sees if you bring it back.

Claude Code treats file-backed instructions differently from conversation history. After compaction, the project-root CLAUDE.md, unscoped rules, and the auto-memory index are re-injected from disk, while earlier hook context is summarized. A chat-only correction may be summarized away or omitted from a fresh conversation, but retained history can return when you resume one. A stale rule in a loaded instruction file can be injected again in later sessions.

Auto memory makes the point sharply. A background sweep can remove eligible session transcripts once they pass the cleanupPeriodDays retention period, though settings can pause that cleanup, bare mode skips it, and Desktop and Cowork sessions carry their own exceptions. Auto memory sits outside that sweep entirely. The memory documentation states that the memory directory is excluded from the retention sweep, and that MEMORY.md and its topic files stay until you or Claude edits or deletes them. A transcript may expire. A memory will not.

So audit the durable sources an agent can read again:

  • Instruction files loaded at launch: CLAUDE.md, CLAUDE.local.md, an AGENTS.md where it applies, and .claude/rules/ files with no paths: frontmatter.
  • The auto-memory index (MEMORY.md), which loads at startup, and its topic files, which Claude reads on demand.
  • Your shared knowledge layer, through the retrieval path agents use.

Path-scoped rules and nested instruction files sit in between. They are summarized away with the conversation and reload when the agent next opens a file they match.

Three signals that a memory has gone stale

Treat the following as review signals. Check each flagged memory against its current source before deciding what to change, especially when the memory records a decision or its rationale.

Contradiction against a current source. The strongest of the three. Memory says Yarn, the repository holds a pnpm lockfile and build commands to match. A checker can compare stored claims about dependencies, commands, and conventions against the files that settle them. Decide which side wins before you automate the comparison.

Claude Code ships a check in roughly this shape. /doctor prompt-audit reads your instruction files and reports problems such as instructions written for older models, references to files or commands that do not exist, and files that contradict each other. You get the findings with proposed edits, and nothing changes until you ask for the edits to be applied. The audit needs v2.1.283 or later, and by default covers CLAUDE.md, CLAUDE.local.md, and AGENTS.md, plus the rules, skills, commands, subagents, and output styles under .claude/ and ~/.claude/.

Age weighed against change rate. Age on its own says little. A six-month-old note about a stable release process can be perfectly good, while a six-week-old note about an API being rewritten can be wrong before you think to check it. Claude Code records a modified timestamp in auto-memory frontmatter as of v2.1.214, on files that already carry frontmatter, and the documentation describes it as showing how current the fact is, both to you and to Claude when it reads the memory back. Weigh that date against how often the paths the memory describes have changed since.

Provenance that no longer resolves. Record the source path and commit when you capture a fact. When history shows the source was deleted or substantially rewritten, the memory needs another look. A moved file may need nothing more than a provenance update. A deleted implementation usually means the guidance went with it.

There is also a crude forcing function worth knowing about. Only the first 200 lines or 25KB of MEMORY.md load at the start of a conversation. Near that limit Claude Code reminds Claude to shorten the index by merging or dropping stale entries, and past it the write still succeeds but the index has to be rewritten, because everything beyond the limit is dropped on the next load. Useful pressure to prune, but it tells you the file is long rather than which entries are wrong.

Retire, correct, or keep as history

A stale flag should open a decision with three outcomes, not trigger one blanket delete.

Retire when the claim is false and nothing takes its place under the same heading. An instruction to import a library you removed should stop reaching retrieval once the migration lands. You can keep the record itself, with its old value, the date, and the change that invalidated it, and mark it inactive. Privacy or compliance rules may require deleting it outright instead.

Correct when the subject still matters and only the value moved. Tests went from Jest to Vitest. The command to run migrations was renamed. Store the new statement as the active one and link it to the version it replaced, so instruction files, retrieval, and generated summaries all serve the current answer.

Keep as history when a past decision explains the present design. A queue technology the team evaluated and rejected is worth having around, because it stops an agent proposing it again for reasons the team weighed once. Rewrite it as a dated decision that says what was considered, why it was rejected, and which assumptions applied, so nobody reads "we rejected this in 2024" as "never use this." The controls around shared memory that people write to are covered in Agent Memory Governance for Engineering Teams.

Three questions sort most cases:

  1. Is the claim false with nothing to replace it? Retire it.
  2. Does the same subject need a new value? Correct it.
  3. Does the old reasoning explain why the code looks the way it does? Keep it, dated.

Attach retirement to the event that made the fact false

A periodic audit leaves a window open. Between the migration merging and the quarterly cleanup, every session reads the old instruction and some of them act on it.

The alternative is to fire the review off the change itself. A dependency removal, a convention change, an architectural decision that supersedes an earlier one: each of these is an event that makes specific stored facts false at a known moment. Capture enough provenance for a job to find them. When a pull request that removes Moment.js merges, a repository hook or CI step can search CLAUDE.md, AGENTS.md, .cursorrules, your rules directory, and your shared layer for references to it. The same pull request usually carries the replacement guidance.

Have the event open a review rather than delete anything. A maintainer approves removal, substitutes the new rule, or marks the old one as history, and the record keeps what changed, which event caused the review, and who signed it off.

Claude Code's own guidance points the same direction. Because two instructions that contradict each other leave Claude free to pick either one, the documentation tells you to review your instruction files periodically and remove outdated or conflicting content. Attaching that review to the merge that caused the drift beats relying on anyone to remember it later.

Where a shared layer changes the problem

Repository instruction files can be fixed for the team by committing the change and having teammates pull it. Auto memory is machine-local, so editing your copy does not update anyone else's. If several agents learned the same stale rule independently, each copy needs review, or the team needs an authoritative shared source for that guidance.

Dosu maintains shared Documents using connected Sources, including GitHub and web access at query time, plus Slack and Notion on the Teams plan. Turn Monitor on for a connected GitHub Source and Dosu checks pull-request changes against relevant published Documents, flags guidance that may be stale, and proposes targeted edits. Reviewers can Accept, Edit, or Decline those proposals from the pull request or the Review page, and suggested updates carry citations back to the pull requests, files, and threads they came from. For Git-backed documentation, accepted updates go through a docs pull request or, when configured, the triggering pull request's branch.

Dedicated memory APIs such as Mem0, Zep, and Letta answer an adjacent question. They give you storage and retrieval primitives for an agent you are building, which is what you want when memory behavior is part of your product rather than part of your documentation. Whichever layer you run, the question to put to it is how a fact gets retired, not only how it gets stored.

FAQ

Is retiring a memory risky if the agent needs it later?

Retirement stops a fact being served as current guidance, which is not the same as destroying it. Keep the record with its retirement date, the change that invalidated it, and whatever replaced it, and let retrieval reach it for questions about why the current design looks the way it does. Deleting outright is for cases where privacy or compliance requires it.

Would a bigger context window fix this?

Not this failure. A larger window keeps more true facts within reach, which helps when something valid fell out of context. A false fact in a larger window simply has more room to sit there and be read. The two problems need different work.

Can staleness detection be automated?

Partly. Contradiction against a current source, provenance that no longer resolves, and age weighed against how often the related code changes are all checkable, and /doctor prompt-audit covers part of the first one for instruction files. Leave the ambiguous cases to a person, particularly where a memory records intent rather than an observable technical fact.

Who should approve a retirement?

Whoever owns the affected code or the decision. A dependency owner can confirm that a migration invalidates old library guidance, and whoever maintains your decision records can approve a change to one. Tooling can route the review and keep the record, but ownership decides the call.

Does this apply to a vector store the same way?

The logic generalizes to any store an agent retrieves from. What differs is the mechanism: a vector store needs invalidation metadata and retrieval filters so retired records stop coming back as current, while instruction files need an edit or a deletion. Instruction files are worth doing first, because they load on every session without anything prompting the lookup.

Start with one retiring event

Take the last dependency migration that merged in your main repository. Search your repository's instruction and rule files for the removed package. In Claude Code, open /memory and inspect the auto-memory directory too, including MEMORY.md and its topic files. Run /doctor prompt-audit for the instruction-file check. Compare each hit with the migration and decide whether to retire it, correct it, or keep it as history.

Use the migration as a first sample, then add the check to future dependency-change reviews. Connect a repository to Dosu and enable a Monitor to review proposed Document updates alongside the code changes that prompted them.

Ready to transform your workflow?

Join leading organizations using Dosu to automate documentation, streamline support, and empower development teams to focus on building great products.