Why Coding Agents Forget Between Sessions and How to Fix It
Losing the thread mid-task and starting cold in a new chat are two different failures with two different fixes. Here is how to tell which one you have before you change anything.
You are forty minutes into a refactor when the agent proposes the approach you ruled out at minute ten. Or you open a new chat the next morning and it asks which package manager the repo uses, for the fourth time this week. Those feel like one problem and they are two. One may be your conversation being summarized out from under you. The other is a new request that never carried the answer. They have different fixes, and picking the wrong one means reaching for a shared memory layer to solve something that lives in your compaction settings.
Your agent's memory is whatever ended up in this request
The model answers from the context supplied for each request and keeps none of it afterward. The agent harness around it maintains the conversation and decides which instructions, files, tool results, and saved knowledge go into the next one. Remembering is not something the model does. The harness does it, by keeping a fact somewhere it can find again and putting it back into the request when it becomes relevant.
Rebuilding that request on every turn makes each rebuild a place where knowledge can fall out. There are two such points, and naming them separately is most of the work.
| Failure | When you notice it | What is missing | Where the fix lives |
|---|---|---|---|
| Intra-session loss | Mid-task, after a long conversation | Detail from earlier turns that the summary did not carry | Compaction timing and handoff notes |
| Inter-session cold start | On the first message of a new session | Anything from a previous conversation that the new request does not include | Durable storage the next session can load |
Both rows describe what the next request contains, not what exists anywhere. That distinction is what makes the diagnosis in the last section work.
What compaction keeps
As a Claude Code session approaches its context limit it compacts, summarizing the conversation so far and replacing the original messages with that summary. The session continues, which is exactly why the loss is easy to miss.
Anthropic documents what survives, and the list is more specific than "it keeps the important parts." According to the context window documentation, after compaction:
- Project-root
CLAUDE.md, unscoped rules, and auto memory are re-injected from disk rather than summarized. - Up to five of the files Claude read or edited are re-read, most recently modified first. A file over 5,000 tokens comes back as a path reference without its contents.
- Invoked skill bodies are re-injected, capped at 5,000 tokens per skill and 25,000 tokens in total, oldest dropped first.
- Nested
CLAUDE.mdfiles in subdirectories and rules withpaths:frontmatter are summarized away with everything else. They reload only when Claude next reads a file they match. - Context that a hook added earlier is summarized along with the conversation.
Read that list for what is absent from it. Everything that reached the model as conversation, including the constraint you typed, the approach you rejected, and your reason for rejecting it, has no file the harness re-reads on your behalf. Claude Code does save those messages to the session transcript on disk, so they are not destroyed, but what the next request carries after compaction is the summary, at whatever fidelity the summarizer chose. Anthropic's troubleshooting lists the same conditions from the other direction. An instruction missing after compaction was given only in conversation, or it lives in a nested CLAUDE.md or a path-scoped rule that has not reloaded yet.
Other tools compact too, and the mechanism differs by tool, by model, and by version. Check the documentation for the agent you run rather than assuming Claude Code's behavior generalizes.
Instruction files carry the least of what crosses both boundaries
Committed instruction files clear both failures without extra tooling, and they are not the only layer that does. Auto memory loads at the start of every session and is re-injected from disk after compaction as well. What committed files add is source control, which makes them the layer you can hand to the team.
Claude Code loads CLAUDE.md and CLAUDE.local.md from your working directory and every directory above it, concatenated from the filesystem root down, and the project-root file is re-injected after compaction. On v2.1.277 and later, a repository's AGENTS.md is read directly when there is no CLAUDE.md, .claude/CLAUDE.md, or CLAUDE.local.md in your working directory or above it. A personal CLAUDE.local.md is enough to change that discovery. Cursor keeps its own rules under .cursor/rules/ or a .cursorrules file, which Claude Code reads when you run /init and folds into the generated CLAUDE.md rather than loading every session.
The binding limit is size. Anthropic's guidance is to keep each CLAUDE.md under 200 lines, because longer files consume more context and reduce adherence. Splitting one with @path imports helps you organize it without reducing its cost, since imported files also load at launch. Path-scoped rules under .claude/rules/ are the lever that does reduce startup context, because a rule with paths: frontmatter loads only when Claude reads a matching file.
So instruction files are for what holds across every task, such as build commands, conventions, directory boundaries, and review requirements. They are the wrong home for the state of the work in front of you right now. A file that accumulates a diary of recent decisions gets longer every week, and the instructions that mattered get diluted by the ones that no longer apply.
They are also input, not enforcement. Claude Code delivers CLAUDE.md as a user message after the system prompt, and the documentation is explicit that there is no guarantee of strict compliance. When a rule has to hold regardless of what the model decides, it belongs in a hook, a permission setting, a test, or CI.
Local resume restores a session, not shared team knowledge
Coding agents do persist sessions, and it is worth knowing what you already have before you add anything. Claude Code writes transcripts continuously. claude --continue reopens the most recent conversation in the current directory, claude --resume opens a picker over this machine's local history, and /resume switches conversations from inside a session. A resumed session restores the full history, tool calls and results included.
Sessions are not confined to one machine either. A cloud session runs on cloud infrastructure and keeps going after you close your laptop, and claude --teleport pulls one into your terminal, checking out its branch and loading its conversation history. Anthropic documents the two as distinct for this reason, since --resume reads local history and does not list cloud sessions.
None of that hands one session's lessons to a teammate. Local transcripts sit at ~/.claude/projects/<project>/<session-id>.jsonl on the machine that ran the session and age out on a 30-day default retention you can change with cleanupPeriodDays. Auto memory is scoped the same way, loading the first 200 lines or 25KB of MEMORY.md at the start of every session, with topic files read on demand, and the documentation states the files are not shared across machines or cloud environments. Moving or reopening a session relocates it. It does not make what that session learned available in everyone else's next session.
Retrieval recovers what somebody wrote down
The layer above resume trades holding context for fetching it. An agent that can read files, search the repository, and query connected tickets and pull requests reconstructs the relevant facts on demand instead of needing the previous conversation preserved. We worked through that tradeoff against memory services and larger windows in Agent Memory vs RAG vs Bigger Context Windows, so this section stays on the part that decides your diagnosis.
Retrieval reaches what someone recorded somewhere it will look. If an engineer explained a compatibility constraint in a chat window and never put it in a comment, an issue, a pull request description, or a decision record, that reasoning is not in the sources retrieval searches. The agent may infer a plausible reason from the code, and an inference is not the decision. Retrieval quality tracks source quality too, since missing permissions hide relevant tickets and a stale document will confidently lead an agent to last quarter's answer.
A bigger context window moves the boundary, not what crosses it
Claude Code supports a 1 million token context window on several models, and compaction works the same way at the larger limit. More capacity buys later compaction. It does not change what the summary keeps once compaction runs, and a fresh session still starts without the previous conversation.
Capacity is also not the same as attention. Liu et al. found in 2023 that model performance on retrieval tasks was highest when the relevant information sat near the beginning or the end of the input, and degraded when the model had to use information in the middle of a long context, including in models built for long contexts. That result is in Lost in the Middle, and it matters for diagnosis. A fact can sit inside the window and still go unused, which is why a mid-session mistake is not by itself evidence that compaction dropped anything.
What to do at each boundary
Before compaction catches you mid-task
Compact on purpose at a task boundary instead of waiting for the automatic pass. /compact preserve the root cause, the files changed, and the approaches we ruled out keeps what you name rather than what the summarizer infers. If you keep getting caught anyway, /autocompact 500k sets how full the window gets before the automatic pass runs.
Write the handoff into a tracked file rather than leaving it in the chat. After compaction, and in the next session, a tracked file is something the agent can open again on purpose. The conversation is saved to the transcript, but nothing loads it back into the next request for you. Record the objective, the files changed, the decisions that have to hold, the failing tests, and the next action.
Use /clear between unrelated tasks. Old conversation crowds out the files you need next and costs tokens on every message until it goes.
Before the next cold start
Keep the instruction file inside the 200-line guidance and push narrow guidance into .claude/rules/ with paths: frontmatter, so it loads when it is relevant instead of every session.
Put rationale where retrieval can reach it. The pull request that made the change is usually the cheapest place, because you are writing there anyway. A decision that constrains future work deserves a record of its own.
Then give the durable facts somewhere shared to live, so the next session and the next engineer's session both start from them instead of re-deriving them in parallel.
Which boundary is costing you
Both tests are things you can run this afternoon.
For intra-session loss, start from the symptom and then confirm it. An agent that is fine for thirty turns and then contradicts a settled decision, reopens a closed investigation, or drops a constraint you set early may have lost that context. Confirm a compaction ran. Use /context to check usage and which memory files loaded, then reopen the source of the missing instruction. If it lived only in chat, restate it or put it in the handoff before continuing. If the guidance is loaded but the agent still ignores it, investigate conflicting instructions or model behavior.
For a cold start, open a fresh session and hand it the task your current session is halfway through. Write down every question the new session asks that the running one could already answer. Then split that list. Where the answer is sitting in the repository or a connected tool and the agent never reached it, you have a retrieval problem, and the fix is making it reachable rather than writing it down a second time. What remains, the answers that exist in no durable source, is your inter-session gap, and it is the same gap for every engineer on the team.
FAQ
Do my CLAUDE.md instructions survive compaction?
The project-root CLAUDE.md and unscoped rules are re-injected from disk, so they survive. Nested CLAUDE.md files in subdirectories and rules with paths: frontmatter are summarized away with the conversation and come back when Claude next reads a file they match. Anything you only said in chat has no file the harness restores for you, so what carries forward is whatever the summary kept.
Should I use /clear or /compact?
Use /compact when you want to keep working on the same task with a smaller request, and give it instructions about what to preserve. Use /clear when you are switching to unrelated work, because carrying the old conversation forward costs tokens on every subsequent message without helping.
Is auto memory shared with my team?
No. Auto memory is per repository and machine-local. All worktrees of the same repository share one directory, and the documentation states the files are not shared across machines or cloud environments. Treat it as a convenience for you, not as team memory.
Does Claude Code read AGENTS.md?
By default it reads AGENTS.md only when there is no CLAUDE.md, .claude/CLAUDE.md, or CLAUDE.local.md in your working directory or above it, and it requires v2.1.277 or later. You can change that with the Project instructions setting, or keep one shared file by importing @AGENTS.md from a CLAUDE.md next to it.
Why does a long session get expensive even when I barely type?
The full conversation is sent with every request, so a one-line question late in a long session still carries the whole history. Prompt caching makes the re-read cheaper rather than free, and the first message after the cache lifetime expires reprocesses everything.
Start the next session from what this one learned
Run the cold-start test first, then check which answers are missing from durable sources and which the agent failed to retrieve.
If the missing answers are mostly rationale, prior fixes, and conventions that live in pull requests, issues, and chat rather than in the code, that is the gap Dosu is built to close. Dosu reads connected Sources, which are GitHub and live web access, plus Slack and Notion on the Teams plan, and turns that activity into Documents. Two paths run alongside each other and they are not the same. Monitors check ready pull requests and their later commits against published Documents and send a proposed edit to Review, where someone approves it unless you turn on auto-accept. Separately, coding agents read that knowledge over the Dosu MCP Server with read_knowledge, and write_knowledge saves a note rather than editing a Document. When the working repository is connected to the selected Library, the note is scoped to that repository and branch, and other members of that Library working on the same branch can read it back. Without that connection, or when the client does not supply git context, the note is still saved, without the repository anchor. npx @dosu/cli setup configures the client for you.
Connect a repository to Dosu and test whether a fresh agent session can retrieve one decision your team has already recorded.