BACK TO BLOG

Procedural Memory for Coding Agents: Turning Knowledge Into Action

Learn how to turn a coding agent's discovery into a shared procedure, choose where the instructions belong, and check that the next agent uses them correctly.

Oct 2, 2026/16 min read
Procedural Memory for Coding Agents: Turning Knowledge Into Action

Procedural memory gives your coding agent a method it can use again. A useful procedure tells it when to apply what it learned and how to check the outcome. For your team, the method also needs to reach the next agent and remain valid as the repository changes.

For example, suppose your repository's make test-int target starts the services the integration tests need, while bare pytest tests/integration fails without them. An agent can discover the working command and save a note about it. A later agent still needs to choose that command when you ask it to run the integration suite.

We think a procedure your team relies on belongs in a shared file you can review and maintain. In this example, the instruction explains how to satisfy the suite's dependencies through the repository's workflow. The next agent needs enough context to recognize when that method applies, including which assumptions to revisit when the code changes.

One lesson reaches another agent through repository instructions and an on-demand skill. File sizes are illustrative.

What counts as procedural memory

An agent's session log can record a failed pytest run and the setup that made a later run succeed. You can use that history as episodic memory, a record of what happened. A note that integration tests need running services gives the next agent semantic knowledge about the repository. A procedure turns that knowledge into guidance for running the suite when the task calls for it.

Procedural memory can guide decisions and the order of dependent steps. For example, if your repository generates an API client from a schema, a reusable procedure can explain when a schema change requires regeneration and how to check the client against the code that uses it. A later agent can apply the method to a different schema change because the dependency between source, generated code, and its consumers still applies.

CoALA, a framework for describing language agents, places procedural memory in the model's weights and the agent's code, including its prompt templates. We apply that idea to the editable instructions a coding agent's software loads. Your team can change those instructions without retraining the model. Anthropic also calls the contents of a skill procedural knowledge.

Saving an instruction and loading it are separate operations. A note in a memory store may help an agent after retrieval, while a repository instruction can enter its context at startup. Automatic skill selection adds another step, because the agent needs to recognize that the task calls for the skill before loading its procedure. You can also invoke a skill directly. The filename alone doesn't establish which path the agent takes.

For example, Claude Code's auto memory stays on your machine. A useful command in that memory won't reach your teammate's checkout on its own. Our guide to team memory for coding agents covers that sharing problem.

Writing a procedure for another agent

Start with the conditions that made the method work. A command that succeeds with services left over from another task may fail in a clean environment. The procedure needs to explain the setup it relies on and which parts of that setup the command handles. Try the method with the starting conditions you expect your teammate to use.

Name the task that should trigger the procedure. An integration-test workflow belongs to requests that need that suite, while a documentation edit may need no service setup. A procedure for regenerating an API client should say which source changes require it. Naming those conditions gives the agent a basis for deciding whether to use the method.

For a short convention, an addition to your shared instruction file can be enough. The test example makes the scope and dependency explicit.

 ## Testing
+ When running the integration suite, run `make test-int` from the repository root. The target starts the services the tests need. Bare `pytest tests/integration` fails without them.

The reason beside the command helps your teammate assess the instruction. If the setup changes later, the explanation points to the assumption that needs revisiting. Details such as the working directory help another agent apply the method outside the session where you discovered it.

Keep the evidence for the change in the pull request. For the test example, link the Makefile target and a test run that confirms it starts the required services. The shared instruction can focus on applying the procedure. You can preserve the discovery's history separately so the next agent can follow the method without reconstructing the original investigation.

Where the agent loads the instruction

Once you have the instruction, choose where it belongs by asking when the agent needs it. A repository-wide convention can live in the root AGENTS.md or CLAUDE.md that your coding tool supports. Advice for one service may belong closer to that service, but the tool's loading behavior determines whether the agent sees it.

Codex builds its initial project instructions from the repository root down to its working directory. An AGENTS.override.md takes the place of the AGENTS.md beside it. A file below the starting directory isn't part of that initial chain.

Consider instructions for a payments service in this layout.

repository/
├── AGENTS.md
├── Makefile
└── services/
    └── payments/
        └── AGENTS.md

Starting Codex in services/payments includes the instruction chain through that directory. Starting at the repository root doesn't include the payments file in the initial chain. Put guidance where the agent encounters it before making the decision it governs. A service-specific procedure in a descendant file can't guide a root-level decision through the initial instruction chain.

Codex's project instructions share a default budget of 32 KiB. A large root file can consume space before Codex reaches a more specific file. If a service-specific instruction is missing from the session, check whether it fits within the remaining budget.

Claude Code combines guidance from the working directory and its parents. Its docs warn that conflicting instructions can lead Claude to choose arbitrarily between them. When a root file and a service file prescribe different workflows, check whether they describe different conditions or contradict each other. If the old instruction is wrong, edit it at its source instead of adding another reminder.

Claude Code also supports path-scoped rules. A rule in .claude/rules/integration-tests.md can start with this frontmatter.

---
paths:
  - "tests/integration/**"
---

A matching file read triggers the rule. Scoping helps when the advice belongs to work on those files, but an agent may need a procedure before reading a matching file. Choose the loading mechanism around the activity that needs the guidance. A task-oriented skill can make a workflow available through its description, while a path-scoped rule depends on file access.

When a procedure belongs in a skill

A procedure may need several steps or a decision based on an intermediate result. You might need to explain what to do when setup fails, which output establishes success, and when cleanup applies. Putting that entire workflow in the root instruction file would add it to tasks that don't need the procedure.

A skill lets you keep the detailed procedure available without loading its full text into every session. In Claude Code, a repository skill can live at .claude/skills/test-integration/SKILL.md. The Agent Skills specification defines a skill's name and description as required metadata. The description explains what the skill does and when the agent should use it.

The integration-test example shows how a skill can include failure handling and completion criteria alongside the command.

---
name: test-integration
description: Run this repository's integration suite with make test-int and diagnose test-service startup failures. Use when asked to run integration tests or validate changes to tests/integration.
---

# Integration tests

Run the suite from the repository root with `make test-int`.
The target starts the services the tests need. Bare
`pytest tests/integration` fails without them.

1. Run `make test-int` from the repository root.
2. If service startup fails, report the failing service and its error.
   Don't report a test result when the suite didn't run.
3. If tests fail, preserve the failing test names and relevant output.
   Don't switch to bare pytest, which omits the required service setup.
4. Report the command you ran and the suite's result. Mention skipped
   tests so the person reviewing the change knows what ran.

The skill explains how to interpret a failed run as well as how to start one. A service failing to start and a test assertion failing need different follow-up work. If your repository needs explicit cleanup, add the verified cleanup command and explain which services the procedure owns before asking an agent to stop them.

When you expand a short convention into a skill, replace the original root-file instruction with a pointer that names the skill and when to use it. Keep the detailed steps in the skill. In this example, the root file can direct integration-test tasks to test-integration, while the skill explains how to run the suite and interpret failures. You then have one place to maintain that workflow.

Claude loads skill metadata before the full procedure. The description needs to connect the method to the task the agent receives. Include the kind of request or source change that should trigger it. A broad description such as help with testing leaves the agent to infer which test tasks need that workflow.

Skills can also include scripts and reference files. Use a script for deterministic steps your repository needs, and keep the skill's instructions focused on when to run the script and how to interpret its result. If the repository's task runner handles those steps, the skill can call it directly. Reference files can explain less common failures without adding that detail to every use of the skill.

A procedure's location determines when its instructions can enter an agent's session.

Checking the next agent's work

Test the procedure in a fresh session, away from the conversation where you wrote it. An agent that helped you author the skill may remember the command from that conversation even if the skill's instructions leave something out.

Give the fresh agent a task that should call for the procedure, without supplying the method or naming the skill. Watch how it finds the guidance and applies it. For a skill, check selection separately from execution. Invoking a skill by name helps you test its steps, but it doesn't show how the agent selects it from an ordinary task request.

If the agent misses the instruction, inspect the loading path before expanding the prose. Was the shared file present in the checkout? Did the session start in a directory that includes it? Did an override file replace it? For a skill, did the agent receive the description? Claude Code limits its skill listing, so a file on disk doesn't prove the model received the metadata.

If the agent loads the procedure but follows a different workflow, check for conflicting guidance or an unclear trigger. For example, an agent might edit a generated API client directly despite instructions to regenerate it from the source schema. If the agent runs the generator but a required tool is missing, check whether the procedure names the required tool. A validation failure after regeneration may instead identify a problem in the schema change. Reporting that validation failure is part of following the procedure.

Following the instruction also doesn't establish that the instruction helps the task. Gloaguen and colleagues evaluated repository-level context files and found no statistically significant change in task success with generated or developer-written files compared with no file. The files added inference cost, even though agents generally followed their instructions. The authors recommend using context files for non-standard practices and evaluating their effect.

The study doesn't tell you whether a particular procedure helps in your repository. Compare fresh sessions with and without the procedure, keeping the task and environment comparable. Record whether the agent chose the expected method and produced the required result. Try a task outside the procedure's scope too, to see whether the instruction causes unnecessary steps.

Vary the task within the procedure's scope to check whether the method applies beyond the original discovery. For the API-client example, a different schema change should still prompt regeneration and a check of the code that uses the client. The comparison gives you evidence about the tasks and environments you tried, including whether the agent recognized the dependency when the details changed.

Keeping the procedure current

Update a procedure in the pull request that changes the requirement it describes. A renamed command is an obvious trigger, but changes to dependencies or generated outputs can also invalidate the method. The code change gives you the reason to revisit the instruction before the next agent relies on it.

Search your instruction files and skills for the old command or requirement. If the root file directs the agent to a skill, check that the pointer still describes when to use the workflow. When several files repeat the full procedure, choose a place to maintain the detailed steps and have the other instructions point to it. You then have fewer copies to reconcile when the setup changes.

The explanation can go stale even when the command name stays the same. A task runner might stop handling setup, or a generator might start producing files that require a new validation step. Revisit the procedure's assumptions and the evidence it asks the agent to check. If automation takes over a manual step, remove the instruction that tells the agent to repeat it.

Retire instructions when the requirement goes away. A workaround for an old dependency can keep influencing the agent after the code no longer needs it. Our agent-memory maintenance post explains how new knowledge can prompt you to revise related instructions.

A running session may retain an earlier version of a skill. Claude Code keeps an invoked skill's instructions in the conversation and doesn't re-read the file on later turns. After changing the procedure, use a fresh session to check the revised instructions. Keep the repository and environment comparable so an unrelated setup change doesn't hide whether the revision works.

Hooks and CI checks

Your procedure helps the agent choose how to complete a task. You can require CI checks for passing tests or confirming that generated code matches its source before a change merges. The checks still apply when an agent skips the procedure.

A Claude Code PreToolUse hook can block matching tool calls before they run. In the test example, a hook could reject a shell call that bypasses the required service setup. Its coverage would depend on the calls it matches, while a required CI check would verify the suite's result. You can guide the agent through the method and enforce a required outcome separately.

Changes to instructions and permissions

A change to your shared instructions can affect later sessions that load them. CoALA warns that procedural changes can introduce bugs or undermine an agent's intended behavior. In April 2026, Cisco demonstrated a persistent memory compromise in Claude Code through an untrusted repository's install script. After the user trusted the repository and approved package installation, the script overwrote MEMORY.md and added a settings hook. Cisco reports that Anthropic later moved user memories out of the system prompt in response.

When a procedure adds a command or changes which files the agent should modify, include those changes in the review. Keep changes to permissions and hooks explicit in the diff too. Your teammate can see which actions the procedure asks the agent to take and whether the change also expands its permissions.

Share what your agent learns with Dosu

At Dosu, we help you preserve knowledge from coding sessions so another agent can retrieve it through your team's Library. A discovery gives you material for a shared procedure, including the conditions that made the method work.

Our session-end hook queues sessions for study, which can save durable lessons as notes. Your agent can use read_knowledge to retrieve shared knowledge before a task. Notes from work on a connected repository can stay tied to the repository and branch, so use the same context when checking whether another agent can retrieve a new lesson.

A useful lesson explains why the method worked and where your agent found the requirement. Our CLI's study process reads your messages and the agent's replies and skips tool output. A discovery that appears only in tool output won't become a note unless a message states the lesson. The CLI filters the session locally and redacts detected secrets such as API keys and tokens before sending conversation text to Dosu's LLM gateway.

Saving the lesson in the Library doesn't automatically edit AGENTS.md or create a skill. Your team chooses whether the procedure should become startup guidance or a workflow the agent loads for a particular task. Dosu gives you shared knowledge to make that decision, and the fresh-session check tells you whether the instruction changes the next agent's work.

Install Dosu with one easy step.

curl -fsSL https://cli.dosu.dev/install | sh

Create your Dosu account and connect a repository to your team's Library. The CLI guides you through connecting your coding agent and enabling session study.

After Dosu studies a session, inspect any lesson it saves and use that knowledge to update your shared instruction file or a skill. Start a fresh session in the same repository and branch, ask your agent to run the task, and check whether it follows the method and produces the expected result. You can then share a procedure you've checked, with the explanation your teammate needs to keep it current.

Found this article helpful?

Share it with your network to help others discover valuable insights.

Want more like this? Subscribe via RSS

Related Articles

Ready to transform your workflow?

Join leading organizations using Dosu to automate documentation, streamline support, and empower development teams to focus on building great products.