BACK TO BLOG

Semantic memory for coding agents needs a history

Your coding agent needs to know which facts still apply. Our Celery-to-DBOS migration shows why semantic memory needs scope, dates, and source evidence.

Oct 1, 2026/10 min read
Semantic memory for coding agents needs a history

If your coding agent keeps recommending a tool your team replaced, saving the new tool's name isn't enough. Your agent needs to know where the change applies and when the old fact stopped being true. Semantic memory stores reusable facts about your project, but a relevant fact can still describe an earlier version of the code.

We found an example in our own repository. A comment on main still said "Wrap as single step until we migrate it from Celery to DBOS" when we checked it for this post. We added the comment on October 26, 2025, moved the thread QA extraction code beneath it off Celery on November 2, and removed our Celery app on November 15. The comment survived all of those changes.

A memory extracted from that comment would need the November 2 change and a source to check. Otherwise, your next session could inherit the comment's outdated claim even though it came from main.

Celery and DBOS overlapped during the migration

Our June 2025 post, Celery Preserializers: A low-friction path to Pydantic support, describes code we used at the time. We added our first DBOS ingestion workflow on June 16 and moved GitHub initial sync and its index trigger in July. By September 2, when we published From Celery to DBOS: Scaling Dosu’s Pipelines to 20k Workflows per Hour, we could write that we'd migrated pipelines to DBOS. Slack and Confluence indexing still used Celery.

We didn't remove our Celery app until November 15. The timeline shows the overlap using code-history dates, not deployment times.

Dosu's code migration off Celery at exact day scale. Commit dates show code changes, not deployment times.

For work during that migration, you need to know which pipeline the fact describes. For work after the removal, you need to know that the Celery implementation no longer exists. Saving the June Celery description and the September DBOS announcement as undated facts leaves your agent to reconstruct both distinctions.

In the CoALA framework for language agents, semantic memory holds knowledge about the world and the agent. Here, that's which system a particular pipeline uses. The migration events explain when the fact changed.

When did the fact change?

If you save our June post in December, December is when you saved the fact, not when we used Celery. Your memory system needs to keep those dates separate.

Mem0's add() uses a model to extract facts from your text. Its extraction prompt asks for the new state and what it replaced. That helps when your source explains the transition, including which pipeline changed.

In the open source Mem0 implementation we checked, add() uses today as the extraction reference date and rejects the Platform-only timestamp argument. Putting the event date in the input text gives the extractor information it otherwise lacks.

Graphiti, the temporal knowledge graph behind Zep, accepts a reference_time for each episode. Its extraction prompt asks when a fact became true and when it stopped, using valid_at and invalid_at. For our migration, September 2 would identify the announcement. It wouldn't establish when a particular pipeline stopped using Celery.

Once you've saved both descriptions, retrieval still has to decide what to return.

A fact moves through extraction, ranking, and dated context. The similarity scores come from our measurement. The spatial layout and ranking examples are illustrative.

What our similarity search test found

We tested whether a month in the question would change which undated fact ranked first. We compared "Dosu's ingestion pipelines run on Celery." with the same sentence naming DBOS, using cosine similarity in three small open embedding models. Each sentence makes a broad claim about the same pipelines, so both can match an ingestion question.

The check used five fact sentences and five questions in one run. Each cell below names the undated fact the model ranked first and its lead in cosine similarity.

Questionbge-smallall-MiniLMnomic
What do Dosu's ingestion pipelines run on?DBOS by 0.0343DBOS by 0.0968Celery by 0.0038
What do Dosu's ingestion pipelines run on now?DBOS by 0.0322DBOS by 0.0984DBOS by 0.0072
What did Dosu's ingestion pipelines run on in December 2025?DBOS by 0.0180DBOS by 0.0268Celery by 0.0013
What did Dosu's ingestion pipelines run on in June 2025?DBOS by 0.0356DBOS by 0.0394Celery by 0.0016
Which task queue does Dosu use for ingestion?DBOS by 0.0380DBOS by 0.1093DBOS by 0.0194

For the June and December questions, all three models kept the facts in the same order. The nomic model's leads for Celery were only 0.0016 and 0.0013. Adding the word now flipped nomic to DBOS by 0.0072, but neither fact supplied a date that could establish when it applied.

Our measurement. The month in the question never changed which fact ranked first.

We then added dates to the facts, using "As of June 2025" for Celery and "As of November 2025" for DBOS. With those sentences, bge-small ranked Celery first for June and DBOS first for December. MiniLM still put DBOS first for both months, and nomic still put Celery first for both.

We also tried "Dosu moved its ingestion pipelines from Celery to DBOS." Every model scored that change sentence below both undated facts for all five questions. For the plain question in bge-small, the change sentence scored 0.7258, against 0.8763 for Celery and 0.9106 for DBOS. The sentence that explained the transition was less similar to the question than either description of a state.

This is a small illustration, not a benchmark of memory services or agent answers. Celery and DBOS overlapped in June, so there's no single correct June backend for every pipeline. Your service might use another embedding model, filters, or a reranker. Mem0's hosted Platform, for example, adds a temporal ranking signal using time metadata it extracts when saving memories.

The change sentence explains the migration, but these models ranked it below the shorter descriptions. Keeping history in storage won't help if your retrieval leaves it out of the agent's context. Our Search & Retrieval post covers choosing search methods for different questions.

How to handle conflicting agent memories

For the comment we found, your agent needs the history of thread QA extraction. You could give it this record, with the save date kept separately.

Fact: Thread QA extraction uses DBOS.
Changed: 2025-11-02, when thread QA extraction moved off Celery.
Source: Migration commit ecdffedc72.
Check: backend/internal/dbos_workflows/jobs.py and the extraction code it calls.
Caveat: The scheduling comment still says the migration is pending.

The record names the affected code and gives your agent evidence to check against the branch it's working on. The broader migration timeline can explain the June post without treating every pipeline as if it changed on the same day. Zep's paper describes putting each fact's start and end dates beside its text when building the agent's context. If you assemble the prompt yourself, check that your code passes those dates through.

When a newer fact contradicts an older one, Graphiti can give the older fact an end date. Its default search leaves time filters unset, so your code still needs to filter by period or give the dates to the agent. Storing an end date doesn't guarantee that your agent sees it.

A newest-wins rule also needs judgment. Our September announcement could have led a model to retire the Celery fact more than two months early, because the announcement didn't mean we'd finished every pipeline. Dosu's Topic writer reads related notes newest first and prefers newer notes where they conflict. Dosu keeps the earlier notes, and the model still has to distinguish a contradiction from two notes about different pipelines. Keeping those notes lets you inspect and correct a mistaken synthesis, much as you can keep an architecture decision record after a newer decision supersedes it.

A code citation gives your agent another way to check. GitHub Copilot stores code citations with repository facts and prompts the agent to verify them against the current branch. A citation to our deleted Celery app would no longer point to existing code. Our leftover comment requires a closer read, because the comment exists but the implementation beneath it has changed.

An expiry timer doesn't replace that check. Copilot deletes unused memories after 28 days, and successful validation and use can refresh the timer. If a check accepts a stale comment without reading the code, repeated use can keep renewing the wrong fact. Our Maintenance post explains the evidence and events we use to revisit stored knowledge.

Check which facts your coding agent retrieves

Pick a tool your agent still recommends after your team moved away from it. Ask about that tool in a fresh session and inspect the memory text behind the answer. Could you tell which code the fact describes, when it changed, and how to check it? Our storage post explains how versions help you trace a correction.

With Dosu, your agent calls read_knowledge to retrieve knowledge from a Library shared with your team. Our agent memory governance guide covers choosing who can read it.

Connect your coding agent to Dosu. If you already use the Dosu CLI, run dosu upgrade, then dosu setup, and accept setup's offer to study your past sessions. Choose a migration you've worked on and ask your agent what changed. Check the facts it retrieves against the code, then save any correction to your team's Library with the migration date and a source your teammates' agents can verify.

Found this article helpful?

Share it with your network to help others discover valuable insights.

Want more like this? Subscribe via RSS

Related Articles

Ready to transform your workflow?

Join leading organizations using Dosu to automate documentation, streamline support, and empower development teams to focus on building great products.