Designing AI memory that stays accurate
Cabin and Mo both remember things about the people who use them. Someone mentions they are changing jobs, or that their sister lives in another city, and a week later the conversation can pick that up without being told again. Done well, this is the feature people notice most. Done badly, it is the one that makes them stop trusting the product.
The common mistake is to treat memory as storage. Save everything, fetch the closest matches, put them in the prompt. It works in a demo. It fails over months.
Why storing everything fails
People change. They move, switch jobs, end relationships and change their minds. A memory that only ever adds will, in time, hold both “works at a bank” and “left the bank in March”. Retrieval then brings back whichever one happens to match the current message best, and that is often the old one.
The model cannot tell which is true. It sees two facts and does its best, which can mean asking about the old job as if nothing happened. For a companion, that is worse than forgetting. Forgetting feels like a gap. Remembering the wrong thing feels like the product was never listening.
So we design for accuracy over volume. The aim is a memory where everything it holds is the current truth, as far as it knows.
Recall: bring back what matters now
Recall is the read side. When a message arrives, the question is which few memories would change the reply. Usually that is a handful: the person’s situation, the decision they are working through, how they like to be spoken to.
Putting too much into the prompt has real costs. Replies get slower and more expensive, and the model gets more chances to latch onto something irrelevant. Five memories that matter beat fifty that are loosely related.
Supersession: replace facts that changed
Supersession is the write side, and it is where accuracy is won or lost. When a new fact arrives, it is not simply stored. First we check whether it updates something already held. If it does, the old fact is superseded and stops being recalled, and the new one takes its place.
Some cases are clear, like a new job replacing an old one. Others are not. “I’m thinking about leaving my job” does not replace “works at a bank”. It adds a plan on top of it. Getting that distinction right is most of the work, which is why the write path deserves more design effort than the retrieval path.
It also helps to keep superseded facts instead of deleting them. The history shows what changed and when, and it lets you undo an update that turned out to be wrong.
Contradictions cost trust
Each contradiction a person sees is a small sign that the product does not really know them. A few are enough to make them share less, and a companion that people do not confide in has nothing to work with.
That is why memory changes go through the same offline evaluations as prompt and model changes. The most useful test cases are conversations where something changes partway through. The check is simple: does the reply reflect the latest state, or an earlier one?
What we would tell another team
Decide early what counts as a fact worth keeping. Spend more effort on writing memory correctly than on retrieving it cleverly. Treat an update as a replacement, not an addition, and keep the history so a wrong update can be reversed. Judge memory by whether today’s replies are right, not by how much is stored.
Memory is a promise to the person on the other side: what you told us, we kept, and we kept it current. It is worth building carefully.
Work with usNeed something like this built? Say hello