This is what I've been saying for years. I attended a tech demo of an AI assistant for a car's owners' manual. But they trained the manual into the model. Which means not only does it need retraining every edition, but it's imperfect. The models need to be trained to fetch and use documentation not vaguely recall infinitely many concepts. I would much rather have 27B parameters on how to code than 26B on stuff like numpy function listings. It's like they approached the problem from a closed notes hand written coding exam. Everyone hates those.
I wrote this, after many iterations consulting for various companies. Documentation is queryable in single digit ms, append only log, etc. Has worked exceptionally well for my projects
Something I started doing recently was writing out principles instead of memories.
Essentially patterns the agents need to always think in. I also implemented a versioning system to the principles that need to be quoted in any comments which are there in the code. That way, when my principles evolve, so does the code.
I did package it up in a way that I can share it with friends [0]. System still evolving, but the last two-ish months that I've used it has served me really, really well, And it's been even better with the latest models.
I've had surprisingly good adherence from agents on this technique.
It sounds interesting, but this looks too much like a marketing website (huge text everywhere) and not enough like a documentation website, so it was hard for me to see how it might work.
My thought though has always been that I don't want there to be agent-only designated documentation.
I use mattpocock/skills and that generates ADRs (Architectural Decision Records). That only uses skills, including a setup skill that will write a few pointers in AGENTS.md. I always have a CONTRIBUTING.md to document development flow and a CODING_STANDARDS.md. Between those and the README.md and architecture documentation and commit messages the agents seem to be able to find and use docs and keep them up to date. We are also writing a lot of specs and putting those in Github issues.
I never understood memory solutions for coding agents.
You have the whole session history right there. One recall skill and some JSON parsing gets you grep over perfect memory. Why would you ever use more tools to spend more tokens to construct a imperfect memory next to your session history?
I will say, having built something similar for tracking 'memory' and items at home, it can quickly consume your tokens when dealing with both reading and updating, keeping stale info relevant etc.. when the amount of data starts to grow. Smaller tasks can balloon in their token cost as documents a read, updated, collated, refreshed etc..
however, I have found keeping a good solid reference to my home infrastructure, services, ci/cd setup, hosts, storage , networking etc.. really works wonders as a set of 'memories' to share across projects that I expect to be tested / deployed / acceptance tested etc.. using the home infra bits and pieces.
For every prompt and response, I extract each semantic statement. Map its reasons in a Whybase proposition tree -- a recursive proposition tree where each atomic statement is proposition with one or more premises (atomic statements, which also stand alone as propositions). Then I map each statement to the relevant code, hinted at by tool calls and git commits. Every time an agent touches that file or directory, a hook triggers in Claude Code that queries the codegraph db for the mapped statements. This helps the agent remember something I said in June when it revisits the code in July.
What if the June conversation is outdated and is no longer applicable by July? Do you have a mechanism for dropping older, subsumed propositions in your database?
Github Copilot had the idea of attaching memory to files, and if the file hash changes the memory is automatically dropped (not sure if they still do it). This means they are overly eager to drop stuff (even if the file change is just cosmetic), but at least they don't accumulate outdated cruft too much. (a memory can still be outdated if it was invalidated by a change in another file though)
Yes, the code and spec system are refined until they agree. How this is done is a work in progress. Think of my statements over time as the raw material into an evolving spec. The spec is refined as you learn and the code evolves. You and the agent loop until the code matches the spec.
Yes, this makes sense. Memory is an uncurated and often opaque system of arbitrary past discussions. It can help, it can harm. Accurate documentation in the other hand is only beneficial.
Peter Naur argued in his famous essay Programming as Theory Building that documentation alone cannot fully capture or preserve the complete mental model behind a program.
However, AI works differently from humans in that much more of its working context has to be made explicit. Because of that, there may be some fundamentally different way for AI to maintain or reconstruct a program’s overall model.
This is what I've been saying for years. I attended a tech demo of an AI assistant for a car's owners' manual. But they trained the manual into the model. Which means not only does it need retraining every edition, but it's imperfect. The models need to be trained to fetch and use documentation not vaguely recall infinitely many concepts. I would much rather have 27B parameters on how to code than 26B on stuff like numpy function listings. It's like they approached the problem from a closed notes hand written coding exam. Everyone hates those.
I wrote this, after many iterations consulting for various companies. Documentation is queryable in single digit ms, append only log, etc. Has worked exceptionally well for my projects
https://github.com/isaachinman/encephalon
Something I started doing recently was writing out principles instead of memories.
Essentially patterns the agents need to always think in. I also implemented a versioning system to the principles that need to be quoted in any comments which are there in the code. That way, when my principles evolve, so does the code.
I did package it up in a way that I can share it with friends [0]. System still evolving, but the last two-ish months that I've used it has served me really, really well, And it's been even better with the latest models.
I've had surprisingly good adherence from agents on this technique.
[0] https://principledriven.dev/
It sounds interesting, but this looks too much like a marketing website (huge text everywhere) and not enough like a documentation website, so it was hard for me to see how it might work.
Agreed, and this seems better.
My thought though has always been that I don't want there to be agent-only designated documentation.
I use mattpocock/skills and that generates ADRs (Architectural Decision Records). That only uses skills, including a setup skill that will write a few pointers in AGENTS.md. I always have a CONTRIBUTING.md to document development flow and a CODING_STANDARDS.md. Between those and the README.md and architecture documentation and commit messages the agents seem to be able to find and use docs and keep them up to date. We are also writing a lot of specs and putting those in Github issues.
Why would I need to install your tool for that? It could be an instruction living in AGENTS.md or with some sort of hook to remind the agent.
I never understood memory solutions for coding agents.
You have the whole session history right there. One recall skill and some JSON parsing gets you grep over perfect memory. Why would you ever use more tools to spend more tokens to construct a imperfect memory next to your session history?
I just don't get it.
I will say, having built something similar for tracking 'memory' and items at home, it can quickly consume your tokens when dealing with both reading and updating, keeping stale info relevant etc.. when the amount of data starts to grow. Smaller tasks can balloon in their token cost as documents a read, updated, collated, refreshed etc..
however, I have found keeping a good solid reference to my home infrastructure, services, ci/cd setup, hosts, storage , networking etc.. really works wonders as a set of 'memories' to share across projects that I expect to be tested / deployed / acceptance tested etc.. using the home infra bits and pieces.
For every prompt and response, I extract each semantic statement. Map its reasons in a Whybase proposition tree -- a recursive proposition tree where each atomic statement is proposition with one or more premises (atomic statements, which also stand alone as propositions). Then I map each statement to the relevant code, hinted at by tool calls and git commits. Every time an agent touches that file or directory, a hook triggers in Claude Code that queries the codegraph db for the mapped statements. This helps the agent remember something I said in June when it revisits the code in July.
What if the June conversation is outdated and is no longer applicable by July? Do you have a mechanism for dropping older, subsumed propositions in your database?
Github Copilot had the idea of attaching memory to files, and if the file hash changes the memory is automatically dropped (not sure if they still do it). This means they are overly eager to drop stuff (even if the file change is just cosmetic), but at least they don't accumulate outdated cruft too much. (a memory can still be outdated if it was invalidated by a change in another file though)
Yes, the code and spec system are refined until they agree. How this is done is a work in progress. Think of my statements over time as the raw material into an evolving spec. The spec is refined as you learn and the code evolves. You and the agent loop until the code matches the spec.
Yes, this makes sense. Memory is an uncurated and often opaque system of arbitrary past discussions. It can help, it can harm. Accurate documentation in the other hand is only beneficial.
Peter Naur argued in his famous essay Programming as Theory Building that documentation alone cannot fully capture or preserve the complete mental model behind a program.
However, AI works differently from humans in that much more of its working context has to be made explicit. Because of that, there may be some fundamentally different way for AI to maintain or reconstruct a program’s overall model.