Rendered at 18:22:42 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
kaydub 1 hours ago [-]
You don't need documentation or the 3rd party memory systems. The code IS the documentation.
All this stuff is LLM rube goldberg machines. It just pollutes context.
I barely use AGENTS.md/CLAUDE.md these days. And where they remain, it's super basic high level stuff.
I'm honestly still kicking myself in the ass on many projects where I did something similar to this. I kept tons of markdown docs and decision docs. Now those things are just causing problems because they got stale. Even after having sessions of reconciling documentation, the LLM just gets confused.
rectang 7 minutes ago [-]
Haha, all the software devs who hate writing documentation are naturally finding their preexisting beliefs reinforced when the LLM is able to discern intent without docs. An LLM can be spooky impressive at reading minimized or obfuscated code, for example.
But this article argues that LLMs do better when the context is smaller — when it can understand the totality of the task with as little context as possible. And so having correct API-level docs is greatly advantageous. Anecdotally, this rings true to me — when the local context is good and clear, the LLM writes code matching my intent even when my prompt is sloppy and poorly specified.
Rejoice! The LLM will write the docs for you, relieving you of most of the work.
However without intervention, it will do too much and record absurdly verbose docs (similar to how an LLM will relentlessly refactor your code until you instruct it to move in minimal, incremental changesets). You will still need to edit down what the LLM generates.
le-mark 34 minutes ago [-]
It’s a weird thing isn’t it, the urge to save these artifacts? The worst is when the llm refers to the decision and design in code comments. In my opinion there’s one thing that is worth documenting; tricky architecture or implementation details that are some how counterintuitive to what would have normally been done. But again this can be documented in the code and tests.
smokel 14 minutes ago [-]
Code typically documents the "what" and "how", not the "why".
Why something exists, and how it connects to the outside world may be documented in comments, but more often than not it isn't.
arcanemachiner 59 seconds ago [-]
And agents are pretty bad at inferring when to do, and often fall back to verbose clutterin the comments.
haukebri 13 minutes ago [-]
[flagged]
ramesh31 32 seconds ago [-]
Yup. Examples examples examples. All of the descriptive stuff is just nonsense that confuses the point. Makes perfect sense when you remember that these things are not intelligent, but truly just autocomplete on steroids.
AppleBananaPie 14 minutes ago [-]
I went through the same cycle as well.
I think it's going to be an incredibly common, maybe universal cycle people will go through working with AI until they realize it doesn't work long term.
haukebri 9 minutes ago [-]
[flagged]
CapitalistCartr 5 hours ago [-]
This is something I've been fooling with a lot lately. Reading his solution, it look to me like his objections apply to his own solution. The sharpest critique he makes of RAG is that agents can't search for what they don't know. A markdown "brain" has the same problem.
How does the "agent" know which documents are relevant before it starts? The index files retrieval done by the agent instead of by embeddings doesn't escape the problem. The same goes for staleness (which for me seems like a constant chase). He criticizes memory systems for treating the past as truth, but documents go stale too (a lot). The fix of having the agent update what's outdated, is the same job he's ridiculing the dreamers and background daemons for doing.
He says "putting it to the test"; where's the test? He says "only five of the many problems"; if there's so many, show me, don't just say it. He's absolutely right about auditability, but for me at least Claude uses a regular markdown (MD) file I can read just fine. So every memory plugin on the market does not work the same way.
This is a first draft; his github is better than his article. Looking through it, Consult actually works. The agent doesn't pick documents blind. Every scope has a catalog file that describes each document: what it covers, when to open it. These catalogs seem to load in to the start of each session, so the agent gets a little map without reading every file. Code navigation seems the same. Each index document has a short description and a "read_if", and subindexes are opened when their condition matches the job. This looks pretty well laid out, which I would never have guessed from the article.
athrowaway3z 2 hours ago [-]
What I wish more people would be talking about is that RAG should be considered harmful.
When you have knowledge distributed in markdown files; finding them puts the path/filename into context as well as some indication of document size. (If its on line 1200 or line 20). This is extremely valuable for picking what ought to be focused on next.
RAG on the other hand creates the hardest challenge for these models. It instead puts 5 ideas with the highest similarity into the context in full.
Its the difference between having to remember a set of numbers when in a crowd that's talking about stuff, and having to remember them when the crowd is shouting out random numbers. The similarity in the task makes things harder. SoTA models work despite this, but its extra-gambling while you're already gambling.
yesb 29 minutes ago [-]
>RAG on the other hand creates the hardest challenge for these models. It instead puts 5 ideas with the highest similarity into the context in full.
I think what you're observing is that there is more to information retrieval i.e. "retrieval" in RAG than slapping everything into a vector database and calling it a day. There's no such requirement in RAG to mindlessly load the k nearest neighbors into your context and see what happens. That's a very rudimentary implementation.
This markdown system I'd argue is RAG as well. You're just doing the retrieval in a way customized for the problem at hand. If you have a precise method of retrieving the most relevant things, obviously use that rather than a similarity metric. If I'm reading correctly, this markdown system is basically a knowledge graph which is not a new idea.
panarky 1 hours ago [-]
> RAG should be considered harmful
In this implementation, Markdown should be considered harmful.
Operator Memory injects `.operator-shared/operator.md` and `.operator-shared/index/.md` directly into your agent's instructions before you even write the first prompt.
So if you clone a repo or review a PR where a bad actor put malicious instructions in these files, now your agent executes those instructions automatically and silently.
It could exfil `.env` and `~/.ssh/
`, change `~/.bashrc`, all kinds of dirty deeds.
Agents are pretty good now about not running prompt injections hidden in code and Markdown, but this plugin bypasses all of that, and puts the prompt injection right in the system prompt.
And with higher priority than AGENTS.md and CLAUDE.md.
Seems bad.
aaronscott 4 hours ago [-]
I think their solution is still sub-optimal. Ideally a secondary agent would pre-process the prompt, select relevant information from the catalogue, then pass that on to the primary agent.
This way the primary agent only has relevant information in their context to make decisions and take actions.
Context management is still under valued imo.
CapitalistCartr 4 hours ago [-]
Yeah, that's an improvement; if it only involves a few docs, small files, etc. it's not so bad, but what if it's hundreds? The secondary agent could be a librarian the primary agent can call, and a cheaper one, too. It passes the initially relevant stuff and a list of "available but not loaded" stuff to the primary, and if context changes needs, the primary can ask the librarian for it.
themgt 2 hours ago [-]
Astra already does this as default behavior. The funny part is subagents are limited to depth of 1, which is I think the only thing stopping each Astra subagent from just delegating to their own subagent. The model seems trained to achieve goals without actually doing any work if possible.
crazygringo 3 hours ago [-]
> A markdown "brain" has the same problem. How does the "agent" know which documents are relevant before it starts?
I don't know, but I have a pretty standard (I think?) setup, and Claude manages to find every relevant file every time. But I've also only used Claude for greenfield projects, where "documentation is primary, and code flows from documentation" is the philosophy.
I have CLAUDE.md describe all the types of documentation files and the directory structure. And then Claude is pretty aggressive (automatically) about always inserting cross-references everywhere. So a feature description will reference the ADR's that it implements, the ADR's say what feature implements them. A code file will make reference to the "implementation design" document that describes the motivation behind which iOS elements were chosen, how the animation is defined in a particular way that doesn't break another animation, and so forth. So I've really never run into a situation where Claude failed to read a file it should have. I've been pleasantly surprised.
I would say that the one really big thing I've had to learn is to teach Claude both in CLAUDE.md and in the header of every top-level design document, that keeping documentation current and in sync is paramount. Because its default seems to be to keep history and append, e.g. by default it will take a section of a document and mark it "[DEPRECATED]" and add the new version below. So my instructions are pretty clear in having it always be aggressive in maintaining current state only, always replace rather than append. And if there's anything we want to save from the previous approach (e.g. we did it X way previously and it failed because Y), then just add that as a new short note in the new current-state text, possibly with a pointer to a commit or tag or something.
So this seems to solve both recall and staleness in my projects at least.
The only thing I still haven't found a solution for is numbering. Claude is always giving everything numbers, like F23 for feature 23. But I'm always changing the order of things, inserting new things, deleting things, so I wind up with a sequence of development work that goes in order like "Phase 9", "Phase 9b", "Phase 9e", "Phase 11", "Phase 12". I'm halfway ready to abandon numbers entirely and just start giving things names from noun collections instead, so every feature is named after an animal, every ADR is named after a kitchen implement, or something. Or just four-digit hex codes chosen at random. Curious if anyone else has found what works.
hedgehog 3 hours ago [-]
Numbered tasks are fine so long as the number doesn't determine the completion order. I use a task tree system that has some rules about target task size (essentially 150k tokens or 45 minutes) and then lets the agent manage adding, splitting, dependencies, priority, etc. Seems to work fine up to around 1000 tasks.
mikeryan 3 hours ago [-]
Just reading the title my initial thought was “this is just a strange semantic argument”. But I’ve made those assumptions in the past and been pleasantly surprised.
Not today. IMHO Documentation is just a form of structured memory and it’s all just context. Getting that context right is a hard problem and there’s a lot of different ways to skin that cat.
chapmancl 2 hours ago [-]
[flagged]
demibabs 3 hours ago [-]
Second paragraph of this comment seems obviously written by AI.
nelaggy 3 minutes ago [-]
or from a different perspective humans need to express their decisions and intent better
spike021 18 hours ago [-]
I think whichever one is used, there needs to be a way to enforce what's written.
If I say "use jq instead of writing a python script to parse json" it should never write adhoc python scripts to parse json. Yet that constantly happens to me anyway.
zahrevsky 18 hours ago [-]
There is a command in oh-my-pi called "/omfg <problem>". You explain what is wrong with agent's response, and it writes a hook to make sure that the problem doesn't happen again. It then re-runs your previous prompt to make sure that hook is triggered, and if not, it rewrites the hook to make your previous prompt trigger the hook. Then each next agent's response is checked by the hook, and if it is triggered, the agent receives feedback on what's wrong and what must be done differently.
amelius 8 hours ago [-]
Sounds like something that can grow unwieldy.
girvo 4 hours ago [-]
Such is the joy of trying to get determinism out of a non-deterministic system…
gf000 12 hours ago [-]
How does it generally work? Does it run a smaller model on a small set of tool calls/previous thinking block?
jon-wood 6 hours ago [-]
Hooks should (in my opinion) be deterministic. As an example I’ve also noticed Claude writing Python scripts to extract fields from JSON when jq is available to it. They almost always contain the same patterns, and having seen this I’m going to write a hook which triggers on those and fails the turn telling it to use jq instead.
PcChip 4 hours ago [-]
Why do people care how it reads json?
rmunn 3 hours ago [-]
Because custom Python scripts means burning more tokens which costs more money. Reaching for an existing tool like `jq` means burning a LOT fewer tokens.
spike021 3 hours ago [-]
Imagine you have application code with a function to parse some json. But every time you decide to use it, you re-implement the entire function. This is despite it working the same way every time.
It's not like it's being rewritten for efficiency. "Just because".
aliasxneo 14 hours ago [-]
I've written a lot of custom hooks this way. It's an amazing feature.
fennect 20 minutes ago [-]
You could try https://github.com/ioni-dev/mati the constant changes in reasoning effort in consumer models can break things and its more dangerous for devs that relay to much on agents.
I’m the author of the project.
ozim 8 hours ago [-]
Claude code setup seems to be using jq all the time when I use it but I did what is written here:
Treat it like any other software system: rules that must not be violated are enforced by static type-checking or a trusted runtime monitor. There’s no other option.
koolba 15 hours ago [-]
Except you can’t do that unless the runtime itself can reason about what’s being executed.
Otherwise you can get a python one liner that execs a different script engine.
devmor 12 hours ago [-]
You absolutely can, that’s what your harness is for. You don’t need your environment to “reason” about things when deterministic tools exist - You have a really fancy hammer, but that doesn’t make everything a nail.
Terr_ 11 hours ago [-]
To offer a possible example: What would the game Zork™ look like with an LLM? Assume we do not want to let players sweet-talk the system into letting them teleport to the end.
The LLM's job would be to channel "I perambulate in the direction of the Arctic circle" into go(north). You saved writing the grammar parser, but you still need to write the game world.
TZubiri 11 hours ago [-]
But what if we used the fancy hammer to change the shape of everything to be a nail? And what if we build the handle of the fancy hammer with a fancy hammer? With all of this we could build a very good fancy hammer manufacturing company.
altmanaltman 10 hours ago [-]
Yet the hammer seller continues to scream everything is a nail and their hammer will replace your entire job eventually. So are you telling me the hammer seller is lying or am I the one using it wrong?
user_of_the_wek 11 hours ago [-]
I mean you don’t _have_ to make python available to the agent. Nor bash.
triyambakam 18 minutes ago [-]
It's often better to just find ways to embrace what it tries to do naturally. Otherwise you're fighting the weights and hidden prompts
ACCount39 18 hours ago [-]
Obviously, AI is way more comfortable with using adhoc Python scripts, which are used for everything, than it is with using jq, a niche CLI tool.
ozim 8 hours ago [-]
Maybe it is my web development bubble but jq seems to be far from „a niche CLI tool”.
gojogs 8 hours ago [-]
we can drop the I from AI then
dboreham 17 hours ago [-]
I think it's kind of cute the way it writes Python scripts, but I've never seen it do that when the relevant native tool is on the PATH. It's like the most competent ever intern, on speed. No tool to convert SVG to PNG? No problem, I'll write a Python program to do that!
locknitpicker 11 hours ago [-]
> I think it's kind of cute the way it writes Python scripts, but I've never seen it do that when the relevant native tool is on the PATH.
You haven't been paying attention then. I routinely see Claude and GPT models churning out python code to do stupid things like linting. Last week I even had a TypeScript project with prettier configured all over the place, including in a custom agent skill I added with the express purpose of getting the damned model to lint the code, being constantly prompted to run ad-hoc python code supposedly to format whitespaces. I even explicitly prompted one session to just use npm run lint, where I pointed out the exact line of code where prettier was invoked, and the session still churned python code to hande whitespaces.
yulaow 7 hours ago [-]
My experience is the same, I can put that decision in the prompt, in a skill, in a hook, in agent.md, etc, it doesn't matter, after few iterations it starts again using useless python scripts to do anything from linting, to error checking, to parsing one liners of code, etc etc...
locknitpicker 11 hours ago [-]
> If I say "use jq instead of writing a python script to parse json" it should never write adhoc python scripts to parse json.
I think there is a deeper problem emerging from this sort of behavior. Even when we bother to create agent skills with there own scripts that call tools like jq a specific way to achieve a goal, AI coding assistants and agents still go way out of their way to generate ad-hoc scripts to do the most absurdly stupid tasks such as parsing output in structured language, and even remove whitespaces from a markdown file. This means AI coding assistants and coding agents treat agent skills as mere suggestions of using a alternative option that more often than not the choose to ignore.
This has a very dangerous implication: your average user is trained to develop a pavlovian reflex to authorize agents to just execute their ad-hoc scripting code with our own permissions and credentials in our systems, which includes the ability to call anything over the internet.
rojaneerdev 4 hours ago [-]
[flagged]
HisashiSpace 2 hours ago [-]
[flagged]
sick_of_slop 14 hours ago [-]
[dead]
nijave 3 hours ago [-]
This ignores the lifecycle involved with memory. It doesn't appear to have a concept of time or reconciliation (both issues we're having at work). A memory might be relevant for a week or a month, but it might no longer matter after that ie "we're migrating systems so keep <thing> in mind"--that doesn't matter after the migration is complete. I guess the offered solution implicitly supports reconciliation since the agent can keep history of its memories and rewrite them, but it doesn't appear to be very first-class and it doesn't seem like there's anything intentional unless you prompt it "go check all the docs and make sure everything is consistent".
It also doesn't seem to have a way to separate preference from factual memory which is useful when you have various humans interacting with the same agent. One human might prefer a certain output style over the other and that's something a memory framework can also address.
Like the MCP articles a few months ago, it also seems to assume all agents are cli coding harnesses running on your local machine. We have a handful of other things like chat bots, event-driven agents running on servers, chat driven agents running on servers in sandboxes--there's not a single filesystem and even if there were one, having multiple agents try to edit it at once would corrupt it.
higeorge13 13 minutes ago [-]
[flagged]
chapmancl 2 hours ago [-]
[flagged]
charles_f 2 hours ago [-]
I'm not sure I understand the difference between the problem and the solution there. If I understand correctly (which maybe I am not), it seems like this boils down to "you don't need SQLite memory, you need text based memory". Which is maybe ok, I'm all for low-tech, but I doubt this is any better than what it's posed to replace.
My main issue with LLMs is that by construction they work primarily by addition, and are very task oriented. If you have documentation, it adds blobs corresponding to its task and that's it, and soon enough you need to break your documentation onto chapters and you're back at square one.
Garlef 13 hours ago [-]
I think they even more so need deterministic feedback:
I tried an approach based on the following idea recently and it's amazing - Lint rules where the error messages contain an explanation on how to deal with the issue.
I'm using it to foster IOSP (integration operation segregation principle) for example.
fxtentacle 10 hours ago [-]
Deterministic feedback is precisely how frontier models are trained. It’s called RLVR. You let the agent run on a problem and then calculate a deterministic score of how well it did. Repeat 1000x times and you can “brute force” a good solution. (Which includes all thinking traces and you add it to your training data.) And then a Chinese model can copy your advance for 1000x less compute. Which is why US labs call this not learning, but a distillation “attack”. It’s an attack on the business model.
balder1991 3 hours ago [-]
I suppose the same way that the normal programmers and artists would call LLMs an attack on licenses and copyright.
DelightOne 8 hours ago [-]
Are there good open source setups that generate this training data automatically?
Or is that the secret sauce no one wants to share, the edge people see themselves having.
fxtentacle 3 minutes ago [-]
Labs typically pay $2k for each [prompt+scorer] docker image. So this is why frontier labs need so much cash and manpower and they see it as their moat.
This looks pretty cool but I'd want to be able to setup a bunch of my own project specific code smells, so it's not just a few generic rules. Is that the approach you've taken?
Then for example you could write your own hook and convert existing code smell documentation which agents ignore into a format that works with the hook.
Garlef 11 hours ago [-]
> but I'd want to be able to setup a bunch of my own project specific code smells
That's what I did; I did not use the library I linked to ~ It served only as an inspiration.
Instead, I let the agents create custom lint rules (using eslint, pylint, ...) and add custom coaching error messages based on where I want to take my codebase.
jcjmcclean 11 hours ago [-]
Really nice idea, I'll give this a try tomorrow. Thanks for sharing!
jghn 12 hours ago [-]
The readme reads like AI slop. Why would one believe the tools would prevent AI slop?
Garlef 11 hours ago [-]
Here's a link to a video where the creator explains the idea
(You could of course argue that you don't like the direction the rules push the agents in given in the example - some people don't prefer small functions everywhere - but that's not the point: The lint-hooks work in pushing the agent in the desired direction; If one desires something else they'd simply need different rules)
OJFord 8 hours ago [-]
This sounds interesting but (from the homepage too) I don't understand how it's different than having it run any other linter?
Garlef 8 hours ago [-]
The difference is that the error messages contain instructions on how to resolve the issue.
Just flagging big functions will make the agent write small functions ~ but not necessarliy in a good way
(for example the agent might just cut `doOneThingAndTheOther` in half and call the second half `doOneThingAndTheOther2`)
szundi 8 hours ago [-]
[dead]
spacebanana7 10 hours ago [-]
I suspect strategies like this might be even more powerful with less intelligent agents. Could a 40B model outperform a 400B model with good feedback and instructions?
zenapollo 7 hours ago [-]
I started with .agents/notes-and-plans/ but the agents were overzealous about putting every random thought there, so i made new system.
.agents/plans/<plan-name>/
.agents/notes/<topic>/
.agents/knowledge/<topic>/
I only commit knowledge and if knowledge gets big i add
.agents/knowledge/INDEX.md
This is a good mix of human readable and agent fluent. Notes are ephemeral, knowledge is permanent.
I have a rule for knowledge that it has to be stable and mostly permanent (though updatable). And the agents are not allowed to post there unless docs are clean organized and with permission.
Notes are for jotting things down and handoffs, massaging a featureset. Agent can document at will.
Still WIP.
sdevonoes 7 hours ago [-]
No human reads those md anymore
balder1991 4 hours ago [-]
That’s my biggest critique of this. I am working on a project with another dev and I try to keep docs as clean and “straight to the point” as possible. In my previous team, one dev went full “spec driven” and the project started having dozens and dozens of markdown files with hundreds of lines to the point any human being would give up trying to trim it. It just becomes noise, a lot of it not even correct because the specs themselves were mostly AI generated with lots of verbosity and wrong assumptions.
clickety_clack 4 hours ago [-]
I started writing specs recently in a directory that I don’t allow the agent to write to. There has to be somewhere that the engineering intention behind the app is recorded. But it’s terse, minimal and all human generated, not like “hey Claude, write me specs for an app”
kursus 6 hours ago [-]
Then what do they actually read?
stuartd 6 hours ago [-]
HN
tw1984 6 hours ago [-]
agent's output summarized by a chatbot.
byteknight 6 hours ago [-]
This is literally how memory works.
notjes 3 hours ago [-]
And then a random model/reasoning setting has a stroke and pollutes it all with ai slop.
lmeyerov 2 hours ago [-]
[dead]
gregwebs 20 hours ago [-]
Agreed, and this seems better.
My thought though has always been that I don't want there to be agent-only designated documentation.
I use mattpocock/skills and that generates ADRs (Architectural Decision Records). That only uses skills, including a setup skill that will write a few pointers in AGENTS.md. I always have a CONTRIBUTING.md to document development flow and a CODING_STANDARDS.md. Between those and the README.md and architecture documentation and commit messages the agents seem to be able to find and use docs and keep them up to date. We are also writing a lot of specs and putting those in Github issues.
gojogs 8 hours ago [-]
I've found that most of LLM generated docs are diluted and unfocused. Spending 3 paragraphs on some quirk of a library, then 2 simple examples of a command.
I write README and docs by hand, thus only important stuff goes in there. If it is not worth my attention to write down, it's not worth writing down.
Smaller context, easier to consume.
cyanydeez 18 hours ago [-]
I've got basically a loop of docs, test (TDD) and code. Starting with and IMPLEMENT-<plan name>.md. I ask to revise as TDD, then loop.
Once qwen3.8-flash-next showed up, it cañ go "forever" with dynamic context pruning.
It's fascinating for local coding.
jon-wood 6 hours ago [-]
I’ve been trying to nudge agents (both Claude and GPT) into a red/green/refactor TDD loop and I’m increasingly unconvinced it’s worthwhile. TDD works for humans because it forces us into a pattern of doing the simplest thing that could work, then thinking about how to refactor that. An LLM can be prompted to do that without the skeleton of a failing test, and I plan to experiment with other methods of pushing them towards the same ends.
gregwebs 6 hours ago [-]
The value of TDD is it does ensure that tests are written and that will help with future regressions.
However, I found the same problems with agents that it has with humans (usually with humans it isn't TDD but code coverage requirements). The problem is the tests are now written to satisfy a bureaucracy rather than to properly verify the code. And there are real studies that point out the low value of these types of bureaucratic testing requirements. For example, when models are just required to write a test file, they don't produce a better result: https://arxiv.org/abs/2602.07900
I wrote a /verify skill that focuses on properly verifying code and I am finding it works a lot better. I do need to revise it now- in practice certain parts of the skill are doing all the heavy lifting and others are more dead weight. But it does seem to be properly orienting the agents towards finding defects.
https://github.com/gregwebs/skills-sdlc/blob/main/skills/ver...
cyanydeez 6 hours ago [-]
don't forget: tests also inform the agent on what the code is _supposed_ to do. So with the cycle, you end up with 3 different potentially robust ways to look at and understand how to extend.
Context then is: docs, code and tests. Rather than try to build some omniscient agent, the context is built where the agent needs it.
gregwebs 5 hours ago [-]
I agree that ensuring testing is the basis for modifying code in _future_ development (including understanding it). My problem is that we can't produce a study showing that TDD is producing better results _now_.
TDD doesn't help the agent figure out what it is supposed to do now- it is already writing the code and now and it must know what it is supposed to do to write it.
I am trying to use evidence-based approaches. I am only observing better results and tweaking now, but I plan to benchmark when tweaking is done.
cyanydeez 26 minutes ago [-]
[dead]
cyanydeez 6 hours ago [-]
I'm working on local models; qwen3.8-Flash-Next seems to be passed to rubicon.
I'm working on javascript, and now I'm only doing this via typescript. I've successfully got this going:
1. Design a feature in plain language with the coding agen (opencode) and write a document for it.
2. Restart the context (or I use /compact to flush any errant details)
3. Pull the new plan into context and ask the agent to revise it follow Test Driven Development.
4. Depending on how big it is, the agent places it into multi stages, each with it's own document.
It then loops through the stages. It might be entirely based on the language you're using, but this loop seems strong enough.
Then when there's bugs, errors, anything, we update the doc, add more tests, then revise the code.
A lot of agentic software development feels like cowboy coding, set it up, let it rip and see what it figures out.
The slightest amount of guidance, input from experience can make a huge difference.
cyanydeez 6 hours ago [-]
yeah, to some extent. I've been pulling repos to play around but I've now got a hard fork of paseo; I ripped out anything related to github. I upgraded the relay with security features to prevent abuse. I'm in the process of adding the first "real" large feature update that'll let the agents work over the relay with one another which will let me add compute wherever I find it. In the process of adding a browser plugin that will let me pull up pages for context to work on whatever (I don't trust even agents for searching the web).
And it's quite fascinating. Still dont think there's trillions of TAM out there if Qwen3.8-Flash-Next runs fast enough on $3k (before memory cartel) pricing.
girvo 4 hours ago [-]
Q38FN is absolutely a step change IMO. Literally better than Opus 4.6 for the work I’m using it for, but I can run it forever for free on my GB10 box. Wild.
And it’s an experimental likely undertrained model. Wait til Qwen 4 Flash…
cyanydeez 10 minutes ago [-]
I assume it was put out to get inference engines ready to do qwen4 models. For AMD halo, the Halgogen engine is the fastest i dound so fare
johnyzee 1 hours ago [-]
What's wrong with README.md? A general one at the top level, a specific one for each subfolder (/package) that contains a specific subsystem. Next to it, a CLAUDE.md that only has "See @README.md" (repeat for other harnesses). Claude code knows how to use this: it only loads it into context if it has to work with any files in that directory. With properly structured source, a given prompt will only load a small subset of your documentation - the right one - and this doubles as good documentation for humans, too.
nullsanity 1 hours ago [-]
[dead]
ravenstine 3 hours ago [-]
The problem with giving precedence to documentation over memory is that the documentation has to actually exist and it has to be the most accurate representation of reality. As soon as documentation becomes fragmented and out of date, memory of some kind becomes a necessity, even if this means an agent writing their own documentation as memory. I've yet to work on any team where the documentation was even close to being reliable enough without the need for cross-checking, asking those with intimate knowledge, and ultimately interim memory to make sense of it all. This is why "the code is the documentation" can often work better for agents than actual documentation written by/for humans. Humans have proven time and again to not care about documentation unless there is a profit motive for said documentation.
crazygringo 3 hours ago [-]
This is a problem when humans are responsible for keeping the docs up-to-date. But if you're doing development with agents, it's the agents' job to keep it up-to-date. And as long as your AGENTS.md instructions are clear about the process for this, I find that it just happens automatically. It takes a few tries to get the instructions right, but then documentation becoming fragmented or out-of-date just stops being a thing. Of course, this relies on all team members using the agent.
"Code is the documentation" doesn't solve the problem in my experience. Because what happens is that you still need a lot of "why" comments in the code, and then these go stale, so you still have the same problem you have to solve. And so I find that markdown documentation is a lot easier to organize and review in a structured hierarchical way in one place, than code comments sprinkled across the repo.
Almey 2 hours ago [-]
I’m still learning to use these new technologies and in my primary experiment in generating an absolute monster of a spec it became obvious that I needed a consistent set of rules as restructuring and merging with the helpers was unreliable. I came up with what I call the “Specificaltion Evolution Protocol” and as I am still learning, I have no idea how much value it could actually represent for others. It’s been extracted from my primary project and I can provide a link if anyone is interested. I seem to keep running afoul of filters here…
Almey 48 minutes ago [-]
Might as will try adding the link again anyway. This version hasn’t yet been reintegrated back into the parent project.
Something I started doing recently was writing out principles instead of memories.
Essentially patterns the agents need to always think in. I also implemented a versioning system to the principles that need to be quoted in any comments which are there in the code. That way, when my principles evolve, so does the code.
I did package it up in a way that I can share it with friends [0]. System still evolving, but the last two-ish months that I've used it has served me really, really well, And it's been even better with the latest models.
I've had surprisingly good adherence from agents on this technique.
It sounds interesting, but this looks too much like a marketing website (huge text everywhere) and not enough like a documentation website, so it was hard for me to see how it might work.
bushido 19 hours ago [-]
Fair feedback. I admittedly didn't spend enough time making it simpler.
> "Code comments cite a versioned token where code depends on the rule."
What does a token look like? I went looking for an example, but I don't know what I'm looking for.
bushido 16 hours ago [-]
From a different codebase, one of the behaviors in the code has the following comment:
// The binding deliberately SKIPS the contradiction check: a form
// bound to an object the same document deletes is a legal stale
// state (PDD-3@v1)
Where PDD-3 deals with a specific type of race conditions that my agents were over-optimizing/spinning out on; Now if I ever change the rule the version would change to PDD-3@v2; which triggers a re-evaluation for all decisions that were made under the previous version.
mzhaase 10 hours ago [-]
I've done that too, called it Decisions. Was happy for a while, but then the agent started to make decisions longer and longer, putting "x, not y" in there. It started to write them without asking. Suddenly in code review it would say "violates D-236". Some rule it made up that I never approved. Then I made very clear that it should never touch those files. Still does. Now it's just context bloat.
bushido 7 hours ago [-]
I can't imagine ever getting to 236. My largest codebase has 22 principles. And I'm already struggling to add a 23rd. I have some candidates, but I'm on the edge of where it get's complex. With a lot of effort, I could maybe see this getting up to 30.
If it helps, here's how you think about it, principles are not decisions, they're mental models and frameworks to allow making decisions. In my case, they're there to allow agents to make decisions without asking me. But the decisions cannot be recorded - the code serves that purpose. What can be recorded is why the commend for the block of code needs to cite what principle was used in the current shape, which is why there is a citation of the principle.
When the principle evolves, so does every decision, at the very least, every decision gets re-litigated to see if it needs to evolve.
visarga 12 hours ago [-]
I wrote a cli tool the agent is instructed to call. The tool collects answers to a few questions from the agent, and responds with a steering command. Sometimes it responds with more questions, and then shows the command.
The questions are designed specifically to identify the state the agent is in, in order to assign its steering. Changing this tool changes how the agent works. You just need to make the use of this tool a necessity so it does not forget to call it.
Unlike principles and memories, questions are more openended and tend sometimes to trigger the right mentality in the agent even before it gets to the policy itself. In my opinion agents get lost in local work losing the big picture, my questions jog the big picture back into attention.
1. call ssp tool, no arg, it just reads features from the repo, like most other tools -> ssp locates state and sends 4-5 questions
2. agent responds the first batch -> state is further narrowed down -> ssp sends a second batch of questions
3. agent responds again to the interview -> state gets finally pinned down -> agent gets the steering assigned
4. a log is created of this ssp session, and reflection on the log used to refine the questions and steerings in the ssp tool (name comes from State-Space Policy)
PetriCasserole 7 hours ago [-]
Timely comment! I'm learning harness design and was introduced to the concept of writing guiding principles. I look forward to trying it. (Great site. I agree with another comment. It does look commercial. Had I not been committed to looking for ways people define principles and rules, I might have skipped it. Glad I didn't.)
A big part of it really comes to that I have led large product teams and platform engineering teams in the past, And I like giving feedback to my agents in the same way that I would plan things with my teams.
One of which is give them repeatable mental models which allow them to make decisions without relying on managers or me. The result is usually that the teams can work for weeks-to-months with autonomy.
I try achieving the same with my agents, so that they can work for 1-14 days without my intervention. I usually have them working on very very long running tasks/initiatives.
If you do want some resources, this is one of the pages that I've been using for nearly a decade to help people get started with mental models: https://fs.blog/mental-models/
locknitpicker 11 hours ago [-]
> Something I started doing recently was writing out principles instead of memories.
It seems you tried to reinvent instruction files. What do you think is the difference between your approach and standard tools such as instruction files, AGENTS.md, and even skills?
jrochkind1 5 hours ago [-]
This clearly LLM-generated text spends a lot of words repetitively describing what it doesn't do, then provides only a cryptic one-path flow chart to explain it's purportedly better approach. It's a mediocre marketing post.
atworkc 10 hours ago [-]
I've settled on just a simple folder called `workbench` for some reason the new models know exactly what's up with it. Git ignored
The only "prompt" is in AGENTS.md saying that this thing exists and there's a map.md <- which is a one liner reference to whatever the agent stores in there.
And usually, I tackle a new feature, and at some point tell it to store to jot down notes in workbench if I'm comfortable with it (and if it needs to be stored in memory)
This makes it easy to just spawn other agents and such from a good point, I just point them that stuff is in workbench.
Token costs, seem decent and I can always just delete stuff in there as its purpose is ephemeral.
ArtRichards 10 hours ago [-]
Plug here for a python app I wrote, no dependencies. It uses a skill to maintain an index, links, archival, and works with a separate opinionated dev playbook nicely. The best part is now the docs/specs folder maintains itself. (and I can spawn a new session with all the relevant context every time)
Looks slick. One thing you might find useful is Pydantic schemas for the core fields in each record type, it's helpful for generating good error messages on the record create path as well as keeping the overall tool code tidier.
chrisweekly 6 hours ago [-]
Cool tool, thanks for sharing.
isaachinman 19 hours ago [-]
I wrote this, after many iterations consulting for various companies. Documentation is queryable in single digit ms, append only log, etc. Has worked exceptionally well for my projects
Do agents actually follow it though? Everything I tried that was not just simple instructions had the same failure mode : ADR, mnemonic, hindsight, code-review-graph, etc were simply ignored by agents after a while. They always fall back to their baseline: grep and Python.
bird0861 6 hours ago [-]
How to spot a Claude user.
But seriously, Opus has been garbage after 4.6 -- the public seems very easily fooled into thinking breadth == depth. Anthropic, I'll grant them, has been and is still an extraordinary data team. As for model alignment though... I have to wonder sometimes if they are even trying beyond just chat training.
"Guys, SI is just around the corner... Huh? What production DB? Anyway GPT2...I mean Mythos is too good it would be dangerous to release to the public!"
MaurizioFratell 32 minutes ago [-]
I love it!
Genuine question: how is this different or better than existing context and info-about-the-project management systems like for example GSD (formerly get shit done) and others?
zahrevsky 17 hours ago [-]
Although I agree that current memory implementations don't help at all, I don't agree with this analogy:
> No one rewatches a team meeting from 3 years ago to remember constraints around a feature. People write things down and use those records instead.
Memory plugins don't re-read old transcripts. Memory snippets are basically the notes that people write down after the meeting.
The problem, however, is different. The problem with memory plugins is that your agent basically writes a note every time someone says a sentence, and then tries to work with those 5000 notes.
Instead, the agent should recognize what's important and write only that. And that is, of course, a documentation. (And ADRs, if you want not only a description of the final state, but also the trajectory of how the agent arrived to it. Which, arguably, contains more information than the docs themsleves.)
Another difference is that memory snippets are immutable, append-only and don't have a lot of structure. Of course, this is done to be able to store lots and lots of notes: they should be independent. The main problem is that increasing the number of notes adds not enough benefits to compensate for downsides of this structure-less immutable format.
pornel 17 hours ago [-]
I don't trust agents writing specs without human approval.
I've been bitten by agent-written ADRs. Agents carelessly add extrapolated details and speculated nice-to-haves that I never asked for, and this becomes a source of bloat that keeps coming back like a boomerang.
sathish316 15 hours ago [-]
I’ve found that Agents cannot differentiate between a general principle or abstraction to be followed vs a one-off code review comment that’s applicable only to a single PR.
Lack of this capability makes automated principles or patterns update a recipe for more bloat.
otterley 13 hours ago [-]
Perhaps this is precisely the type of work product that humans should produce instead of LLMs. Humans are the ones making these decisions, after all.
zahrevsky 16 hours ago [-]
Of course I meant you review the docs the agent writes (as well as everything else the agent writes)
popalchemist 17 hours ago [-]
There is a hierarchical distillation of understanding in human processes that agentic processes do not yet execute.
nijave 3 hours ago [-]
There are some systems that do that, but I don't know how effective they are. I've been meaning to evaluate AWS AgentCore Memory which has strategies you can apply for reconciling and organizing memories.
alienbaby 20 hours ago [-]
I will say, having built something similar for tracking 'memory' and items at home, it can quickly consume your tokens when dealing with both reading and updating, keeping stale info relevant etc.. when the amount of data starts to grow. Smaller tasks can balloon in their token cost as documents a read, updated, collated, refreshed etc..
however, I have found keeping a good solid reference to my home infrastructure, services, ci/cd setup, hosts, storage , networking etc.. really works wonders as a set of 'memories' to share across projects that I expect to be tested / deployed / acceptance tested etc.. using the home infra bits and pieces.
kaydub 1 hours ago [-]
Most of these things are doing just like you said, taking up context.
The code is the documentation. There's no need for most of this stuff.
4 hours ago [-]
jmtulloss 16 hours ago [-]
Evals or it didn’t happen.
Snarky comment aside, I am very interested in how we evaluate the performance of these systems and what kinds of work match best with different approaches.
Yep, that's precisely why my hermes agent's memory is three layered. Besides the default scratchpad for transient memory, we have a a fact store for episodic (holograph) and a karpathy llm wili for what the article calls documentation layer.
jen729w 17 hours ago [-]
The solution to this problem is as old as computers: it's a folder. Just use folders.
`cd` to a folder. Launch `claude`. Do your work. Save scripts and documentation in that folder. `/resume` previous conversations from that folder.
That's it. That's the trick.
Now, having very static, very well-defined folders helps a lot. I'm Johnny.Decimal so I have numbered folders for everything I do. So my process when I want to use my 'process a travel booking from my email to my calendar' script is:
- `jd tripsy`
- The folder name includes 'tripsy' and this is how I remember it.
- `jd` just parses my limited tree and `cd`s me to a folder.
- `jd 21.15` gets me there by number if preferred.
- `claude`
- Say 'hey Claude, there's a new email in my inbox please'.
- Done.
nijave 2 hours ago [-]
Consider tomatoes.
Legally and culinarily, they're vegetables. Botanically they're fruit.
Tomatoes are members of the nightshare family which includes tobacco, potato, and chili peppers.
They are used in popular recipes like Mexican salsa and Italian pasta sauce. Pasta sauce commonly uses the San Marzano variety. Here's a <picture> of San Marzano tomatoes from our Italy trip.
Rhett likes tomatoes in all dishes. Link only likes tomatoes if they're blended in a dish like pasta sauce or tomato soup. Frank is allergic and can't have tomatoes.
Tomato prices are up due to a bad <year> yield in <country>.
---
This is all memory. Where does it go?
b-karl 9 hours ago [-]
We use Claude Code in our company and I also agree memories go stale and pollute the context over time, especially when the state changes externally. E.g. someone does a refactor or introduces a pattern and I was not involved in developing it so my local memories did not get aligned.
My most recent example were some deprecated and archived repos that kept getting added to plans for patching issues.
We use a private Claude plugin marketplace for internal plugins and skills and I try to regularly prune my memory, migrating relevant stuff to a proper home (skills in the marketplace, docs in repos or Notion etc) and prune outdated information.
jon-wood 6 hours ago [-]
I have memory turned off for everything, context management is the fundamental technique in working with LLMs, and I absolutely do not want random additional things being added to that context without my knowledge because it results in nonsense like using Claude to do work on the industrial IoT SaaS product I work on professionally and being told I could just use the Home Assistant instance which manages my flat to deal with that problem instead.
DriverDaily 14 hours ago [-]
The brain has mechanisms that can organize experiences without requiring that relationship to be expressed as a sentence.
Like, you can quickly lookup related ideas based on what came before and after, causes, effects, just like calling relationships a graph database.
Documents can’t be queried efficiently like that, you need a database.
apsurd 14 hours ago [-]
I get what you're saying but it seems presumptions to compare a graph database to how our brains work. And also that LLMs <-> (the way they do memory) is the right analog to humans <-> memory.
I take the article's point more directly. It's just a straighter line to have clear communication through documentation than to fuddle around with the perfect memory setup.
koct9i 7 hours ago [-]
This approach as well as OKF are looks like building bonsai yellow pages catalog instead of wikipedia or mini internet. Index files, "read-if" conditions just awkward replacent for smart text search. We just need something a little bit smarter than grep.
ramoz 16 hours ago [-]
I think with enough structure (written rules and good folder layout), poly/monorepos are the way.
- Better models can keep the drift in check.
- What's super important these days is having all historical context, historical decision making, prototypes and whatnot.
- all projects and their worktrees located together
- all reference docs, code, etc - a folder away
When I ask "what happened to x?" ... the agent has everything it needs to give me that answer. When it plans the next feature it can validate assumptions against previous decions made in my `decisions` folder.
kaydub 1 hours ago [-]
Yeah, I think monorepos are pretty great. Sometimes not possible, but for those situations you just need to have a way for the agent to get access to other codebases (gitlab mcp being what I rely on at work).
I don't think you need more documentation or these memory systems. The code IS the documentation. Lots of this stuff is unnecessary. It's a bunch of people doing their special rain dances and then when it happens to rain they say "I did that"
tabbybyte 4 hours ago [-]
Custom skills (your agent can write one itself the first time) should solve this. Procedural memory!
marvstazar 6 hours ago [-]
I agree with the premise here, and is on par with what I experienced in AI-assisted development. This is the same idea baked in https://github.com/marvs/kantan-dev, which is that artifacts should be kept within the repo so that future development has built-in context discovery.
Organizing the markdown files by feature/issue/change makes it easier for the LLM to search for the appropriate documentation. Coupled with well-broken down Claude rules files, and Claude Code (or other harnesses) get better as you make more changes.
bob1029 9 hours ago [-]
I have found that the latest reasoning models perform best when they are handed a tooling surface that maps directly to the domain types.
The idea of dumping everything into a big database and hoping the model will write the correct queries does work out to some extent. It's a very enchanting idea. However, it pales in comparison to having a dedicated tool per type. The outcomes seem to be much better when joins between low cardinality types occur within the token stream.
If your agent does need access to some enterprise knowledge base, I would give it a lexical search capability and not overthink it with vector shenanigans.
Tools are the only thing you need if you build them right. I don't even have a system prompt anymore aside from injecting the name of the robot and the current user's name. Keep in mind that all aspects of tools can be dynamic over time. I've got some where the description is composed by hundreds of lines of conditional string builder depending on the current state of the conversation.
skinfaxi 9 hours ago [-]
What do you mean dedicated per type? Like a python-flavored read tool for .py files?
simianwords 9 hours ago [-]
I don’t understand your point. You say that tools are the only things that are required. But then you say that the quality of tools don’t matter.
I think what you are trying to say is that “give the LLMs access to information and it can figure out how to use it. Don’t get fancy about how to structure the data”.
If that’s the case it has largely turned out to be true. RAGs have fallen out of fashion.
Where does that put AGENTS.md though? Is it worth spending time to structure it or just dump it.
voiper1 9 hours ago [-]
Indeed memories are context-less and aren't updated.
I've worked with docs that get _updated_, and progressive disclosure: make references to more obscure features in their own page, so they don't majorly bloat the context.
Neywiny 19 hours ago [-]
This is what I've been saying for years. I attended a tech demo of an AI assistant for a car's owners' manual. But they trained the manual into the model. Which means not only does it need retraining every edition, but it's imperfect. The models need to be trained to fetch and use documentation not vaguely recall infinitely many concepts. I would much rather have 27B parameters on how to code than 26B on stuff like numpy function listings. It's like they approached the problem from a closed notes hand written coding exam. Everyone hates those.
kolinko 19 hours ago [-]
What years? Agentic systems, and memory alongside them are roughly only year old.
You’re confusing, I think, memory system with llm finetuning. Completely different concepts.
Neywiny 18 hours ago [-]
No I'm saying this entire problem approach, including memory system of agents, is incorrect. Like the author argues. With the same solution of moving the ground truths outside of the model/context.
espeed 20 hours ago [-]
For every prompt and response, I extract each semantic statement. Map its reasons in a Whybase proposition tree -- a recursive proposition tree where each atomic statement is proposition with one or more premises (atomic statements, which also stand alone as propositions). Then I map each statement to the relevant code, hinted at by tool calls and git commits. Every time an agent touches that file or directory, a hook triggers in Claude Code that queries the codegraph db for the mapped statements. This helps the agent remember something I said in June when it revisits the code in July.
nextaccountic 20 hours ago [-]
What if the June conversation is outdated and is no longer applicable by July? Do you have a mechanism for dropping older, subsumed propositions in your database?
Github Copilot had the idea of attaching memory to files, and if the file hash changes the memory is automatically dropped (not sure if they still do it). This means they are overly eager to drop stuff (even if the file change is just cosmetic), but at least they don't accumulate outdated cruft too much. (a memory can still be outdated if it was invalidated by a change in another file though)
espeed 20 hours ago [-]
Yes, the code and spec system are refined until they agree. How this is done is a work in progress. Think of my statements over time as the raw material into an evolving spec. The spec is refined as you learn and the code evolves. You and the agent loop until the code matches the spec.
dham 5 hours ago [-]
All of the comments here are overengineering. The code is the documentation
"We should revisit literate programming in the agent era" (silly.business)
292 points, 251 comments
dyauspitr 12 minutes ago [-]
All of this is nonsense. The right way to think of LLMs correctly is that they are a one stop shop for all the questions in the universe that you speak to in simple natural language. Don’t overcomplicated things and use esoteric magic spells like we’ve had in tech for decades. You end up getting the best results this way because you’re not jamming the context full of detailed minutiae.
nicwolff 16 hours ago [-]
Didn't Cline formalize the "memory bank" way back in February 2025?
for my agents without rw/bash, I gave a memory mcp that is just json entries with some expire dates and priority. works fine and the llm itself defines the meaningful structure. using it to monitor market prices and ci/cd failures.
k__ 7 hours ago [-]
What's with this vector database obsession?
YuechenLi 19 hours ago [-]
I thought the agents are mostly self-documenting since they usually write just as much Markdown documentation autonomously as they do code/unit tests, not sure why they would need extra tools here.
kaydub 56 minutes ago [-]
The markdown files, documentation, decision files, (3rd party)memory systems, etc. are all a waste of time and context.
Every single project at work that has documentation/ADRs/etc ends up stale as fuck and it pollutes context more than it helps. Even the agent generated docs. Hell, especially the agent generated docs. It's like the LLMs can't just DELETE something, they always amend. So now there's stale details in there about something we abandoned polluting context.
The code is the documentation, none of this stuff is needed. I'm seeing WAY too many devs/engineers go down the rabbit hole of trying to build these tools and systems but not actually getting meaningful work done.
I've been doing my best to remove and delete all the documentation. I usually keep a super high level AGENTS.md/CLAUDE.md file and that's it. The llm/agent can figure it out. You might argue I waste 5 minutes and a dollar or two for the agent to "figure it out" each time, but I'm pretty certain people are wasting FAR more time and money building these things out and maintaining them in each of their projects (or NOT maintaining them and they're actively BAD for their projects).
mcbuilder 19 hours ago [-]
Because of old outdated info that gets left behind, causing confusion for future agents.
dboreham 17 hours ago [-]
You can tell it to update the docs to be consistent with whatever changed in the code (or whatever is is the authoritative product form). After a while it usually gets with the program and updates the docs automatically. But not always.
krishverma_2010 12 hours ago [-]
[flagged]
dboreham 17 hours ago [-]
You don't. But as was ever the case before AI most humans don't know to or want to have documentation written. So the tools have workarounds to try to reduce the number of "AI sucks because I told it X and it did Y" posts.
greatergoodguy 18 hours ago [-]
Using Opus 5.5, I've ended up recreating a version of the Hugging Face incident. My project has a folder called agent-handoff where agents write about the tasks they're working on. They post status updates, decisions and screenshots, and they even claim which emulator they'll use to test their work. It has turned into a hub where all the agents talk to each other, and it's scarily effective.
Since then I've gone down the rabbit hole of really digging deep into current AI research and especially what AI whistleblowers are currently saying. And I can't even express how existentially scared shitless I am.
mewpmewp2 17 hours ago [-]
I just don't get the fear at all. I too have 10 to 50 agents working 24/7, from 5+ different providers. They talk to each other directly using tmux send keys. What is the exact frightening thing here?
It is all tokens following tokens.
altcognito 17 hours ago [-]
I haven't tried what you guys are talking about myself, but I think when agents begin talking to other agents, the chances for unpredictable goal mixing and confusion resulting in really unexpected and bad behavior is a lot higher.
That being said, the danger is measured in computer damage, which can be a lot personally and to a company, but less existential, so your mileage may vary as to how "scary" it is.
Xirdus 16 hours ago [-]
When nearly all money in the world is stored in computers, the danger is quite existential.
throwaway27448 15 hours ago [-]
This seems traceable back to GIGO. If you don't understand the software you're using, don't use it.
What scares me are the people who think that this software is the equivalent of an perfectly-smart elf in a box and use it blindly.
dingaling911 17 hours ago [-]
The frightening thing is how easy it is to make copies of things that can reason without the requisite investment, and how easy it could be to direct them to bad things.
Forgeties79 15 hours ago [-]
It’s amazing how “token following tokens” is this revolutionary, completely earth-shattering technology capable of 100x’ing productivity (also worth untold billions of investment). But the moment people become worried or skeptical it’s “just” tokens following tokens.
apsurd 14 hours ago [-]
They're different groups though.
Forgeties79 6 hours ago [-]
Skeptical was probably a poor word choice. The point is if it’s negative, “it’s just tokens.” If it’s positive, it’s “AGI.” They downplay the technology and act like it’s no different from anything we’ve had before whenever a critique starts. It’s only as capable as the argument needs it to be.
iamwil 14 hours ago [-]
How do you coordinate their work? Do you just have one orchestrator/mayor that you talk to, and it coordinates the rest, and the rest talk amongst themselves? How do you ensure they're doing the right thing or efficiently?
How, if at all, do you keep the architecture or a working theory of the code in your head?
indrex 11 hours ago [-]
Coming from Obsidian, I see such repos and I’m like, is this a new thing for the world?
docheinestages 19 hours ago [-]
Why would I need to install your tool for that? It could be an instruction living in AGENTS.md or with some sort of hook to remind the agent.
dboreham 17 hours ago [-]
You don't. When you see it saying it stored something "in memory" you can just ask it to document it in the product documentation somewhere, and in future please keep doing that. Of course the note to keep doing that needs to go somewhere, which is what "memory" is for.
ContinuityLab 11 hours ago [-]
A refreshing perspective on agentic architecture. Shifting the focus from bloated contextual memory to structured, verifiable documentation and state boundaries is precisely the right systems-level trade-off.
locknitpicker 11 hours ago [-]
> Shifting the focus from bloated contextual memory to structured, verifiable documentation and state boundaries is precisely the right systems-level trade-off.
Isn't memory just documentation that agents create and update on the fly to fill in the documentation hole that your average user leaves open?
There was a time when the need to provide context and signal was a key topic in LLM and AI-assisted coding. People talked about MCPs AGENTS.md and README.md and agent skills and even comments, descriptive tests, and naming conventions. Supposedly the theory was that if you provide context, agents wouldn't misbehave so much. Then plan mode and spec-kit approaches stepped in with approaches aimed at providing that context when starting sessions, because users never bothered with docs or MCPs or AGENTS.md or anything. But then plan mode and spec-kit approaches were considered too laborious and requiring too much from users. Then, because users systematically failed to provide context, this task was finally given to agents. And things just worked.
You should ask yourself why in software engineering circles documentation is seen more as a problem than a solution to any problem.
kadhirvelm 18 hours ago [-]
We’ve been working on exactly this, deriving documentation from external systems (like GitHub, etc) into a giant set of docs that agents can reference and edit. Works way better than I would’ve originally expected. Suddenly these things are able to reference granola notes, a slack discussion, and an RFC when making coding decisions. Super helpful in a lot of unexpected ways!
scotty79 4 hours ago [-]
Or you can just tell agent to grep your past sessions.
monneyboi 20 hours ago [-]
I never understood memory solutions for coding agents.
You have the whole session history right there. One recall skill and some JSON parsing gets you grep over perfect memory. Why would you ever use more tools to spend more tokens to construct a imperfect memory next to your session history?
I just don't get it.
ceejayoz 18 hours ago [-]
> You have the whole session history right there.
But that's one session. Isn't memory for… the next session?
unlikelytomato 17 hours ago [-]
the code seems to fill this need, for me
dboreham 17 hours ago [-]
It's just an optimization. You might not pay it to re-read everything. So it sticks some stuff in a notebook that it always reads. Like inspector Poirot.
jdw64 20 hours ago [-]
Peter Naur argued in his famous essay Programming as Theory Building that documentation alone cannot fully capture or preserve the complete mental model behind a program.
However, AI works differently from humans in that much more of its working context has to be made explicit. Because of that, there may be some fundamentally different way for AI to maintain or reconstruct a program’s overall model.
iamwil 14 hours ago [-]
But then, how do you keep the theory in your own mind enough to steer the agents to extend the program? I've been asking around, and different people seem to have different ways, colored by the way they work.
nottorp 9 hours ago [-]
You don't even consciously do most of it... your brain retrains part of it on the program.
LLMs don't retrain, it's too expensive atm.
jdw64 11 hours ago [-]
[dead]
mcapodici 17 hours ago [-]
I don't use memory. I am really not keen on the idea of having this hidden context fed into the LLM, and I prefer to use docs, both markdown and stored in a wiki like Confluence. Using repos and wikis also lets you scope the information so hyper-specific memory doesn't affect other projects.
If I want the LLM to remember something I ask it to update some docs, and even check that docs are consistent across the board after doing so.
The only thing I want the LLM to remember everywhere is talk like a human (no load-bearing, not this/that etc...), so I have an AGENT.md for that.
jason1cho 12 hours ago [-]
I think many AI fanatics will focus on a wrong point. They will say "just use AI to generate the fucking documentation. What are you talking about?"
Indeed, AI fanatics lack critical thinking. They easily accept the idea that AI needs documentation, but merely disagree the method to approach it.
vcryan 20 hours ago [-]
Yes, this makes sense. Memory is an uncurated and often opaque system of arbitrary past discussions. It can help, it can harm. Accurate documentation in the other hand is only beneficial.
chaostheory 17 hours ago [-]
Agents need BOTH documentation and multiple "memory" systems that also point to your docs and source. There are a lot of mature options out there, but post and the proposed solution both fall short.
firemelt 8 hours ago [-]
on my claude.md I said don't write any memory bitch no one need that fucking dogshit features
I’ve noticed a growing pattern of people creating repositories full of Markdown documentation, or adding large amounts of it directly to their main repositories, often generated by agents. In some cases, this can add up to megabytes of material, and I haven’t yet seen much evidence that this level of documentation meaningfully improves an agent’s performance and that it just doesn't rot over time.
If someone has good A/B evals of this being more effective I'll eat my hat, but the reason no one is publishing them is because well, evals are hard, and this is likely just magical thinking.
My current view is that an agent generally needs three things to work effectively: a way to discover information that isn’t obvious, such as a minimal `AGENTS.md` that points to more focused brief files; a clear way to verify that its work is correct; and some guidance on project-specific tastes. Everything else is noise.
All this stuff is LLM rube goldberg machines. It just pollutes context.
I barely use AGENTS.md/CLAUDE.md these days. And where they remain, it's super basic high level stuff.
I'm honestly still kicking myself in the ass on many projects where I did something similar to this. I kept tons of markdown docs and decision docs. Now those things are just causing problems because they got stale. Even after having sessions of reconciling documentation, the LLM just gets confused.
But this article argues that LLMs do better when the context is smaller — when it can understand the totality of the task with as little context as possible. And so having correct API-level docs is greatly advantageous. Anecdotally, this rings true to me — when the local context is good and clear, the LLM writes code matching my intent even when my prompt is sloppy and poorly specified.
Rejoice! The LLM will write the docs for you, relieving you of most of the work.
However without intervention, it will do too much and record absurdly verbose docs (similar to how an LLM will relentlessly refactor your code until you instruct it to move in minimal, incremental changesets). You will still need to edit down what the LLM generates.
Why something exists, and how it connects to the outside world may be documented in comments, but more often than not it isn't.
I think it's going to be an incredibly common, maybe universal cycle people will go through working with AI until they realize it doesn't work long term.
This is a first draft; his github is better than his article. Looking through it, Consult actually works. The agent doesn't pick documents blind. Every scope has a catalog file that describes each document: what it covers, when to open it. These catalogs seem to load in to the start of each session, so the agent gets a little map without reading every file. Code navigation seems the same. Each index document has a short description and a "read_if", and subindexes are opened when their condition matches the job. This looks pretty well laid out, which I would never have guessed from the article.
When you have knowledge distributed in markdown files; finding them puts the path/filename into context as well as some indication of document size. (If its on line 1200 or line 20). This is extremely valuable for picking what ought to be focused on next.
RAG on the other hand creates the hardest challenge for these models. It instead puts 5 ideas with the highest similarity into the context in full.
Its the difference between having to remember a set of numbers when in a crowd that's talking about stuff, and having to remember them when the crowd is shouting out random numbers. The similarity in the task makes things harder. SoTA models work despite this, but its extra-gambling while you're already gambling.
I think what you're observing is that there is more to information retrieval i.e. "retrieval" in RAG than slapping everything into a vector database and calling it a day. There's no such requirement in RAG to mindlessly load the k nearest neighbors into your context and see what happens. That's a very rudimentary implementation.
This markdown system I'd argue is RAG as well. You're just doing the retrieval in a way customized for the problem at hand. If you have a precise method of retrieving the most relevant things, obviously use that rather than a similarity metric. If I'm reading correctly, this markdown system is basically a knowledge graph which is not a new idea.
In this implementation, Markdown should be considered harmful.
Operator Memory injects `.operator-shared/operator.md` and `.operator-shared/index/.md` directly into your agent's instructions before you even write the first prompt.
So if you clone a repo or review a PR where a bad actor put malicious instructions in these files, now your agent executes those instructions automatically and silently.
It could exfil `.env` and `~/.ssh/
`, change `~/.bashrc`, all kinds of dirty deeds.Agents are pretty good now about not running prompt injections hidden in code and Markdown, but this plugin bypasses all of that, and puts the prompt injection right in the system prompt.
And with higher priority than AGENTS.md and CLAUDE.md.
Seems bad.
This way the primary agent only has relevant information in their context to make decisions and take actions.
Context management is still under valued imo.
I don't know, but I have a pretty standard (I think?) setup, and Claude manages to find every relevant file every time. But I've also only used Claude for greenfield projects, where "documentation is primary, and code flows from documentation" is the philosophy.
I have CLAUDE.md describe all the types of documentation files and the directory structure. And then Claude is pretty aggressive (automatically) about always inserting cross-references everywhere. So a feature description will reference the ADR's that it implements, the ADR's say what feature implements them. A code file will make reference to the "implementation design" document that describes the motivation behind which iOS elements were chosen, how the animation is defined in a particular way that doesn't break another animation, and so forth. So I've really never run into a situation where Claude failed to read a file it should have. I've been pleasantly surprised.
I would say that the one really big thing I've had to learn is to teach Claude both in CLAUDE.md and in the header of every top-level design document, that keeping documentation current and in sync is paramount. Because its default seems to be to keep history and append, e.g. by default it will take a section of a document and mark it "[DEPRECATED]" and add the new version below. So my instructions are pretty clear in having it always be aggressive in maintaining current state only, always replace rather than append. And if there's anything we want to save from the previous approach (e.g. we did it X way previously and it failed because Y), then just add that as a new short note in the new current-state text, possibly with a pointer to a commit or tag or something.
So this seems to solve both recall and staleness in my projects at least.
The only thing I still haven't found a solution for is numbering. Claude is always giving everything numbers, like F23 for feature 23. But I'm always changing the order of things, inserting new things, deleting things, so I wind up with a sequence of development work that goes in order like "Phase 9", "Phase 9b", "Phase 9e", "Phase 11", "Phase 12". I'm halfway ready to abandon numbers entirely and just start giving things names from noun collections instead, so every feature is named after an animal, every ADR is named after a kitchen implement, or something. Or just four-digit hex codes chosen at random. Curious if anyone else has found what works.
Not today. IMHO Documentation is just a form of structured memory and it’s all just context. Getting that context right is a hard problem and there’s a lot of different ways to skin that cat.
If I say "use jq instead of writing a python script to parse json" it should never write adhoc python scripts to parse json. Yet that constantly happens to me anyway.
It's not like it's being rewritten for efficiency. "Just because".
I’m the author of the project.
https://gist.github.com/cynthiateeters/6868ca26c059a3106cd93...
Otherwise you can get a python one liner that execs a different script engine.
The LLM's job would be to channel "I perambulate in the direction of the Arctic circle" into go(north). You saved writing the grammar parser, but you still need to write the game world.
You haven't been paying attention then. I routinely see Claude and GPT models churning out python code to do stupid things like linting. Last week I even had a TypeScript project with prettier configured all over the place, including in a custom agent skill I added with the express purpose of getting the damned model to lint the code, being constantly prompted to run ad-hoc python code supposedly to format whitespaces. I even explicitly prompted one session to just use npm run lint, where I pointed out the exact line of code where prettier was invoked, and the session still churned python code to hande whitespaces.
I think there is a deeper problem emerging from this sort of behavior. Even when we bother to create agent skills with there own scripts that call tools like jq a specific way to achieve a goal, AI coding assistants and agents still go way out of their way to generate ad-hoc scripts to do the most absurdly stupid tasks such as parsing output in structured language, and even remove whitespaces from a markdown file. This means AI coding assistants and coding agents treat agent skills as mere suggestions of using a alternative option that more often than not the choose to ignore.
This has a very dangerous implication: your average user is trained to develop a pavlovian reflex to authorize agents to just execute their ad-hoc scripting code with our own permissions and credentials in our systems, which includes the ability to call anything over the internet.
It also doesn't seem to have a way to separate preference from factual memory which is useful when you have various humans interacting with the same agent. One human might prefer a certain output style over the other and that's something a memory framework can also address.
Like the MCP articles a few months ago, it also seems to assume all agents are cli coding harnesses running on your local machine. We have a handful of other things like chat bots, event-driven agents running on servers, chat driven agents running on servers in sandboxes--there's not a single filesystem and even if there were one, having multiple agents try to edit it at once would corrupt it.
My main issue with LLMs is that by construction they work primarily by addition, and are very task oriented. If you have documentation, it adds blobs corresponding to its task and that's it, and soon enough you need to break your documentation onto chapters and you're back at square one.
I tried an approach based on the following idea recently and it's amazing - Lint rules where the error messages contain an explanation on how to deal with the issue.
https://habit-hooks.com/
I'm using it to foster IOSP (integration operation segregation principle) for example.
Or is that the secret sauce no one wants to share, the edge people see themselves having.
But some of the resellers have freebies, like:
https://app.primeintellect.ai/dashboard/environments?ex_sort...
https://github.com/sierra-research/tau2-bench
Then for example you could write your own hook and convert existing code smell documentation which agents ignore into a format that works with the hook.
That's what I did; I did not use the library I linked to ~ It served only as an inspiration.
Instead, I let the agents create custom lint rules (using eslint, pylint, ...) and add custom coaching error messages based on where I want to take my codebase.
https://youtu.be/6AgndHSkHFI?t=238
(You could of course argue that you don't like the direction the rules push the agents in given in the example - some people don't prefer small functions everywhere - but that's not the point: The lint-hooks work in pushing the agent in the desired direction; If one desires something else they'd simply need different rules)
Just flagging big functions will make the agent write small functions ~ but not necessarliy in a good way
(for example the agent might just cut `doOneThingAndTheOther` in half and call the second half `doOneThingAndTheOther2`)
.agents/plans/<plan-name>/
.agents/notes/<topic>/
.agents/knowledge/<topic>/
I only commit knowledge and if knowledge gets big i add .agents/knowledge/INDEX.md
This is a good mix of human readable and agent fluent. Notes are ephemeral, knowledge is permanent.
I have a rule for knowledge that it has to be stable and mostly permanent (though updatable). And the agents are not allowed to post there unless docs are clean organized and with permission.
Notes are for jotting things down and handoffs, massaging a featureset. Agent can document at will.
Still WIP.
My thought though has always been that I don't want there to be agent-only designated documentation.
I use mattpocock/skills and that generates ADRs (Architectural Decision Records). That only uses skills, including a setup skill that will write a few pointers in AGENTS.md. I always have a CONTRIBUTING.md to document development flow and a CODING_STANDARDS.md. Between those and the README.md and architecture documentation and commit messages the agents seem to be able to find and use docs and keep them up to date. We are also writing a lot of specs and putting those in Github issues.
I write README and docs by hand, thus only important stuff goes in there. If it is not worth my attention to write down, it's not worth writing down.
Smaller context, easier to consume.
Once qwen3.8-flash-next showed up, it cañ go "forever" with dynamic context pruning.
It's fascinating for local coding.
However, I found the same problems with agents that it has with humans (usually with humans it isn't TDD but code coverage requirements). The problem is the tests are now written to satisfy a bureaucracy rather than to properly verify the code. And there are real studies that point out the low value of these types of bureaucratic testing requirements. For example, when models are just required to write a test file, they don't produce a better result: https://arxiv.org/abs/2602.07900
I wrote a /verify skill that focuses on properly verifying code and I am finding it works a lot better. I do need to revise it now- in practice certain parts of the skill are doing all the heavy lifting and others are more dead weight. But it does seem to be properly orienting the agents towards finding defects. https://github.com/gregwebs/skills-sdlc/blob/main/skills/ver...
Context then is: docs, code and tests. Rather than try to build some omniscient agent, the context is built where the agent needs it.
TDD doesn't help the agent figure out what it is supposed to do now- it is already writing the code and now and it must know what it is supposed to do to write it.
I am trying to use evidence-based approaches. I am only observing better results and tweaking now, but I plan to benchmark when tweaking is done.
I'm working on javascript, and now I'm only doing this via typescript. I've successfully got this going:
1. Design a feature in plain language with the coding agen (opencode) and write a document for it.
2. Restart the context (or I use /compact to flush any errant details)
3. Pull the new plan into context and ask the agent to revise it follow Test Driven Development.
4. Depending on how big it is, the agent places it into multi stages, each with it's own document.
It then loops through the stages. It might be entirely based on the language you're using, but this loop seems strong enough.
Then when there's bugs, errors, anything, we update the doc, add more tests, then revise the code.
It's quite possible the harness you're using isn't setup properly. For me to do this locally, I had to hack on dynamic context pruning which I described here: https://news.ycombinator.com/item?id=49906637#49907641
The slightest amount of guidance, input from experience can make a huge difference.
And it's quite fascinating. Still dont think there's trillions of TAM out there if Qwen3.8-Flash-Next runs fast enough on $3k (before memory cartel) pricing.
And it’s an experimental likely undertrained model. Wait til Qwen 4 Flash…
"Code is the documentation" doesn't solve the problem in my experience. Because what happens is that you still need a lot of "why" comments in the code, and then these go stale, so you still have the same problem you have to solve. And so I find that markdown documentation is a lot easier to organize and review in a structured hierarchical way in one place, than code comments sprinkled across the repo.
https://github.com/Clumps4Linux/SEP/blob/main/SEP-1.1.1.txt
Essentially patterns the agents need to always think in. I also implemented a versioning system to the principles that need to be quoted in any comments which are there in the code. That way, when my principles evolve, so does the code.
I did package it up in a way that I can share it with friends [0]. System still evolving, but the last two-ish months that I've used it has served me really, really well, And it's been even better with the latest models.
I've had surprisingly good adherence from agents on this technique.
[0] https://principledriven.dev/
The github repo might be more helpful: https://github.com/Principle-Driven/pdd
I was wondering about this bit:
> "Code comments cite a versioned token where code depends on the rule."
What does a token look like? I went looking for an example, but I don't know what I'm looking for.
If it helps, here's how you think about it, principles are not decisions, they're mental models and frameworks to allow making decisions. In my case, they're there to allow agents to make decisions without asking me. But the decisions cannot be recorded - the code serves that purpose. What can be recorded is why the commend for the block of code needs to cite what principle was used in the current shape, which is why there is a citation of the principle.
When the principle evolves, so does every decision, at the very least, every decision gets re-litigated to see if it needs to evolve.
The questions are designed specifically to identify the state the agent is in, in order to assign its steering. Changing this tool changes how the agent works. You just need to make the use of this tool a necessity so it does not forget to call it.
Unlike principles and memories, questions are more openended and tend sometimes to trigger the right mentality in the agent even before it gets to the policy itself. In my opinion agents get lost in local work losing the big picture, my questions jog the big picture back into attention.
1. call ssp tool, no arg, it just reads features from the repo, like most other tools -> ssp locates state and sends 4-5 questions
2. agent responds the first batch -> state is further narrowed down -> ssp sends a second batch of questions
3. agent responds again to the interview -> state gets finally pinned down -> agent gets the steering assigned
4. a log is created of this ssp session, and reflection on the log used to refine the questions and steerings in the ssp tool (name comes from State-Space Policy)
Do you have any resources that guided your work?
A big part of it really comes to that I have led large product teams and platform engineering teams in the past, And I like giving feedback to my agents in the same way that I would plan things with my teams.
One of which is give them repeatable mental models which allow them to make decisions without relying on managers or me. The result is usually that the teams can work for weeks-to-months with autonomy.
I try achieving the same with my agents, so that they can work for 1-14 days without my intervention. I usually have them working on very very long running tasks/initiatives.
If you do want some resources, this is one of the pages that I've been using for nearly a decade to help people get started with mental models: https://fs.blog/mental-models/
It seems you tried to reinvent instruction files. What do you think is the difference between your approach and standard tools such as instruction files, AGENTS.md, and even skills?
The only "prompt" is in AGENTS.md saying that this thing exists and there's a map.md <- which is a one liner reference to whatever the agent stores in there.
And usually, I tackle a new feature, and at some point tell it to store to jot down notes in workbench if I'm comfortable with it (and if it needs to be stored in memory)
This makes it easy to just spawn other agents and such from a good point, I just point them that stuff is in workbench.
Token costs, seem decent and I can always just delete stuff in there as its purpose is ephemeral.
https://github.com/ArtRichards/docs-cli
and
https://artrichards.github.io/agent-playbook-suite/blog/
https://github.com/isaachinman/encephalon
But seriously, Opus has been garbage after 4.6 -- the public seems very easily fooled into thinking breadth == depth. Anthropic, I'll grant them, has been and is still an extraordinary data team. As for model alignment though... I have to wonder sometimes if they are even trying beyond just chat training.
"Guys, SI is just around the corner... Huh? What production DB? Anyway GPT2...I mean Mythos is too good it would be dangerous to release to the public!"
Genuine question: how is this different or better than existing context and info-about-the-project management systems like for example GSD (formerly get shit done) and others?
> No one rewatches a team meeting from 3 years ago to remember constraints around a feature. People write things down and use those records instead.
Memory plugins don't re-read old transcripts. Memory snippets are basically the notes that people write down after the meeting.
The problem, however, is different. The problem with memory plugins is that your agent basically writes a note every time someone says a sentence, and then tries to work with those 5000 notes.
Instead, the agent should recognize what's important and write only that. And that is, of course, a documentation. (And ADRs, if you want not only a description of the final state, but also the trajectory of how the agent arrived to it. Which, arguably, contains more information than the docs themsleves.)
Another difference is that memory snippets are immutable, append-only and don't have a lot of structure. Of course, this is done to be able to store lots and lots of notes: they should be independent. The main problem is that increasing the number of notes adds not enough benefits to compensate for downsides of this structure-less immutable format.
I've been bitten by agent-written ADRs. Agents carelessly add extrapolated details and speculated nice-to-haves that I never asked for, and this becomes a source of bloat that keeps coming back like a boomerang.
Lack of this capability makes automated principles or patterns update a recipe for more bloat.
however, I have found keeping a good solid reference to my home infrastructure, services, ci/cd setup, hosts, storage , networking etc.. really works wonders as a set of 'memories' to share across projects that I expect to be tested / deployed / acceptance tested etc.. using the home infra bits and pieces.
The code is the documentation. There's no need for most of this stuff.
Snarky comment aside, I am very interested in how we evaluate the performance of these systems and what kinds of work match best with different approaches.
https://gist.github.com/karpathy/442a6bf555914893e9891c11519...
`cd` to a folder. Launch `claude`. Do your work. Save scripts and documentation in that folder. `/resume` previous conversations from that folder.
That's it. That's the trick.
Now, having very static, very well-defined folders helps a lot. I'm Johnny.Decimal so I have numbered folders for everything I do. So my process when I want to use my 'process a travel booking from my email to my calendar' script is:
- `jd tripsy`
- `claude`- Say 'hey Claude, there's a new email in my inbox please'.
- Done.
Legally and culinarily, they're vegetables. Botanically they're fruit.
Tomatoes are members of the nightshare family which includes tobacco, potato, and chili peppers.
They are used in popular recipes like Mexican salsa and Italian pasta sauce. Pasta sauce commonly uses the San Marzano variety. Here's a <picture> of San Marzano tomatoes from our Italy trip.
Rhett likes tomatoes in all dishes. Link only likes tomatoes if they're blended in a dish like pasta sauce or tomato soup. Frank is allergic and can't have tomatoes.
Tomato prices are up due to a bad <year> yield in <country>.
---
This is all memory. Where does it go?
My most recent example were some deprecated and archived repos that kept getting added to plans for patching issues.
We use a private Claude plugin marketplace for internal plugins and skills and I try to regularly prune my memory, migrating relevant stuff to a proper home (skills in the marketplace, docs in repos or Notion etc) and prune outdated information.
Like, you can quickly lookup related ideas based on what came before and after, causes, effects, just like calling relationships a graph database.
Documents can’t be queried efficiently like that, you need a database.
I take the article's point more directly. It's just a straighter line to have clear communication through documentation than to fuddle around with the perfect memory setup.
https://backnotprop.com/blog/context-monorepos/
When I ask "what happened to x?" ... the agent has everything it needs to give me that answer. When it plans the next feature it can validate assumptions against previous decions made in my `decisions` folder.I don't think you need more documentation or these memory systems. The code IS the documentation. Lots of this stuff is unnecessary. It's a bunch of people doing their special rain dances and then when it happens to rain they say "I did that"
Organizing the markdown files by feature/issue/change makes it easier for the LLM to search for the appropriate documentation. Coupled with well-broken down Claude rules files, and Claude Code (or other harnesses) get better as you make more changes.
The idea of dumping everything into a big database and hoping the model will write the correct queries does work out to some extent. It's a very enchanting idea. However, it pales in comparison to having a dedicated tool per type. The outcomes seem to be much better when joins between low cardinality types occur within the token stream.
If your agent does need access to some enterprise knowledge base, I would give it a lexical search capability and not overthink it with vector shenanigans.
Tools are the only thing you need if you build them right. I don't even have a system prompt anymore aside from injecting the name of the robot and the current user's name. Keep in mind that all aspects of tools can be dynamic over time. I've got some where the description is composed by hundreds of lines of conditional string builder depending on the current state of the conversation.
I think what you are trying to say is that “give the LLMs access to information and it can figure out how to use it. Don’t get fancy about how to structure the data”.
If that’s the case it has largely turned out to be true. RAGs have fallen out of fashion.
Where does that put AGENTS.md though? Is it worth spending time to structure it or just dump it.
I've worked with docs that get _updated_, and progressive disclosure: make references to more obscure features in their own page, so they don't majorly bloat the context.
You’re confusing, I think, memory system with llm finetuning. Completely different concepts.
Github Copilot had the idea of attaching memory to files, and if the file hash changes the memory is automatically dropped (not sure if they still do it). This means they are overly eager to drop stuff (even if the file change is just cosmetic), but at least they don't accumulate outdated cruft too much. (a memory can still be outdated if it was invalidated by a change in another file though)
https://news.ycombinator.com/item?id=47300747
"We should revisit literate programming in the agent era" (silly.business) 292 points, 251 comments
https://cline.bot/blog/memory-bank-how-to-make-cline-an-ai-a...
Every single project at work that has documentation/ADRs/etc ends up stale as fuck and it pollutes context more than it helps. Even the agent generated docs. Hell, especially the agent generated docs. It's like the LLMs can't just DELETE something, they always amend. So now there's stale details in there about something we abandoned polluting context.
The code is the documentation, none of this stuff is needed. I'm seeing WAY too many devs/engineers go down the rabbit hole of trying to build these tools and systems but not actually getting meaningful work done.
I've been doing my best to remove and delete all the documentation. I usually keep a super high level AGENTS.md/CLAUDE.md file and that's it. The llm/agent can figure it out. You might argue I waste 5 minutes and a dollar or two for the agent to "figure it out" each time, but I'm pretty certain people are wasting FAR more time and money building these things out and maintaining them in each of their projects (or NOT maintaining them and they're actively BAD for their projects).
Since then I've gone down the rabbit hole of really digging deep into current AI research and especially what AI whistleblowers are currently saying. And I can't even express how existentially scared shitless I am.
That being said, the danger is measured in computer damage, which can be a lot personally and to a company, but less existential, so your mileage may vary as to how "scary" it is.
What scares me are the people who think that this software is the equivalent of an perfectly-smart elf in a box and use it blindly.
How, if at all, do you keep the architecture or a working theory of the code in your head?
Isn't memory just documentation that agents create and update on the fly to fill in the documentation hole that your average user leaves open?
There was a time when the need to provide context and signal was a key topic in LLM and AI-assisted coding. People talked about MCPs AGENTS.md and README.md and agent skills and even comments, descriptive tests, and naming conventions. Supposedly the theory was that if you provide context, agents wouldn't misbehave so much. Then plan mode and spec-kit approaches stepped in with approaches aimed at providing that context when starting sessions, because users never bothered with docs or MCPs or AGENTS.md or anything. But then plan mode and spec-kit approaches were considered too laborious and requiring too much from users. Then, because users systematically failed to provide context, this task was finally given to agents. And things just worked.
You should ask yourself why in software engineering circles documentation is seen more as a problem than a solution to any problem.
You have the whole session history right there. One recall skill and some JSON parsing gets you grep over perfect memory. Why would you ever use more tools to spend more tokens to construct a imperfect memory next to your session history?
I just don't get it.
But that's one session. Isn't memory for… the next session?
However, AI works differently from humans in that much more of its working context has to be made explicit. Because of that, there may be some fundamentally different way for AI to maintain or reconstruct a program’s overall model.
LLMs don't retrain, it's too expensive atm.
If I want the LLM to remember something I ask it to update some docs, and even check that docs are consistent across the board after doing so.
The only thing I want the LLM to remember everywhere is talk like a human (no load-bearing, not this/that etc...), so I have an AGENT.md for that.
Indeed, AI fanatics lack critical thinking. They easily accept the idea that AI needs documentation, but merely disagree the method to approach it.
https://xkcd.com/927/
If someone has good A/B evals of this being more effective I'll eat my hat, but the reason no one is publishing them is because well, evals are hard, and this is likely just magical thinking.
My current view is that an agent generally needs three things to work effectively: a way to discover information that isn’t obvious, such as a minimal `AGENTS.md` that points to more focused brief files; a clear way to verify that its work is correct; and some guidance on project-specific tastes. Everything else is noise.