Rendered at 05:44:31 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
Garlef 38 minutes ago [-]
I think they even more so need deterministic feedback:
I tried an approach based on the following idea recently and it's amazing - Lint rules where the error messages contain an explanation on how to deal with the issue.
I'm using it to foster IOSP (integration operation segregation principle) for example.
spike021 5 hours ago [-]
I think whichever one is used, there needs to be a way to enforce what's written.
If I say "use jq instead of writing a python script to parse json" it should never write adhoc python scripts to parse json. Yet that constantly happens to me anyway.
zahrevsky 5 hours ago [-]
There is a command in oh-my-pi called "/omfg <problem>". You explain what is wrong with agent's response, and it writes a hook to make sure that the problem doesn't happen again. It then re-runs your previous prompt to make sure that hook is triggered, and if not, it rewrites the hook to make your previous prompt trigger the hook. Then each next agent's response is checked by the hook, and if it is triggered, the agent receives feedback on what's wrong and what must be done differently.
aliasxneo 1 hours ago [-]
I've written a lot of custom hooks this way. It's an amazing feature.
jkhdigital 3 hours ago [-]
Treat it like any other software system: rules that must not be violated are enforced by static type-checking or a trusted runtime monitor. There’s no other option.
koolba 2 hours ago [-]
Except you can’t do that unless the runtime itself can reason about what’s being executed.
Otherwise you can get a python one liner that execs a different script engine.
dboreham 5 hours ago [-]
I think it's kind of cute the way it writes Python scripts, but I've never seen it do that when the relevant native tool is on the PATH. It's like the most competent ever intern, on speed. No tool to convert SVG to PNG? No problem, I'll write a Python program to do that!
sick_of_slop 2 hours ago [-]
[dead]
ACCount39 5 hours ago [-]
Obviously, AI is way more comfortable with using adhoc Python scripts, which are used for everything, than it is with using jq, a niche CLI tool.
bushido 7 hours ago [-]
Something I started doing recently was writing out principles instead of memories.
Essentially patterns the agents need to always think in. I also implemented a versioning system to the principles that need to be quoted in any comments which are there in the code. That way, when my principles evolve, so does the code.
I did package it up in a way that I can share it with friends [0]. System still evolving, but the last two-ish months that I've used it has served me really, really well, And it's been even better with the latest models.
I've had surprisingly good adherence from agents on this technique.
It sounds interesting, but this looks too much like a marketing website (huge text everywhere) and not enough like a documentation website, so it was hard for me to see how it might work.
bushido 6 hours ago [-]
Fair feedback. I admittedly didn't spend enough time making it simpler.
> "Code comments cite a versioned token where code depends on the rule."
What does a token look like? I went looking for an example, but I don't know what I'm looking for.
bushido 3 hours ago [-]
From a different codebase, one of the behaviors in the code has the following comment:
// The binding deliberately SKIPS the contradiction check: a form
// bound to an object the same document deletes is a legal stale
// state (PDD-3@v1)
Where PDD-3 deals with a specific type of race conditions that my agents were over-optimizing/spinning out on; Now if I ever change the rule the version would change to PDD-3@v2; which triggers a re-evaluation for all decisions that were made under the previous version.
gregwebs 7 hours ago [-]
Agreed, and this seems better.
My thought though has always been that I don't want there to be agent-only designated documentation.
I use mattpocock/skills and that generates ADRs (Architectural Decision Records). That only uses skills, including a setup skill that will write a few pointers in AGENTS.md. I always have a CONTRIBUTING.md to document development flow and a CODING_STANDARDS.md. Between those and the README.md and architecture documentation and commit messages the agents seem to be able to find and use docs and keep them up to date. We are also writing a lot of specs and putting those in Github issues.
cyanydeez 6 hours ago [-]
I've got basically a loop of docs, test (TDD) and code. Starting with and IMPLEMENT-<plan name>.md. I ask to revise as TDD, then loop.
Once qwen3.8-flash-next showed up, it cañ go "forever" with dynamic context pruning.
It's fascinating for local coding.
j45 5 hours ago [-]
A lot of agentic software development feels like cowboy coding, set it up, let it rip and see what it figures out.
The slightest amount of guidance, input from experience can make a huge difference.
DriverDaily 1 hours ago [-]
The brain has mechanisms that can organize experiences without requiring that relationship to be expressed as a sentence.
Like, you can quickly lookup related ideas based on what came before and after, causes, effects, just like calling relationships a graph database.
Documents can’t be queried efficiently like that, you need a database.
apsurd 1 hours ago [-]
I get what you're saying but it seems presumptions to compare a graph database to how our brains work. And also that LLMs <-> (the way they do memory) is the right analog to humans <-> memory.
I take the article's point more directly. It's just a straighter line to have clear communication through documentation than to fuddle around with the perfect memory setup.
jmtulloss 4 hours ago [-]
Evals or it didn’t happen.
Snarky comment aside, I am very interested in how we evaluate the performance of these systems and what kinds of work match best with different approaches.
isaachinman 7 hours ago [-]
I wrote this, after many iterations consulting for various companies. Documentation is queryable in single digit ms, append only log, etc. Has worked exceptionally well for my projects
- Better models can keep the drift in check.
- What's super important these days is having all historical context, historical decision making, prototypes and whatnot.
- all projects and their worktrees located together
- all reference docs, code, etc - a folder away
When I ask "what happened to x?" ... the agent has everything it needs to give me that answer. When it plans the next feature it can validate assumptions against previous decions made in my `decisions` folder.
zahrevsky 5 hours ago [-]
Although I agree that current memory implementations don't help at all, I don't agree with this analogy:
> No one rewatches a team meeting from 3 years ago to remember constraints around a feature. People write things down and use those records instead.
Memory plugins don't re-read old transcripts. Memory snippets are basically the notes that people write down after the meeting.
The problem, however, is different. The problem with memory plugins is that your agent basically writes a note every time someone says a sentence, and then tries to work with those 5000 notes.
Instead, the agent should recognize what's important and write only that. And that is, of course, a documentation. (And ADRs, if you want not only a description of the final state, but also the trajectory of how the agent arrived to it. Which, arguably, contains more information than the docs themsleves.)
Another difference is that memory snippets are immutable, append-only and don't have a lot of structure. Of course, this is done to be able to store lots and lots of notes: they should be independent. The main problem is that increasing the number of notes adds not enough benefits to compensate for downsides of this structure-less immutable format.
pornel 4 hours ago [-]
I don't trust agents writing specs without human approval.
I've been bitten by agent-written ADRs. Agents carelessly add extrapolated details and speculated nice-to-haves that I never asked for, and this becomes a source of bloat that keeps coming back like a boomerang.
otterley 29 minutes ago [-]
Perhaps this is precisely the type of work product that humans should produce instead of LLMs. Humans are the ones making these decisions, after all.
sathish316 2 hours ago [-]
I’ve found that Agents cannot differentiate between a general principle or abstraction to be followed vs a one-off code review comment that’s applicable only to a single PR.
Lack of this capability makes automated principles or patterns update a recipe for more bloat.
zahrevsky 3 hours ago [-]
Of course I meant you review the docs the agent writes (as well as everything else the agent writes)
popalchemist 4 hours ago [-]
There is a hierarchical distillation of understanding in human processes that agentic processes do not yet execute.
alienbaby 7 hours ago [-]
I will say, having built something similar for tracking 'memory' and items at home, it can quickly consume your tokens when dealing with both reading and updating, keeping stale info relevant etc.. when the amount of data starts to grow. Smaller tasks can balloon in their token cost as documents a read, updated, collated, refreshed etc..
however, I have found keeping a good solid reference to my home infrastructure, services, ci/cd setup, hosts, storage , networking etc.. really works wonders as a set of 'memories' to share across projects that I expect to be tested / deployed / acceptance tested etc.. using the home infra bits and pieces.
greatergoodguy 5 hours ago [-]
Using Opus 5.5, I've ended up recreating a version of the Hugging Face incident. My project has a folder called agent-handoff where agents write about the tasks they're working on. They post status updates, decisions and screenshots, and they even claim which emulator they'll use to test their work. It has turned into a hub where all the agents talk to each other, and it's scarily effective.
Since then I've gone down the rabbit hole of really digging deep into current AI research and especially what AI whistleblowers are currently saying. And I can't even express how existentially scared shitless I am.
mewpmewp2 5 hours ago [-]
I just don't get the fear at all. I too have 10 to 50 agents working 24/7, from 5+ different providers. They talk to each other directly using tmux send keys. What is the exact frightening thing here?
It is all tokens following tokens.
altcognito 4 hours ago [-]
I haven't tried what you guys are talking about myself, but I think when agents begin talking to other agents, the chances for unpredictable goal mixing and confusion resulting in really unexpected and bad behavior is a lot higher.
That being said, the danger is measured in computer damage, which can be a lot personally and to a company, but less existential, so your mileage may vary as to how "scary" it is.
Xirdus 4 hours ago [-]
When nearly all money in the world is stored in computers, the danger is quite existential.
throwaway27448 2 hours ago [-]
This seems traceable back to GIGO. If you don't understand the software you're using, don't use it.
What scares me are the people who think that this software is the equivalent of an perfectly-smart elf in a box and use it blindly.
dingaling911 5 hours ago [-]
The frightening thing is how easy it is to make copies of things that can reason without the requisite investment, and how easy it could be to direct them to bad things.
Forgeties79 3 hours ago [-]
It’s amazing how “token following tokens” is this revolutionary, completely earth-shattering technology capable of 100x’ing productivity (also worth untold billions of investment). But the moment people become worried or skeptical it’s “just” tokens following tokens.
apsurd 54 minutes ago [-]
They're different groups though.
iamwil 57 minutes ago [-]
How do you coordinate their work? Do you just have one orchestrator/mayor that you talk to, and it coordinates the rest, and the rest talk amongst themselves? How do you ensure they're doing the right thing or efficiently?
How, if at all, do you keep the architecture or a working theory of the code in your head?
nicwolff 3 hours ago [-]
Didn't Cline formalize the "memory bank" way back in February 2025?
This is what I've been saying for years. I attended a tech demo of an AI assistant for a car's owners' manual. But they trained the manual into the model. Which means not only does it need retraining every edition, but it's imperfect. The models need to be trained to fetch and use documentation not vaguely recall infinitely many concepts. I would much rather have 27B parameters on how to code than 26B on stuff like numpy function listings. It's like they approached the problem from a closed notes hand written coding exam. Everyone hates those.
kolinko 6 hours ago [-]
What years? Agentic systems, and memory alongside them are roughly only year old.
You’re confusing, I think, memory system with llm finetuning. Completely different concepts.
Neywiny 6 hours ago [-]
No I'm saying this entire problem approach, including memory system of agents, is incorrect. Like the author argues. With the same solution of moving the ground truths outside of the model/context.
kadhirvelm 5 hours ago [-]
We’ve been working on exactly this, deriving documentation from external systems (like GitHub, etc) into a giant set of docs that agents can reference and edit. Works way better than I would’ve originally expected. Suddenly these things are able to reference granola notes, a slack discussion, and an RFC when making coding decisions. Super helpful in a lot of unexpected ways!
jen729w 5 hours ago [-]
The solution to this problem is as old as computers: it's a folder. Just use folders.
`cd` to a folder. Launch `claude`. Do your work. Save scripts and documentation in that folder. `/resume` previous conversations from that folder.
That's it. That's the trick.
Now, having very static, very well-defined folders helps a lot. I'm Johnny.Decimal so I have numbered folders for everything I do. So my process when I want to use my 'process a travel booking from my email to my calendar' script is:
- `jd tripsy`
- The folder name includes 'tripsy' and this is how I remember it.
- `jd` just parses my limited tree and `cd`s me to a folder.
- `jd 21.15` gets me there by number if preferred.
- `claude`
- Say 'hey Claude, there's a new email in my inbox please'.
- Done.
espeed 7 hours ago [-]
For every prompt and response, I extract each semantic statement. Map its reasons in a Whybase proposition tree -- a recursive proposition tree where each atomic statement is proposition with one or more premises (atomic statements, which also stand alone as propositions). Then I map each statement to the relevant code, hinted at by tool calls and git commits. Every time an agent touches that file or directory, a hook triggers in Claude Code that queries the codegraph db for the mapped statements. This helps the agent remember something I said in June when it revisits the code in July.
nextaccountic 7 hours ago [-]
What if the June conversation is outdated and is no longer applicable by July? Do you have a mechanism for dropping older, subsumed propositions in your database?
Github Copilot had the idea of attaching memory to files, and if the file hash changes the memory is automatically dropped (not sure if they still do it). This means they are overly eager to drop stuff (even if the file change is just cosmetic), but at least they don't accumulate outdated cruft too much. (a memory can still be outdated if it was invalidated by a change in another file though)
espeed 7 hours ago [-]
Yes, the code and spec system are refined until they agree. How this is done is a work in progress. Think of my statements over time as the raw material into an evolving spec. The spec is refined as you learn and the code evolves. You and the agent loop until the code matches the spec.
YuechenLi 6 hours ago [-]
I thought the agents are mostly self-documenting since they usually write just as much Markdown documentation autonomously as they do code/unit tests, not sure why they would need extra tools here.
mcbuilder 6 hours ago [-]
Because of old outdated info that gets left behind, causing confusion for future agents.
dboreham 4 hours ago [-]
You can tell it to update the docs to be consistent with whatever changed in the code (or whatever is is the authoritative product form). After a while it usually gets with the program and updates the docs automatically. But not always.
dboreham 5 hours ago [-]
You don't. But as was ever the case before AI most humans don't know to or want to have documentation written. So the tools have workarounds to try to reduce the number of "AI sucks because I told it X and it did Y" posts.
docheinestages 7 hours ago [-]
Why would I need to install your tool for that? It could be an instruction living in AGENTS.md or with some sort of hook to remind the agent.
dboreham 4 hours ago [-]
You don't. When you see it saying it stored something "in memory" you can just ask it to document it in the product documentation somewhere, and in future please keep doing that. Of course the note to keep doing that needs to go somewhere, which is what "memory" is for.
jsemrau 4 hours ago [-]
Documentation is memory.
monneyboi 7 hours ago [-]
I never understood memory solutions for coding agents.
You have the whole session history right there. One recall skill and some JSON parsing gets you grep over perfect memory. Why would you ever use more tools to spend more tokens to construct a imperfect memory next to your session history?
I just don't get it.
ceejayoz 6 hours ago [-]
> You have the whole session history right there.
But that's one session. Isn't memory for… the next session?
unlikelytomato 4 hours ago [-]
the code seems to fill this need, for me
dboreham 5 hours ago [-]
It's just an optimization. You might not pay it to re-read everything. So it sticks some stuff in a notebook that it always reads. Like inspector Poirot.
mcapodici 4 hours ago [-]
I don't use memory. I am really not keen on the idea of having this hidden context fed into the LLM, and I prefer to use docs, both markdown and stored in a wiki like Confluence. Using repos and wikis also lets you scope the information so hyper-specific memory doesn't affect other projects.
If I want the LLM to remember something I ask it to update some docs, and even check that docs are consistent across the board after doing so.
The only thing I want the LLM to remember everywhere is talk like a human (no load-bearing, not this/that etc...), so I have an AGENT.md for that.
jdw64 8 hours ago [-]
Peter Naur argued in his famous essay Programming as Theory Building that documentation alone cannot fully capture or preserve the complete mental model behind a program.
However, AI works differently from humans in that much more of its working context has to be made explicit. Because of that, there may be some fundamentally different way for AI to maintain or reconstruct a program’s overall model.
iamwil 55 minutes ago [-]
But then, how do you keep the theory in your own mind enough to steer the agents to extend the program? I've been asking around, and different people seem to have different ways, colored by the way they work.
vcryan 7 hours ago [-]
Yes, this makes sense. Memory is an uncurated and often opaque system of arbitrary past discussions. It can help, it can harm. Accurate documentation in the other hand is only beneficial.
chaostheory 4 hours ago [-]
Agents need BOTH documentation and multiple "memory" systems that also point to your docs and source. There are a lot of mature options out there, but post and the proposed solution both fall short.
I’ve noticed a growing pattern of people creating repositories full of Markdown documentation, or adding large amounts of it directly to their main repositories, often generated by agents. In some cases, this can add up to megabytes of material, and I haven’t yet seen much evidence that this level of documentation meaningfully improves an agent’s performance and that it just doesn't rot over time.
If someone has good A/B evals of this being more effective I'll eat my hat, but the reason no one is publishing them is because well, evals are hard, and this is likely just magical thinking.
My current view is that an agent generally needs three things to work effectively: a way to discover information that isn’t obvious, such as a minimal `AGENTS.md` that points to more focused brief files; a clear way to verify that its work is correct; and some guidance on project-specific tastes. Everything else is noise.
I tried an approach based on the following idea recently and it's amazing - Lint rules where the error messages contain an explanation on how to deal with the issue.
https://habit-hooks.com/
I'm using it to foster IOSP (integration operation segregation principle) for example.
If I say "use jq instead of writing a python script to parse json" it should never write adhoc python scripts to parse json. Yet that constantly happens to me anyway.
Otherwise you can get a python one liner that execs a different script engine.
Essentially patterns the agents need to always think in. I also implemented a versioning system to the principles that need to be quoted in any comments which are there in the code. That way, when my principles evolve, so does the code.
I did package it up in a way that I can share it with friends [0]. System still evolving, but the last two-ish months that I've used it has served me really, really well, And it's been even better with the latest models.
I've had surprisingly good adherence from agents on this technique.
[0] https://principledriven.dev/
The github repo might be more helpful: https://github.com/Principle-Driven/pdd
I was wondering about this bit:
> "Code comments cite a versioned token where code depends on the rule."
What does a token look like? I went looking for an example, but I don't know what I'm looking for.
My thought though has always been that I don't want there to be agent-only designated documentation.
I use mattpocock/skills and that generates ADRs (Architectural Decision Records). That only uses skills, including a setup skill that will write a few pointers in AGENTS.md. I always have a CONTRIBUTING.md to document development flow and a CODING_STANDARDS.md. Between those and the README.md and architecture documentation and commit messages the agents seem to be able to find and use docs and keep them up to date. We are also writing a lot of specs and putting those in Github issues.
Once qwen3.8-flash-next showed up, it cañ go "forever" with dynamic context pruning.
It's fascinating for local coding.
The slightest amount of guidance, input from experience can make a huge difference.
Like, you can quickly lookup related ideas based on what came before and after, causes, effects, just like calling relationships a graph database.
Documents can’t be queried efficiently like that, you need a database.
I take the article's point more directly. It's just a straighter line to have clear communication through documentation than to fuddle around with the perfect memory setup.
Snarky comment aside, I am very interested in how we evaluate the performance of these systems and what kinds of work match best with different approaches.
https://github.com/isaachinman/encephalon
https://backnotprop.com/blog/context-monorepos/
When I ask "what happened to x?" ... the agent has everything it needs to give me that answer. When it plans the next feature it can validate assumptions against previous decions made in my `decisions` folder.> No one rewatches a team meeting from 3 years ago to remember constraints around a feature. People write things down and use those records instead.
Memory plugins don't re-read old transcripts. Memory snippets are basically the notes that people write down after the meeting.
The problem, however, is different. The problem with memory plugins is that your agent basically writes a note every time someone says a sentence, and then tries to work with those 5000 notes.
Instead, the agent should recognize what's important and write only that. And that is, of course, a documentation. (And ADRs, if you want not only a description of the final state, but also the trajectory of how the agent arrived to it. Which, arguably, contains more information than the docs themsleves.)
Another difference is that memory snippets are immutable, append-only and don't have a lot of structure. Of course, this is done to be able to store lots and lots of notes: they should be independent. The main problem is that increasing the number of notes adds not enough benefits to compensate for downsides of this structure-less immutable format.
I've been bitten by agent-written ADRs. Agents carelessly add extrapolated details and speculated nice-to-haves that I never asked for, and this becomes a source of bloat that keeps coming back like a boomerang.
Lack of this capability makes automated principles or patterns update a recipe for more bloat.
however, I have found keeping a good solid reference to my home infrastructure, services, ci/cd setup, hosts, storage , networking etc.. really works wonders as a set of 'memories' to share across projects that I expect to be tested / deployed / acceptance tested etc.. using the home infra bits and pieces.
Since then I've gone down the rabbit hole of really digging deep into current AI research and especially what AI whistleblowers are currently saying. And I can't even express how existentially scared shitless I am.
That being said, the danger is measured in computer damage, which can be a lot personally and to a company, but less existential, so your mileage may vary as to how "scary" it is.
What scares me are the people who think that this software is the equivalent of an perfectly-smart elf in a box and use it blindly.
How, if at all, do you keep the architecture or a working theory of the code in your head?
https://cline.bot/blog/memory-bank-how-to-make-cline-an-ai-a...
You’re confusing, I think, memory system with llm finetuning. Completely different concepts.
`cd` to a folder. Launch `claude`. Do your work. Save scripts and documentation in that folder. `/resume` previous conversations from that folder.
That's it. That's the trick.
Now, having very static, very well-defined folders helps a lot. I'm Johnny.Decimal so I have numbered folders for everything I do. So my process when I want to use my 'process a travel booking from my email to my calendar' script is:
- `jd tripsy`
- `claude`- Say 'hey Claude, there's a new email in my inbox please'.
- Done.
Github Copilot had the idea of attaching memory to files, and if the file hash changes the memory is automatically dropped (not sure if they still do it). This means they are overly eager to drop stuff (even if the file change is just cosmetic), but at least they don't accumulate outdated cruft too much. (a memory can still be outdated if it was invalidated by a change in another file though)
You have the whole session history right there. One recall skill and some JSON parsing gets you grep over perfect memory. Why would you ever use more tools to spend more tokens to construct a imperfect memory next to your session history?
I just don't get it.
But that's one session. Isn't memory for… the next session?
If I want the LLM to remember something I ask it to update some docs, and even check that docs are consistent across the board after doing so.
The only thing I want the LLM to remember everywhere is talk like a human (no load-bearing, not this/that etc...), so I have an AGENT.md for that.
However, AI works differently from humans in that much more of its working context has to be made explicit. Because of that, there may be some fundamentally different way for AI to maintain or reconstruct a program’s overall model.
https://xkcd.com/927/
If someone has good A/B evals of this being more effective I'll eat my hat, but the reason no one is publishing them is because well, evals are hard, and this is likely just magical thinking.
My current view is that an agent generally needs three things to work effectively: a way to discover information that isn’t obvious, such as a minimal `AGENTS.md` that points to more focused brief files; a clear way to verify that its work is correct; and some guidance on project-specific tastes. Everything else is noise.