My 17 Agents Remember With 1,069 Markdown Skills, Not Vectors
Kevin Liao published "Agents Don't Need Memory. They Need Documentation." on October 3rd. It reached the Hacker News front page at 159 points and 90 comments. That is small by HN standards, and unusually dense with people who actually run agents.
His claim, in his own words: a memory plugin "generates 1,000 isolated snippets and inserts them into a vector database. With your every prompt, it attaches the five most similar snippets." He calls the result "a lottery over RAG snippets, injected on every prompt."
I run 17 agent profiles on this machine, so I went and counted what their durable knowledge actually is. Liao is right about the mechanism. He is thin on the prescription, and that gap is the part worth arguing about.
What the post actually argues
The post names five problems with recall-by-similarity, and each one is worth reading in his phrasing:
- Memories are "surfaced by similarity", so you don't know "which is correct, current, or what's missing."
- They are "stored without context."
- "The past is treated as truth." He asks how accurate each of the 500 snippets about authentication actually is.
- "Agents can't search for what they don't know."
- "The store is unauditable": there are "10,000 embeddings in SQLite. Which memories exist? Which are stale?"
His alternative is a Markdown "brain" and a different loop. Instead of "prompt → build → forget", you run "prompt → consult → build → update."
That distinction holds up when you look at a fleet already built this way.
So I counted mine
I keep 17 named agent profiles running here. There is a researcher, a CTO, backend and frontend engineers, a QA pair, and so on.
Every profile carries a SOUL.md and a skills/ corpus. Most also keep a memories/ directory, though the fleet is not uniform about it:
- SOUL.md: the standing instructions. What this agent is, and what it may not do. All 17 profiles have one.
- skills/: a corpus of SKILL.md files, one per procedure the agent has learned to reuse. Also all 17.
- memories/: MEMORY.md plus USER.md, small and capped on purpose. This part is common but not universal: 12 of 17 profiles have a MEMORY.md, 8 have a USER.md, and 7 have both.
Across all 17 profiles that is 1,069 SKILL.md files, between 58 and 72 per profile. viral-researcher, my most prolific, carries 60.
That is where the durable knowledge lives. Not in an embedding table, and not in a snippet injected on every prompt.
The memory files are small on purpose
The 1,069 files are the visible half of the receipt. The other half is the number nobody quotes:
- viral-researcher, MEMORY.md: 594 bytes
- cto, MEMORY.md: 8.8 KB, the largest per-profile memory file in the fleet
- root MEMORY.md: 9,089 bytes
- root USER.md: 2,924 bytes
A memory file that can grow without limit is a bug, not a feature. "Write everything down" produces exactly the failure the thread is already arguing about, so the character budget is load-bearing on purpose.
The thread's objections are the good part
The best arguments against the post are sitting in its own comments, and I would not skip them.
The first camp says agent-written documentation is just a new kind of bloat:
- stbenjam: "a growing pattern of people creating repositories full of Markdown documentation... often generated by agents. In some cases, this can add up to megabytes... I haven't yet seen much evidence that this level of documentation meaningfully improves an agent's performance and that it just doesn't rot over time"
- pornel: "I don't trust agents writing specs without human approval. I've been bitten by agent-written ADRs... source of bloat that keeps coming back like a boomerang"
- mzhaase built the same thing, "Decisions", until "the agent... would say 'violates D-236'. Some rule it made up that I never approved"
- alienbaby: "it can quickly consume your tokens when dealing with both reading and updating, keeping stale info relevant."
The second camp says the whole category is overbuilt:
- monneyboi: "You have the whole session history right there. One recall skill and some JSON parsing gets you grep over perfect memory"
- ceejayoz replies: "But that's one session. Isn't memory for… the next session?"
- jen729w: "The solution... is a folder"
Both camps are circling the same point. Files are the right substrate, and an uncurated file corpus fails the same way an uncurated vector store does. mzhaase's D-236 is not a Markdown problem. It is an unapproved-rule problem, and it survives whatever storage format you pick.
The mechanism is not new, and the thread says so
Liao should get credit for the framing, not the invention, and the comments hand him the lineage. gregwebs points out that mattpocock/skills "generates ADRs (Architectural Decision Records)" plus "a setup skill that will write a few pointers in AGENTS.md". nicwolff asks: "Didn't Cline formalize the 'memory bank' way back in February 2025?" He did. Cline's own memory-bank guide is a 200-point post.
The post builds on the AGENTS.md convention, which is public and widely used.
The deeper reference in the thread is jdw64 citing Peter Naur's "Programming as Theory Building". That is the strongest version of the argument. Documentation is not a log of decisions; it is how the theory of a system outlives the people who built it. Which is also why it only works when somebody curates it.
So don't delete your vector store
The post is right about the mechanism and incomplete as a prescription. "More Markdown" is not the fix. A system with three parts is:
- Instructions: SOUL.md or AGENTS.md. What this agent is, and what it must never do.
- Procedural skills: the 1,069 versioned SKILL.md files. How to do one job, written once, reused.
- A capped, human-reviewed memory file: 594 bytes to 8.8 KB of the facts that have to outlive a session.
Hermes ships the last two as separate subsystems. Its docs name the split precisely. There is a "Memory System: Persistent memory that grows across sessions" and a "Skills System: Procedural memory the agent creates and reuses." At 17 agents that separation stops being a design nicety. It is the reason the memory files stay small enough to read.
So here is the test for the next memory plugin you are sold. Ask what happens at 10,000 entries. If the answer is similarity search, you have bought retrieval theatre.
Move the durable material into files you can read, diff, and revert. Then cap the memory file, so that a human has to decide what stays.