AI agents · OpenClaw · self-hosting · automation

Quick Answer

Agent Memory vs Documentation: RAG or Markdown? (2026)

Published:

The short answer

For project knowledge — where a feature lives, why it was built, what was agreed — a documented Markdown workspace the agent reads before working and updates after beats a vector store of transcript snippets. Similarity search surfaces what is close, not what is correct or current, and it cannot be audited. Vector memory still earns its place for per-user facts that arrive from conversation rather than from work. Kevin Liao’s October 3, 2026 essay “Agents Don’t Need Memory. They Need Documentation.” put this argument on the Hacker News front page (57 points, 42 comments); this page lays out both sides. Facts verified October 4, 2026.

Two architectures

Vector-store memory (Mem0, Zep, Letta, most “memory” plugins)Document-based memory (Operator Memory, hand-rolled internal/)
Source of truthSnippets extracted from past transcriptsDocuments the agent writes and maintains
RetrievalTop-k by embedding similarity on every prompt, plus a search toolAgent reads an index, then the relevant files
OrganisationFlat (or tiered short/long-term) embeddingsTyped docs: instructions, specs, decisions, research, indexes
FreshnessBackground daemons: dedupe, merge, “dreaming” rewritesUpdated by the agent at the end of each task
AuditabilityThousands of rows in SQLite/pgvectorgit diff
SharingPer-installation storeCommit it; teammates and other agents read the same files
Token costInjection on every prompt + daemon runsReads only what the index points to
Best atFacts about people learned in conversationFacts about the project learned by working on it

The case against recall

Liao’s argument is that every memory plugin shares one pipeline — mine transcripts → generate snippets → embed → inject top-5 → expose a search tool — and that the fancier features (multi-tier memory, overnight “dreamers”, rerankers, continuous compression) are patches on the same flawed assumption: the agent forgets, so capture more and retrieve smarter. The failure modes follow directly:

  1. Similarity is not relevance. Five snippets about authentication from different months rank by closeness to the prompt, not by which one is still true.
  2. Snippets lose context. A 200-token memory drops the motivation, the constraints and the environment that made the decision sensible.
  3. The past is treated as truth in a codebase that changes every day.
  4. Agents cannot search for what they do not know. A search tool only helps if the agent already suspects something exists.
  5. The store is unauditable. Which of 10,000 embeddings are stale, never retrieved, or quietly wrong?

The alternative is how humans already work: nobody rewatches a meeting from three years ago; they read the doc. AGENTS.md files were the first step — they exist because people knew agents needed context — but one file is not a brain. Liao’s loop replaces prompt → build → forget with prompt → consult → build → update: before working, read the index and the relevant documents; after working, fix what is now stale and add what is missing, while the whole task is still in context.

The case for vector memory (the part the essay skips)

The essay is right about coding agents on a single project and wrong if generalised to everything:

  • Per-user memory at scale. A support or companion product with 100,000 users cannot maintain a curated document per user. Facts like “prefers terse answers” and “is on the Pro plan” arrive in conversation, not from work the agent did, and a memory store with extraction and dedupe is the only practical place for them. This is the workload Mem0, Zep and Letta are benchmarked on (compared here).
  • Unstructured history. When the question is “what did the customer say in March,” transcript search is the right tool, not a document someone had to anticipate.
  • The update step is the weak link. Document memory depends on the agent reliably revising docs after every task. Skip it a few times and the brain is as stale as any vector store — just more confidently so, because it reads as authoritative.
  • Documents rot differently, not never. A spec that no longer matches the code is a stale memory with better formatting. Liao’s loop mitigates this by making updates part of every task; it does not eliminate it.

The honest synthesis: documents for the project, a memory store for the people, and never inject either on every prompt without a reason.

Setting up document-based memory today

  1. Create the brain. internal/ (Liao’s original) or docs/agent/, with INDEX.md listing every document and one line on when to read it.
  2. Type the documents. Instructions (how review works here), specs (what was agreed with the user), decisions (what was chosen and why), research (notes on a foreign API), indexes (where things live). Short files beat long ones; the agent reads by file.
  3. Wire the loop into AGENTS.md / CLAUDE.md. “Before any task, read INDEX.md and the documents it points to for this area. After the task, update stale documents and add missing ones.” The after-step is the one teams forget.
  4. Commit it. The brain lives in the repo. Reviewers see memory changes in the same PR as the code they describe.
  5. Or install the plugin. Operator Memory (free, open source) packages exactly this: no vector database, no embeddings, no summarisers or daemons, plain Markdown you can read and share.

If you also need per-user memory, add Mem0 or Zep beside the documents rather than instead of them, and scope what each one is allowed to answer.

Related: agent memory vs chat memory vs RAG, best AI agent memory systems 2026, Anthropic Dreaming vs memory rot.

Last verified: October 4, 2026.

Sources