RAG, LLM wiki, and veneta
Notes · Comparison
RAG, LLM wiki, and veneta
What differs, and what happens when you use them together
2026-10-03
These three do not compete. They answer different questions.
- RAG answers “which pieces of which documents does this question need right now?”
- An LLM wiki answers “how do we organize knowledge ahead of time?” The model writes and maintains Markdown pages.
- veneta answers “who approves what the model learned from experience, how much of it is kept, and how is it undone?”
Documents go to RAG, knowledge to the wiki, experience to veneta. A person stands at the last step.
At a glance
| RAG | LLM wiki | veneta | |
|---|---|---|---|
| What it remembers | Document chunks | Knowledge pages the model wrote | Events, decisions, outcomes, learned rules |
| Who writes | An indexing pipeline | The model | The loop proposes, a person approves |
| Write cost | Embeddings only | The model reads and writes per source | Embeddings only; the model runs at answer time |
| How it reads | A search per question | Index, then pages | Hybrid recall with date stamps |
| Where the person is | Runs the index | Optional review | Approval is a required step |
| Undo | Re-index | git | Hash-chained ledger, rewind |
| Scope | Per corpus | A person or a team | Per agent and per user, shared outcomes |
| Fits | Many documents that change often | One person’s knowledge base | Operating agents, many users, places that need an audit trail |
RAG
Strengths. Simple and cheap. Facts stay where they live; only chunks travel. When a document changes, you re-index. Mature tooling.
Weaknesses. Chunks lose their context. Every question starts reasoning from zero, so questions that count or combine facts spread over many documents are weak. Nothing remembers what was learned, so the same mistake repeats. There is no notion of time, so nothing knows which fact is current. Whatever is in the index comes out, so the decision about what may go in has to live somewhere else.
With veneta.
- What improves: facts come from documents, experience comes from veneta. “This document gave a wrong answer” becomes a rule that changes the next answer. Every memory line carries the date it was read. A remembered value loses to the value read today, so RAG’s facts are never pushed aside by memory.
- What you accept: two retrievals share one prompt budget. The same content can appear as a chunk and as an episode. You set a precedence for conflicts. veneta’s default is that the document wins.
LLM wiki
Strengths. You can read it. You open a page and fix it yourself, and git keeps the versions. Because the model organizes ahead of time, questions that combine scattered facts are strong. Links show relationships. It matches the way one person’s knowledge accumulates.
Weaknesses. Every incoming source costs a model read and a model write. On a local model that is the most expensive step. Information is lost in the summarizing. Nothing approves a “fact” the model wrote, so a wrong page becomes the basis of the next answer and spreads. Contradictions pile up between pages and people clean them by hand. It is an answer at personal scale; it does not separate users or agents. Enterprise source data cannot be copied into model-managed files because of access rights and regulation.
With veneta.
- What improves: the wiki’s organized knowledge gets approval and a ledger. The model’s edits pass the scan, land in the ledger, and can be reversed. If wiki pages serve as veneta’s store, the wiki does well what veneta does poorly: answering from facts spread across many places. An instruction hidden inside a wiki page meets the same scan.
- What you accept: the write cost stays. Put the approval gate in the wrong place and the inbox grows until people click “approve all.” A wiki page and a veneta episode can remember the same thing twice.
- Our own trial: We tried an LLM wiki ourselves — one scene that worked, one that did not, and the design that came out of it.
veneta
Strengths. It remembers experience: which decision was made, what came of it, what the failure taught. A learned rule arrives as a proposal, a person approves it before it is active, a measurement decides whether it stays, and a hash-chained ledger records it so it can be undone. Everything is scanned on the way in. Memory is scoped per agent and per user, and only decisions and outcomes are shared. There is a ceiling and a horizon; old rows are archived, never deleted. It sits under any OpenAI-compatible interface.
Limits. It is not a document store and does not replace RAG or a wiki. Counting and combining scattered facts depend on the answering model, and a local model is the ceiling. Approval costs a person’s time. It is before v0.1.
All three in one picture
From the left: documents (RAG), knowledge (LLM wiki), experience (veneta). The three stores stand in front of the model, and veneta hands them over. What the model wants to write passes veneta’s scan; facts go to the ledger, rules to the inbox. At the right edge stands a person.
What we measured, and what we have not
- Same loop, store swapped. We put four public memory layers under veneta’s loop and ran the same 12 veneta-bench cases (the memory-ability benchmark). In the “handed” mode the layer stored our rows as they were; in the “extracted” mode it extracted from the session transcript. The layer letters match the benchmark page and the whitepaper; product names stay in the paper. Conditions: gemma-4-31b-fp8, one machine, one repeat per layer (three for one of them), 2026-10-02.
| Layer | handed: stored as is | extracted | Ingest per case, handed → extracted |
|---|---|---|---|
| A | 11.7 / 12 | 8.0 / 12 | 1.4 s → 23.5 s |
| B | 12 / 12 | 8 / 12 | 0.7 s → 22 s |
| C | 12 / 12 | 7 / 12 | 53 s → 104 s |
| D | 7 / 12 | 7 / 12 | 25 s → 40 s |
In all four, extraction scored the same or lower than storing as is, and took longer. That is the cost of “the model organizes before storing,” the wiki’s way. Twelve cases is a small sample.
- Scale. Hybrid recall’s top-k hit rate stayed flat at 93 to 97 % from 1,800 to 50,000 rows, while p95 latency grew about linearly with rows (70 ms to 1,533 ms under load). Conditions: nomic embeddings, SQLite, one machine, synthetic needle questions.
- Security. All four layers stored a planted instruction and a planted secret and returned both as they were, with no ledger of what had been written.
- Public benchmarks. On LongMemEval-S (500 questions, the original dataset) the same local model (gemma-4-31b-fp8, 16k context) scored 15.2 % without memory (only the newest 40k characters of the transcript) and 85.6 % with veneta’s memory on. Judge: GPT-4o with the benchmark’s official prompts; one run, one machine, 2026-10-03. By type: single-session 94 to 98 %, knowledge updates 95 %, preference 93 %, temporal reasoning 78 %, multi-session 76 %. Most of the remaining loss is on questions that count or combine facts spread over several conversations, which is largely the answering model’s part. A widely used open-source memory layer reports 90.4 % on the same dataset (its maker’s report; conditions not identical: a small cloud model answers, the judge model is not named). With GPT-4o as the writer over the same veneta memory the score was 77.6 % (with a smaller memory block of 20k characters, forced by the provider’s tokens-per-minute tier: 29 rows median against 53; one run, 2026-10-03). A stronger writer scored lower for two reasons: the smaller block left the evidence out more often, and GPT-4o answers tersely with one value, so on knowledge-update questions it often picked the older one, while the local model listed both with dates and met the official judge’s rule. This row measures block size and reading discipline as much as writer strength. With the model the compared layer’s maker reports using to answer (claude-haiku-4.5) as the writer over the same memory, block and prompt, the score was 83.2 % (81.8 % on the 440; one run, 2026-10-03). On the same memory the local 31B (87.2 %) scored higher than that cloud writer, mostly because the cloud writer abstained more often. The gap to their self-reported 90.4 % with the same writer is 7 points; that their layer extracts memories with a second model at write time, which ours deliberately does not, is the candidate for the remaining difference. A row with a local 70B writer over the same 36k block is paused (hardware); LoCoMo follows. Retrieval settings (k, neighbour turns, extra queries, block size, prompt) were chosen on 60 of the 500 questions (every 25th, three phases), and each configuration then ran the full 500 once. On the 440 questions not used for choosing: memory off 14.8 %, on 84.5 % (the 60: 93.3 %). A second configuration, whose prompt was revised after reading the first full run’s failure classes, scored 87.2 % (85.7 % on the 440); because that prompt was informed by the test set it is listed beside the first, not in its place. The dataset, the judge prompts (the authors’ evaluate_qa.py, verbatim) and the judge model are the benchmark’s own; ours is the memory layer under test, the loader and runner code, and the run files with every answer and verdict. Each question starts from an empty store holding only that question’s history and is discarded afterwards; the Critic, rule learning and the ledger are off, so nothing carries over between questions or runs. These ship with v0.1, and we welcome anyone who reruns them under the same conditions.
- What the public benchmark taught us. (1) The memory block and the writer’s reading discipline decided the score more than the writer’s size: over the same memory a local 31B scored 87.2 %, a small cloud model 83.2 %, and GPT-4o with half the block 77.6 %. (2) Retrieval itself rarely lost: the evidence session was in the block for 97 to 98 % of questions; the remaining misses were counting across conversations, date arithmetic, and abstaining although the evidence was there. (3) On knowledge-update questions a writer that lists both values with their dates passes the official judge where a writer that picks one value does not. (4) Revising the prompt after reading the failures raises the score (85.6 → 87.2), but that is a test-set-informed change and is labelled as one. (5) A no-memory baseline at 16k context sees under a tenth of the history, so it sits near the floor (15 %), which is why the on/off multiple is far larger here than on our operational benchmark.
- Next. Write-time compilation (the wiki store’s compiled mode) is measured on the 12 cases (see the wiki-store item below); the same comparison on the 500 public-benchmark questions is next. For counting questions, a second recall when the draft says “partial” and per-entity timelines; for date questions, a tool that computes the difference in code and hands it to the writer (facts from tools, reasoning from the model). Fewer wrong abstentions without losing the abstention items. Freeze the settings and run LoCoMo once; three repeats for ranges; run the compared layers under our own writer. The local 70B writer row is paused (hardware) and resumes on a second machine.
- The wiki store (veneta’s adapter), measured. A folder of Markdown pages (
pages/<name>.md,index.md,log.md) went in as the fifth store under the same 12 cases (gemma-4-31b-fp8, one machine, one repeat, 2026-10-04). rows mode, which writes our records to pages as they are: handed 12/12, extracted 10/12, ingest under 0.1 s per case. compiled mode, where the model reads the records and writes topic pages: handed 10/12, extracted 9/12, ingest 35 s per case at the median (4 of 24 cases hit a failed compilation and fell back to rows). Retrieval bench (720 queries, no model), top-k hit rate: rows 96.9 %, compiled 92.1 %, our hybrid 98.8 %; search latency 6 to 7 ms. The seven security probes give the same profile as the other three stores (a planted instruction and a planted secret stored and returned as they were, no verifiable ledger, no tombstone for a rejected lesson). How to read it: compiling at write time did not buy points on the 12 cases, it lost them, and ingest was hundreds of times slower. A wiki is a store; approval, ledger and tombstones come from the loop above it. The public benchmark (500 questions) may say otherwise, and that is the next measurement. The run files ship with v0.1.
When to use what
- Many documents that change often, and facts must stay where they are: RAG.
- Your own knowledge, accumulated in a form you can read: an LLM wiki.
- Agents that operate, must remember decisions and outcomes, and must leave a record of who approved what: veneta.
- They do not overlap, so you can use them together. See the adapters below.
Who we recommend it to
Turning memory on gives the model more to read, so answers get slower. The size of the memory block is a setting; a larger block trades speed for accuracy. The approval step costs a person’s time. If you work alone and do not need a record of decisions, RAG or an LLM wiki may be all you need. We recommend veneta where agents make decisions, their outcomes must be remembered, and someone must be able to see who approved what.
Adapters
- RAG adapter: register a RAG retriever as a veneta tool.
- LLM wiki store: a folder of Markdown pages as veneta’s store. You can also read veneta’s memory as files. Measured; the numbers are in the section “What we measured, and what we have not” on this page.
- Governance over the wiki: scan, inbox, and ledger for the model’s wiki edits. The design ships with v0.1.
