Skip to content

We tried an LLM wiki ourselves

Notes · Field notes

We tried an LLM wiki ourselves

One scene that worked, one that did not, and the design that came out of it

2026-10-03

Rather than compare LLM wikis in words only, we built the smallest possible one and fed it our own documents. The model reads sources, writes a Markdown page per topic, keeps an index and a change log, and when a question comes it picks pages from the index, reads them, and answers. This is the pattern, not any product; the code is a script of about 250 lines, ours. The model was a local gemma-4-31b-fp8; the sources were six of our documents (paper draft, experiment records), about 145,000 characters.

What it produced

Pages 41
Facts (one dated, sourced line each) 145
Model calls 68 (one per source chunk, one per page summary)
Time 78 minutes
Change-log lines 96
“Possible contradiction” flags 126

The 78 minutes of ingestion is the wiki’s first cost. Loading the same documents into veneta’s store uses embeddings only and takes under a minute. In return, the wiki’s pages are readable prose, and the organizing is done before any question arrives.

The scene that worked

Question: “What is the adoption rule for the judgment tier?”

The model picked two pages from the index and answered:

The CEO Adoption Rule requires adoption only when a system is clearly better on every set and risk cannot rise [CEO Adoption Rule, judge-tier-2026-10-01.md]. Adoption requires redoing all measurements, including the eight-model matrix, and testing after third-party memory benchmarks [CEO Adoption Rule, judge-tier-2026-10-01.md].

That is the rule we set on October 1, with the page name and the source file attached so it can be checked. The answer was possible because facts spread over several documents had already been gathered onto one page. That is the wiki’s strength, and veneta’s weak spot.

The scene that did not

In the same run the model created a page named constitution.md. In our paper, “constitution” is the list of articles a memory rule may not violate. The page’s summary reads:

The Constitution is the supreme law of the land. It defines the framework of the government and the rights of citizens.

No source contains that sentence. The model saw a word and wrote general knowledge as fact, and the page went into the index with nobody’s approval. The next question containing “constitution” will be answered from it. duplicate-matching.md is the same: our “merging near-duplicates” became a database-textbook definition. One page had a summary and no facts at all.

This is what “a fact the model wrote, with no approval gate” looks like in practice. A local 31B does it more often; a larger model would do it less. Less, not never.

Too many signals are no signal

We had the script flag a new fact as a “possible contradiction” when it shared words with an old fact but differed in numbers or dates. It flagged 126, and most were not contradictions. For example:

“The dense memory adversarial configuration passed 4.0/4 cases.” vs “The dense memory [layer A] records adversarial configuration passed 4.0/4 cases across 3 runs.”

Two configurations of the same model. With this many flags a person ignores all of them, and at that point review no longer exists. It is the same thing as clicking “approve all” in code review.

What we learned

veneta’s design for governance over a wiki came out of these three scenes and ships with v0.1. What the scenes say is plain.

  • Nobody approves 145 facts one by one. A person’s attention belongs somewhere else.
  • A new page with no basis in any source, like the constitution.md above, is something a person has to see.
  • Reading the 96-line change log is the review. There must be no interrupting prompts.
  • Count how many of the 126 flags were real; if most were not, the flagging rule has to change.

Next

This wiki’s file layout (pages/<name>.md with a summary, dated facts and sources, index.md, log.md) is the format veneta’s wiki store will work on. The store is measured first and its numbers go on the comparison page. The script ships with v0.1. Anyone who wants to repeat the experiment with another model is welcome.