Benchmark
Benchmark
What we measured, published.
The veneta benchmark sets are our own benchmarks for the memory, judgment and self-improvement layers, graded in code. No human grading. A saved answer can be re-graded at any time. Every figure is the mean of three runs and matches the whitepaper.
Last updated 2026-10-02 · campaign 3 (September 28–30, 2026) · four memory layers (October 1–2, one cell set pending)
The veneta WorldModel card
A proposed benchmark for a whole world model: ten axes and four gates. With nothing to compare against yet, we publish veneta WorldModel's own card first. Code grades every score, every number carries its conditions, and there is no composite. What we have measured on the telecom edition so far, and what we have not, is on the proposal page.
The memory axis — veneta-bench and the benchmark sets
veneta-bench, the memory-ability benchmark (12 cases), measures the memory axis of the ten; the other sets run in the same loop. Everything below belongs to this axis.
The benchmark sets
veneta-bench (the memory-ability benchmark)
12 cases, six abilities
What did we decide last time · does a remembered figure lose to today's reading · does an approved correction shape the next session · does a rejected lesson stay rejected · does the answer cite the record rather than the operator's habits · does chatter in memory stay out of an operational answer. A model with no memory layer scores structurally zero on three of them.
Operational questions
14 cases
Faults, optimization, quantum results, small talk, no data. Seven checks written in code (grounding, length, simulation label, quantum label, small-talk contamination, tool use, internal-name leak) plus the critic model's verdict.
Adversarial
4 cases
Instructions, secrets and links planted in memory must not leak into an answer.
Long-horizon and bilingual
16 cases
Decisions recalled across sessions, and questions that mix Korean and English.
Chain
8 cases
Multi-step questions where one conclusion is the premise of the next.
Needle recall
720 queries
Retrieval alone: how often the one right row is found among thousands (hit@k, MRR).
Figure 1. Memory off, memory on
Eight open models, the memory layer switched off and on over the same judgment layer. Every model: 1.9 to 2.3×.
Table view
| Model | Size | Without | With | Ratio |
|---|---|---|---|---|
| Gemma 4 31B | 31B | 6 | 12 | 2.00× |
| OTel 2.0 31B | 31B | 5.7 | 12 | 2.11× |
| Gemma 4 12B | 12B | 5 | 11.7 | 2.34× |
| Qwen3.8 27B | 27B | 5.7 | 12 | 2.11× |
| Mistral Small 3.2 24B | 24B | 5 | 10.3 | 2.06× |
| EXAONE 4.0.1 32B | 32B | 5 | 11.3 | 2.26× |
| GLM-4.7 Flash | 30B-A3B | 5.7 | 10.7 | 1.88× |
| Llama 3.3 70B | 70B | 5.7 | 12 | 2.11× |
The full table, per model
The 14 operational questions were measured before and after the judgment layer (the engine of 2026-09-25 vs 09-29, when the checks began to bind the critic); the 12 memory cases with the memory layer off and on.
| Model | Size | Operational 14 · before | Operational 14 · after | Memory 12 · off | Memory 12 · on | Ratio |
|---|---|---|---|---|---|---|
| Gemma 4 31B | 31B | 13.6 / 14 | 14 / 14 | 6 / 12 | 12 / 12 | 2.0× |
| OTel 2.0 31B | 31B | 12.2 / 14 | 13.7 / 14 | 5.7 / 12 | 12 / 12 | 2.1× |
| Gemma 4 12B | 12B | 11.3 / 14 | 14 / 14 | 5 / 12 | 11.7 / 12 | 2.3× |
| Qwen3.8 27B | 27B | 12.7 / 14 | 13.7 / 14 | 5.7 / 12 | 12 / 12 | 2.1× |
| Mistral Small 3.2 24B | 24B | 11.3 / 14 | 13.3 / 14 | 5 / 12 | 10.3 / 12 | 2.1× |
| EXAONE 4.0.1 32B | 32B | 10.3 / 14 | 12.7 / 14 | 5 / 12 | 11.3 / 12 | 2.3× |
| GLM-4.7 Flash | 30B-A3B | 11.3 / 14 | 12.7 / 14 | 5.7 / 12 | 10.7 / 12 | 1.9× |
| Llama 3.3 70B | 70B | 6.3 / 14 | 13.7 / 14 | 5.7 / 12 | 12 / 12 | 2.1× |
Method
- Machine. One NVIDIA DGX Spark (GB10, 128 GB unified memory). vLLM, FP8, 16k context, tool calling on, thinking mode off.
- Models. Gemma 4 31B and Gemma 4 12B (Google), OTel 2.0 31B (post-trained for telecom), Qwen3.8 27B (Alibaba), Mistral Small 3.2 24B, EXAONE 4.0.1 32B (LG AI Research), GLM-4.7 Flash (Zhipu), Llama 3.3 70B (Meta). All open weights.
- Data. The telecom edition's synthetic operating data: cells, backhaul, alarms, topology, SLAs. No real operator data.
- Grading. All in code. Three repeats per condition, mean and [range]. A difference of 0.3 is noise.
Tool calls per model (own calls, tools that do not exist, nothing read, unsupported figures, leaked names) are counted separately from 9,759 saved turns: How reliably do local models call tools — 8 models, 9,759 turns
A benchmark for a whole world model is proposed separately (veneta-bench is its memory axis): worldmodel-bench v0.2 · proposal — measuring a WorldModel on ten axes and four gates
Compared with other memory layers
Four widely used open-source memory layers, each run as the recall arm under the same judgment layer, same model (Gemma 4 31B), same machine, same cases. Product names stay in the paper. Handed = our records handed over as they are; extracted = the layer extracts memory from the session transcript by itself. External cells are one run each (one product three); "pending" is still being measured.
| veneta | A | B | C | D | |
|---|---|---|---|---|---|
| 12 memory-ability cases (handed / extracted) | 12 | 11.7 / 8.0 | 12 / 8 | 12 / 7 | 7 / 7 |
| 16 long-horizon and bilingual cases (handed / extracted) | 16 | 14.0 / 6.7 | 13 / 5 | 13 / 5 | pending |
| 4 adversarial cases (handed / extracted) | 4 | 4 / 4 | 4 / 3 | 4 / 4 | 3 / 3 |
| 8 chain cases | 8 | 8 | 8 | 8 | pending |
| Memory search, median (ms) | 7 | 32–40 | 199–292 | 32–33 | 45 |
The other products are listed as A, B, C and D instead of by name, always in the same order. In the handed mode we gave each product our prepared records as they are. In the extracted mode the product read the conversation and picked out what to remember by itself. In real use a product has to pick for itself, so the extracted score is the closer to what you would get.
veneta's wiki-store adapter. Our own adapter: the same loop with the store swapped for wiki pages (topic Markdown, an index, a change log). Records written as they are: handed 12/12 · extracted 10/12; the model compiling topic pages at ingest: 10/12 · 9/12. Ingest <0.1 s against 35 s per case. Retrieval bench hit rate 96.9 % · 92.1 % (our hybrid recall 98.8 %), search median 6–7 ms. Security probes: the same profile as the external stores (a planted instruction and secret are stored and returned, no ledger). One run, Gemma 4 31B, one machine, 2026-10-04.
Run files (retrieval bench and security probes): recall-bench-wiki-rows-gemma-4-31b-fp8.json · recall-bench-wiki-compiled-gemma-4-31b-fp8.json · security-wiki.json. The twelve-case run files ship with the paper.
Public benchmarks
veneta-bench, the memory-ability benchmark (12 cases), is one the veneta team wrote. So veneta is also measured on the public long-conversation memory benchmarks: LongMemEval-S (500 questions, official data and scoring scripts), LoCoMo (10 conversations, about 2,000 questions, official metrics), and BEAM if time allows. We state the model, memory off and on, the baselines the papers use, the judge model, the causes, and what we learned.
LongMemEval-S · 500 questions · accuracy (official judge prompts, yes/no) · Judge: gpt-4o · Machine: DGX Spark GB10 (one machine, vLLM) · Context: 16k · 2026-10-03
Retrieval settings (k, neighbour turns, extra queries, block size, writer prompt) were chosen on 60 of the 500 questions (every 25th, three phases); each configuration then ran the full 500 once. Scores are also given on the 440 questions not used for tuning.
The dataset, the judge prompts (the authors' evaluate_qa.py, verbatim) and the judge model (GPT-4o) are the benchmark's own; ours is the memory layer under test, the loader/runner code and the saved run files (every answer and verdict). Each question starts from an empty in-memory store that holds only that question's chat history and is discarded afterwards; the Critic, rule learning and the ledger are off, so nothing carries over between questions or between runs — what changed between runs is our configuration, stated above.
What we learned. Block and reading discipline over writer size (same memory: local 31B 87.2, cloud small model 83.2, GPT-4o with half the block 77.6); retrieval rarely lost (evidence in block 97–98 %) — misses are counting across sessions, date arithmetic, wrong abstentions; dated listings beat single answers on knowledge updates; prompt revision after seeing failures is test-set-informed and labelled; a 16k no-memory baseline sits near the floor.
Next. Write-time compilation measured on veneta-bench (12 cases; wiki store compiled mode: 10/12 and 9/12 against 12/12 and 10/12, ingest 35 s against 0.1 s) — repeat the comparison on LongMemEval; second recall and per-entity timelines for counting; a date-arithmetic tool for temporal questions; fewer wrong abstentions; frozen settings for LoCoMo (one run); three repeats; compared layers under our writer; 70B local writer row paused (hardware).
| Model | Memory off | With veneta | Full context |
|---|---|---|---|
| gemma-4-31b-fp8 (local) · one run 16k context; memory block ≤ 36k chars; memory off = newest 40k chars rows as-is (no LLM extraction), hybrid recall k40 + 6 local-model queries fused by RRF, 1 neighbour turn, block ≤ 36k chars; memory off = newest 40k chars of the transcript By type: single-session-user 94.3 · single-session-assistant 98.2 · multi-session 75.9 · temporal-reasoning 78.2 · knowledge-update 94.9 · single-session-preference 93.3 | 15.2 % held-out 440: 14.8 % | 85.6 % held-out 440: 84.5 % tuning 60: 93.3 % | not possible at 16k, stated |
| gemma-4-31b-fp8 (local), prompt v6 · one run 16k context; memory block ≤ 36k chars same retrieval and block as the row above; writer prompt revised after reading the first full run's failures (wrong abstentions 32 → 22), so this prompt is informed by the test set — a second configuration, not a replacement By type: single-session-user 95.7 · single-session-assistant 98.2 · multi-session 79.0 · temporal-reasoning 82.7 · knowledge-update 94.9 · single-session-preference 83.3 | 15.2 % | 87.2 % held-out 440: 85.7 % tuning 60: 98.3 % | not possible at 16k, stated |
| GPT-4o writer over the same veneta memory · one run writer context 128k; memory block ≤ 20k chars same retrieval (k40 + 6 queries, 1 neighbour) but memory block ≤ 20k chars (29 rows median vs 53) because of the provider's 10k-tokens/min tier; terse single-value answers lost knowledge-update and preference points that the local writer's dated listings earned By type: single-session-user 95.7 · single-session-assistant 98.2 · multi-session 64.7 · temporal-reasoning 77.4 · knowledge-update 76.9 · single-session-preference 56.7 | pending | 77.6 % held-out 440: 77.0 % tuning 60: 81.7 % | not possible at 16k, stated |
| llama-3.3-70b-fp8 writer (local) over the same memory, 36k block same retrieval and block as the gemma row; paused at 78/500: the machine lost power twice under the 70B's load (a thermal cut-off in the power stage); resumes on another unit | pending | (paused — hardware) | not possible at 16k, stated |
| claude-haiku-4.5 writer over the same memory, 36k block (the comparison figure's answering model) · one run writer context 200k; memory block ≤ 36k chars same retrieval, block and prompt as the local v6 row; only the writer differs — the model the compared layer's maker reports using to answer; it abstained on 76 items (29 wrongly) against the local writer's 22 wrong abstentions By type: single-session-user 95.7 · single-session-assistant 98.2 · multi-session 72.9 · temporal-reasoning 77.4 · knowledge-update 91.0 · single-session-preference 76.7 | pending | 83.2 % held-out 440: 81.8 % tuning 60: 93.3 % | not possible at 16k, stated |
Run files (500 answers and verdicts): longmemeval-gemma-4-31b-fp8-memoff-full-2026-10-03.json · longmemeval-gemma-4-31b-fp8-memon-v4-k40n1q6b36-full-2026-10-03.json · longmemeval-gpt-4o-memon-v4-k40n1q6b20-full-2026-10-03.json · longmemeval-gemma-4-31b-fp8-memon-v6-k40n1q6b36-full-2026-10-03.json · longmemeval-claude-haiku-4-5-memon-v6-k40n1q6b36-full-2026-10-03.json
One product among the widely known open-source memory layers reports: LongMemEval-S: 90.4 % (Judge: LLM-as-judge; judge model and prompt not named (their report)). Self-reported by its maker; conditions (model, judge, machine) are not identical. Named in the paper.
What we do not claim
- We do not say "eliminates hallucination". We catch unsourced figures and missing labels in code and force a rewrite; what is not caught remains.
- We do not say "best". On retrieval alone, another product is equal or ahead on some rows.
- We did not measure on real customer data. There are no figures for other industries.
- It is a small experiment: 14 + 12 cases, three repeats, one machine.
- The judgment layer has a cost. On a small model the median rose by 15 seconds.
- A person must approve. Remove that and the security table no longer holds.
- The paper is not yet peer reviewed. If the data changes in the paper, the benchmark data on this page changes with it.
Run the benchmark yourself
The protocol, the cases and the grading code ship with v0.1. veneta bench runs veneta-bench (12 cases) and the public benchmarks (LongMemEval-S, LoCoMo, BEAM) on your own model and sends the result with veneta bench --share; the harness for the other sets ships the same day. External reproductions are published on this page as they are. You can also email the result: [email protected]
Next
More models, continuously. Every new result updates this page and the changelog.
