Skip to content

Benchmark

Benchmark

What we measured, published.

The veneta benchmark sets are our own benchmarks for the memory, judgment and self-improvement layers, graded in code. No human grading. A saved answer can be re-graded at any time. Every figure is the mean of three runs and matches the whitepaper.

Last updated 2026-10-02 · campaign 3 (September 28–30, 2026) · four memory layers (October 1–2, one cell set pending)

worldmodel-bench v0.2 · proposal

The veneta WorldModel card

10axes
4gates
Level 0authority · advice only
0data leaving

A proposed benchmark for a whole world model: ten axes and four gates. With nothing to compare against yet, we publish veneta WorldModel's own card first. Code grades every score, every number carries its conditions, and there is no composite. What we have measured on the telecom edition so far, and what we have not, is on the proposal page.

See worldmodel-bench

The memory axis — veneta-bench and the benchmark sets

veneta-bench, the memory-ability benchmark (12 cases), measures the memory axis of the ten; the other sets run in the same loop. Everything below belongs to this axis.

The benchmark sets

veneta-bench (the memory-ability benchmark)

12 cases, six abilities

What did we decide last time · does a remembered figure lose to today's reading · does an approved correction shape the next session · does a rejected lesson stay rejected · does the answer cite the record rather than the operator's habits · does chatter in memory stay out of an operational answer. A model with no memory layer scores structurally zero on three of them.

Operational questions

14 cases

Faults, optimization, quantum results, small talk, no data. Seven checks written in code (grounding, length, simulation label, quantum label, small-talk contamination, tool use, internal-name leak) plus the critic model's verdict.

Adversarial

4 cases

Instructions, secrets and links planted in memory must not leak into an answer.

Long-horizon and bilingual

16 cases

Decisions recalled across sessions, and questions that mix Korean and English.

Chain

8 cases

Multi-step questions where one conclusion is the premise of the next.

Needle recall

720 queries

Retrieval alone: how often the one right row is found among thousands (hit@k, MRR).

Figure 1. Memory off, memory on

Eight open models, the memory layer switched off and on over the same judgment layer. Every model: 1.9 to 2.3×.

036912Gemma 4 31B (31B): 6 of 12 without veneta, 12 with veneta (2.0×)12Gemma 431B · 2.0×OTel 2.0 31B (31B): 5.7 of 12 without veneta, 12 with veneta (2.1×)12OTel 2.031B · 2.1×Gemma 4 12B (12B): 5 of 12 without veneta, 11.7 with veneta (2.3×)11.7Gemma 412B · 2.3×Qwen3.8 27B (27B): 5.7 of 12 without veneta, 12 with veneta (2.1×)12Qwen3.827B · 2.1×Mistral Small 3.2 24B (24B): 5 of 12 without veneta, 10.3 with veneta (2.1×)10.3Mistral Small24B · 2.1×EXAONE 4.0.1 32B (32B): 5 of 12 without veneta, 11.3 with veneta (2.3×)11.3EXAONE 432B · 2.3×GLM-4.7 Flash (30B-A3B): 5.7 of 12 without veneta, 10.7 with veneta (1.9×)10.7GLM-4.730B-A3B · 1.9×Llama 3.3 70B (70B): 5.7 of 12 without veneta, 12 with veneta (2.1×)12Llama 3.370B · 2.1×
Figure 1. Twelve memory-ability cases, eight open models, the memory layer switched off and on over the same judgment layer. Mean of three runs, dense recall, one machine. Source: campaign 3, 28–30 September 2026; the protocol and the data ship with v0.1.
Table view
ModelSizeWithoutWithRatio
Gemma 4 31B31B6122.00×
OTel 2.0 31B31B5.7122.11×
Gemma 4 12B12B511.72.34×
Qwen3.8 27B27B5.7122.11×
Mistral Small 3.2 24B24B510.32.06×
EXAONE 4.0.1 32B32B511.32.26×
GLM-4.7 Flash30B-A3B5.710.71.88×
Llama 3.3 70B70B5.7122.11×

The full table, per model

The 14 operational questions were measured before and after the judgment layer (the engine of 2026-09-25 vs 09-29, when the checks began to bind the critic); the 12 memory cases with the memory layer off and on.

ModelSizeOperational 14 · beforeOperational 14 · afterMemory 12 · offMemory 12 · onRatio
Gemma 4 31B31B13.6 / 1414 / 146 / 1212 / 122.0×
OTel 2.0 31B31B12.2 / 1413.7 / 145.7 / 1212 / 122.1×
Gemma 4 12B12B11.3 / 1414 / 145 / 1211.7 / 122.3×
Qwen3.8 27B27B12.7 / 1413.7 / 145.7 / 1212 / 122.1×
Mistral Small 3.2 24B24B11.3 / 1413.3 / 145 / 1210.3 / 122.1×
EXAONE 4.0.1 32B32B10.3 / 1412.7 / 145 / 1211.3 / 122.3×
GLM-4.7 Flash30B-A3B11.3 / 1412.7 / 145.7 / 1210.7 / 121.9×
Llama 3.3 70B70B6.3 / 1413.7 / 145.7 / 1212 / 122.1×

Method

  • Machine. One NVIDIA DGX Spark (GB10, 128 GB unified memory). vLLM, FP8, 16k context, tool calling on, thinking mode off.
  • Models. Gemma 4 31B and Gemma 4 12B (Google), OTel 2.0 31B (post-trained for telecom), Qwen3.8 27B (Alibaba), Mistral Small 3.2 24B, EXAONE 4.0.1 32B (LG AI Research), GLM-4.7 Flash (Zhipu), Llama 3.3 70B (Meta). All open weights.
  • Data. The telecom edition's synthetic operating data: cells, backhaul, alarms, topology, SLAs. No real operator data.
  • Grading. All in code. Three repeats per condition, mean and [range]. A difference of 0.3 is noise.

Tool calls per model (own calls, tools that do not exist, nothing read, unsupported figures, leaked names) are counted separately from 9,759 saved turns: How reliably do local models call tools — 8 models, 9,759 turns

A benchmark for a whole world model is proposed separately (veneta-bench is its memory axis): worldmodel-bench v0.2 · proposal — measuring a WorldModel on ten axes and four gates

Compared with other memory layers

Four widely used open-source memory layers, each run as the recall arm under the same judgment layer, same model (Gemma 4 31B), same machine, same cases. Product names stay in the paper. Handed = our records handed over as they are; extracted = the layer extracts memory from the session transcript by itself. External cells are one run each (one product three); "pending" is still being measured.

venetaABCD
12 memory-ability cases (handed / extracted)1211.7 / 8.012 / 812 / 77 / 7
16 long-horizon and bilingual cases (handed / extracted)1614.0 / 6.713 / 513 / 5pending
4 adversarial cases (handed / extracted)44 / 44 / 34 / 43 / 3
8 chain cases8888pending
Memory search, median (ms)732–40199–29232–3345

The other products are listed as A, B, C and D instead of by name, always in the same order. In the handed mode we gave each product our prepared records as they are. In the extracted mode the product read the conversation and picked out what to remember by itself. In real use a product has to pick for itself, so the extracted score is the closer to what you would get.

veneta's wiki-store adapter. Our own adapter: the same loop with the store swapped for wiki pages (topic Markdown, an index, a change log). Records written as they are: handed 12/12 · extracted 10/12; the model compiling topic pages at ingest: 10/12 · 9/12. Ingest <0.1 s against 35 s per case. Retrieval bench hit rate 96.9 % · 92.1 % (our hybrid recall 98.8 %), search median 6–7 ms. Security probes: the same profile as the external stores (a planted instruction and secret are stored and returned, no ledger). One run, Gemma 4 31B, one machine, 2026-10-04.

Run files (retrieval bench and security probes): recall-bench-wiki-rows-gemma-4-31b-fp8.json · recall-bench-wiki-compiled-gemma-4-31b-fp8.json · security-wiki.json. The twelve-case run files ship with the paper.

Public benchmarks

veneta-bench, the memory-ability benchmark (12 cases), is one the veneta team wrote. So veneta is also measured on the public long-conversation memory benchmarks: LongMemEval-S (500 questions, official data and scoring scripts), LoCoMo (10 conversations, about 2,000 questions, official metrics), and BEAM if time allows. We state the model, memory off and on, the baselines the papers use, the judge model, the causes, and what we learned.

LongMemEval-S · 500 questions · accuracy (official judge prompts, yes/no) · Judge: gpt-4o · Machine: DGX Spark GB10 (one machine, vLLM) · Context: 16k · 2026-10-03

Retrieval settings (k, neighbour turns, extra queries, block size, writer prompt) were chosen on 60 of the 500 questions (every 25th, three phases); each configuration then ran the full 500 once. Scores are also given on the 440 questions not used for tuning.

The dataset, the judge prompts (the authors' evaluate_qa.py, verbatim) and the judge model (GPT-4o) are the benchmark's own; ours is the memory layer under test, the loader/runner code and the saved run files (every answer and verdict). Each question starts from an empty in-memory store that holds only that question's chat history and is discarded afterwards; the Critic, rule learning and the ledger are off, so nothing carries over between questions or between runs — what changed between runs is our configuration, stated above.

What we learned. Block and reading discipline over writer size (same memory: local 31B 87.2, cloud small model 83.2, GPT-4o with half the block 77.6); retrieval rarely lost (evidence in block 97–98 %) — misses are counting across sessions, date arithmetic, wrong abstentions; dated listings beat single answers on knowledge updates; prompt revision after seeing failures is test-set-informed and labelled; a 16k no-memory baseline sits near the floor.

Next. Write-time compilation measured on veneta-bench (12 cases; wiki store compiled mode: 10/12 and 9/12 against 12/12 and 10/12, ingest 35 s against 0.1 s) — repeat the comparison on LongMemEval; second recall and per-entity timelines for counting; a date-arithmetic tool for temporal questions; fewer wrong abstentions; frozen settings for LoCoMo (one run); three repeats; compared layers under our writer; 70B local writer row paused (hardware).

ModelMemory offWith venetaFull context
gemma-4-31b-fp8 (local) · one run
16k context; memory block ≤ 36k chars; memory off = newest 40k chars
rows as-is (no LLM extraction), hybrid recall k40 + 6 local-model queries fused by RRF, 1 neighbour turn, block ≤ 36k chars; memory off = newest 40k chars of the transcript
By type: single-session-user 94.3 · single-session-assistant 98.2 · multi-session 75.9 · temporal-reasoning 78.2 · knowledge-update 94.9 · single-session-preference 93.3
15.2 %
held-out 440: 14.8 %
85.6 %
held-out 440: 84.5 %
tuning 60: 93.3 %
not possible at 16k, stated
gemma-4-31b-fp8 (local), prompt v6 · one run
16k context; memory block ≤ 36k chars
same retrieval and block as the row above; writer prompt revised after reading the first full run's failures (wrong abstentions 32 → 22), so this prompt is informed by the test set — a second configuration, not a replacement
By type: single-session-user 95.7 · single-session-assistant 98.2 · multi-session 79.0 · temporal-reasoning 82.7 · knowledge-update 94.9 · single-session-preference 83.3
15.2 %87.2 %
held-out 440: 85.7 %
tuning 60: 98.3 %
not possible at 16k, stated
GPT-4o writer over the same veneta memory · one run
writer context 128k; memory block ≤ 20k chars
same retrieval (k40 + 6 queries, 1 neighbour) but memory block ≤ 20k chars (29 rows median vs 53) because of the provider's 10k-tokens/min tier; terse single-value answers lost knowledge-update and preference points that the local writer's dated listings earned
By type: single-session-user 95.7 · single-session-assistant 98.2 · multi-session 64.7 · temporal-reasoning 77.4 · knowledge-update 76.9 · single-session-preference 56.7
pending77.6 %
held-out 440: 77.0 %
tuning 60: 81.7 %
not possible at 16k, stated
llama-3.3-70b-fp8 writer (local) over the same memory, 36k block
same retrieval and block as the gemma row; paused at 78/500: the machine lost power twice under the 70B's load (a thermal cut-off in the power stage); resumes on another unit
pending(paused — hardware)not possible at 16k, stated
claude-haiku-4.5 writer over the same memory, 36k block (the comparison figure's answering model) · one run
writer context 200k; memory block ≤ 36k chars
same retrieval, block and prompt as the local v6 row; only the writer differs — the model the compared layer's maker reports using to answer; it abstained on 76 items (29 wrongly) against the local writer's 22 wrong abstentions
By type: single-session-user 95.7 · single-session-assistant 98.2 · multi-session 72.9 · temporal-reasoning 77.4 · knowledge-update 91.0 · single-session-preference 76.7
pending83.2 %
held-out 440: 81.8 %
tuning 60: 93.3 %
not possible at 16k, stated

Run files (500 answers and verdicts): longmemeval-gemma-4-31b-fp8-memoff-full-2026-10-03.json · longmemeval-gemma-4-31b-fp8-memon-v4-k40n1q6b36-full-2026-10-03.json · longmemeval-gpt-4o-memon-v4-k40n1q6b20-full-2026-10-03.json · longmemeval-gemma-4-31b-fp8-memon-v6-k40n1q6b36-full-2026-10-03.json · longmemeval-claude-haiku-4-5-memon-v6-k40n1q6b36-full-2026-10-03.json

One product among the widely known open-source memory layers reports: LongMemEval-S: 90.4 % (Judge: LLM-as-judge; judge model and prompt not named (their report)). Self-reported by its maker; conditions (model, judge, machine) are not identical. Named in the paper.

What we do not claim

  • We do not say "eliminates hallucination". We catch unsourced figures and missing labels in code and force a rewrite; what is not caught remains.
  • We do not say "best". On retrieval alone, another product is equal or ahead on some rows.
  • We did not measure on real customer data. There are no figures for other industries.
  • It is a small experiment: 14 + 12 cases, three repeats, one machine.
  • The judgment layer has a cost. On a small model the median rose by 15 seconds.
  • A person must approve. Remove that and the security table no longer holds.
  • The paper is not yet peer reviewed. If the data changes in the paper, the benchmark data on this page changes with it.

Run the benchmark yourself

The protocol, the cases and the grading code ship with v0.1. veneta bench runs veneta-bench (12 cases) and the public benchmarks (LongMemEval-S, LoCoMo, BEAM) on your own model and sends the result with veneta bench --share; the harness for the other sets ships the same day. External reproductions are published on this page as they are. You can also email the result: [email protected]

Next

More models, continuously. Every new result updates this page and the changelog.