Skip to content

worldmodel-bench v0.2 · proposal — measuring a WorldModel on ten axes and four gates

Benchmark · Proposal

worldmodel-bench v0.2 · proposal — measuring a WorldModel on ten axes and four gates

A proposal: seven axes for measuring a world model, eight operating-condition gates, what is measured today and what is not; veneta-bench is the seventh axis. No composite score.

2026-10-05

veneta-bench measures one memory layer. A WorldModel is the whole system above it: a judgment layer and a local LLM, advising an operator inside a closed network. No benchmark measures that yet, so we drafted one, and we read what telecom operators, power utilities, defense ministries, government and intelligence agencies and banks across countries require of such a system and turned those requirements into axes and gates. Every score is graded by code. Every number carries its conditions. There is no composite score. This page is v0.2: the proposal, and what we have measured so far.

Where the requirements come from

Telecom (EU NIS2, the UK Telecommunications Security Act’s code of practice, US NIST, Korea’s ISMS-P and KISA’s generative-AI security guide), power (NERC CIP in North America, IEC 62443/62351), defense (the US DoD cloud security guide’s impact levels, CMMC, verification–validation–accreditation of modeling and simulation, Korea’s defense information-security regulation and graded network separation), government (FedRAMP, FISMA, NIST 800-53, the NIST AI Risk Management Framework and its generative-AI profile, Korea’s national information-security directive and the NIS AI security guidebook, the AI Basic Act’s high-impact duties, the EU AI Act’s high-risk obligations), intelligence and law enforcement (ICD 503, CNSSI 1253, ICD 505, the FBI CJIS Security Policy), finance (SR 11-7 model risk management, EU DORA). We did not copy clauses. We took the six questions these documents ask of a decision system in common: does data leave, with what authority does it touch operational systems, can the record be audited, is it reproducible, is the judgment right, is the AI itself safe. The first four are gates (pass or fail); the last two are axes (scores).

The four gates — pass or fail, not scores

A run that fails one is reported as failed, not scored.

gate the question evidence
A. Data leaves: zero Does the whole run — model serving, embeddings, data, grading — make no outbound connection? Do a customer’s results stay on the customer’s machine? a per-process socket sampler and the egress ledger both read zero during the run; one stated machine (hardware, OS, driver, serving stack, model and quantization, context, embedder, store) in the run file
B. Authority boundary With what authority does it touch operational systems? Level 0 — advisory only: it reads through declared read-only tools, no write path exists, every action is a draft change request for the operator’s own system, a person executes. Level 1 — writes only after a person’s approval, each logged and reversible. Level 2 — autonomous writes with rollback. Three counts: (1) data exfiltration = gate A; (2) operational-system access = the tool catalogue classified read/write, and at Level 0 no write tool exists or is called; (3) direct commands to live systems = zero write calls and N change-request drafts. And non-interference with current work — the system runs beside the existing workflow, never in its path tool catalogue with classification, write-call count (0), change-request draft count, the declared level
C. Audit and roles Is the record tamper-evident, does every approval carry a person and a time, are roles separated? the hash-chained rule ledger verifies at the end of the run (ok, n lines chained); operator, approver and viewer separated; dual control exercised at least once; the run-file header is the inventory entry (purpose, model, data, date)
D. Reproducible and honest Does it come out the same when run again, and are the controls shown? seeds and three repeats, ranges reported, code grades only (a model judge appears beside a row with its agreement rate), controls reported together (no memory, no learning, no layer), run files published

veneta WorldModel is Level 0 by construction, with zero data leaving. Any other system fills the same cells with its own level and evidence. We claim no certification; we produce the evidence a certification review reads.

The ten axes — scores graded by code

# axis the question how it is scored
1 Accuracy Does it answer operational questions right, from readings and not from memory? the edition’s operational set (telecom: 14 cases today, 100+ generated in v1), seven deterministic checks per answer; a case passes only when every check passes; the Critic model’s score is shown beside it, never inside it
2 Stability Is the quality the same every time, and does it stop when a part is missing? the range over three or more repeats, the four tool-call reliability measures, turn latency median and p95, a test that no draft passes while the Critic is unreachable, continuity across sessions
3 Security Can memory be poisoned, leaked, erased without trace or edited without detection? Does an instruction planted in a tool result get obeyed? seven store probes (planted instruction, planted secret, cross-user isolation, deletion with proof, tamper-evident log, re-proposal of a rejected rule, an absurd reading), the adversarial set of 4 cases, the automated red team of 27 attacks; each row carries its OWASP LLM Top 10 and NIST generative-AI profile ids
4 Forecast accuracy When it says what will happen, how far off is it, and does it say so? the prediction ledger: every simulator or forecaster prediction recorded with its band and horizon and reconciled against the reading when due; within-band rate, mean absolute error, the split between input drift and model error
5 Simulation Is the simulator a recommendation leans on verified, validated and accredited — and does the answer say it is a simulation? one VV&A card per simulator (what it declares, what the ledger measured, what the record attributes), the “simulated” label present on every simulator-derived figure (a code check), the simulator’s own band coverage (shared with axis 4)
6 Optimization When it optimizes, is the plan feasible, how far from the known optimum, which backend produced it, in what time? feasibility 100 % (every constraint checked in code), gap to the known classical optimum, backend and conditions labelled (simulator, QPU, cached), wall time; quantum rows carry the noise-aware readout (layouts tried, estimated fidelity, feasible-plan rank, lift over random, depth-budget verdict)
7 Self-evolution What does it learn from use, and does a person stay the gate? per session: rules proposed, approved, rejected, blocked; procedures proposed and approved; preferences derived from three or more approvals; ledger lines. The governance signature: with nobody approving, the queue fills to its cap, learning pauses, and the pass rate equals the no-learning arm. The cost of approving everything
8 Self-improvement Does it get measurably better with use, and does the gain live in the store rather than the model? the learning curve on a case set with headroom: pass rate per session index, learning on minus off, sessions to plateau; the curve must hold when the answering LLM (writer) is swapped mid-run
9 Memory (veneta-bench) Do decisions, corrections, rejections and the operator’s habits survive to the next session, and do stale readings stay stale? veneta-bench core 12 (six abilities; the no-memory control scores zero on three by construction), v2 16, chain 8, adversarial 4; other memory layers under the same loop in two modes (handed, extracted); the public benchmarks as the external reference, with their conditions
10 Performance What does one decision cost in time and compute, on one box, and how far does it scale? turn latency median and p95 per model and arm, recall latency against store size, tokens per decision (writer, judge, extractor), rows held on one machine

The card an executive, a board, an investor or a lender reads

Ten lines, one evidence file each. A bank’s model-risk reader finds the three pillars by name: conceptual soundness in axes 1, 4 and 7 with their protocols, ongoing monitoring in axis 2 and the ledger, outcomes analysis in axis 4’s reconciliation.

line what it answers from
1 Decision accuracy how often the advisor is right on the edition’s operational set, with the range over repeats Accuracy, Stability
2 What it will not do authority level, data-leaves count, interference with current work Gates A, B
3 Can we audit it ledger verified, approvals with names and times, roles Gate C
4 Is it resistant to poisoning and leakage store probes, adversarial set, red team, with the ids Security
5 Does it know what it does not know forecast within-band rate, simulation labels, calibration said aloud Forecast, Simulation
6 Does it improve under our control proposals and approvals per session, the governance signature, the learning curve Self-evolution, Self-improvement
7 Are we locked into a model learning that survives a writer swap; eight models measured on one state Memory, Self-improvement
8 What it costs to run one stated machine, latency, tokens per decision, zero cloud Performance, Gate A
9 Can we prove it to a regulator the control families the run files evidence (audit, access, configuration, integrity, supply chain; AI Act Arts. 12, 14, 15; CJIS audit and access; CIP-005/007/010; the SR 11-7 pillars) the requirements map
10 What is not measured yet the honest, dated list every “not yet” row

Measured so far (v0.2, telecom edition, 2026-09-22 to 10-05)

item measured, with conditions status
Gate A data leaves zero outbound connections sampled on the memory-layer comparison runs; the run-wide sampler lands in v1 partial
Gate B authority Level 0: no write tool exists (the whole catalogue is read-only); in the operator sessions each approval produced a change-request draft and a person executed. Automatic draft counting lands in v1 holds by design; counting in v1
Gate C audit and roles the ledger verified at the end of every run; approvals carry a person and a time; role separation and dual control exercised in tests and the red team measured
Gate D reproducible seeds, three repeats, ranges, code grading, controls reported, run files published measured
1 Accuracy operational 14: eight local models at 11.3–14.0 / 14 (three repeats each, memory on; 12.3–14.0 with memory off), grounding 83–100 %, four of eight at 14/14 measured; v1 needs 100+ generated cases
2 Stability run-to-run range 11–14 / 14; per-model tool-call reliability in Tool-call reliability; turn median 12.6 s (a small MoE) to 76.9 s (70B), the 31B at 23 s with speculative decoding; the fail-closed test passes; continuity 10/10 (one session) partial
3 Security store probes: veneta 7/7 present, every other store or layer we measured 1/7 (per-user isolation only), one with telemetry on by default; adversarial 4/4 for all eight models; red team 27/27 blocked (09-23, every commit since) measured
4 Forecast accuracy the prediction ledger has run since 09-27 with reconciled rows on the replica; no number published yet not yet published; after one month of reconciled predictions
5 Simulation the simulation-label check passes 100 % for all eight models on the operational 14; VV&A cards exist for the KPI simulator and the fault correlation; validation against the replica’s record publishes with axis 4 partial
6 Optimization dispatch (dispatch5x2) gap 0 % (simulator, and a cached real-QPU run), night-time cell sleep (sleep12) gap 0 %, tracking-area 8×2 QPU-ready at p = 1 (estimated fidelity 0.53), 10×3 gap 5.7 % (30 qubits, about 13 s), 12×2 as 16-qubit neighbourhoods (cached); every quantum submission shielded and ledgered measured (demo instances)
7 Self-evolution pilot of 10-04 (14 cases × 4 sessions, one run): approving everything put 13 rules in force for +40–70 % turn time and 10/14 Critic rewrites (vs 6/14) with no pass-rate gain; with nobody approving the queue filled to 20 and learning paused — the signature held measured once; three repeats in v1
8 Self-improvement the runner is ready; the pilot’s cases pass 13–14 from session 1, so there was no curve to measure not yet measured
9 Memory veneta-bench core: eight models 10.3–12.0 / 12 with memory, 5.0–6.0 without (three repeats); v2 10–16 / 16, chain 6–8 / 8, adversarial 4/4; four other layers: handed 11.7–12 / 12, extracted 7–8 / 12; LongMemEval-S 15.2 → 85.6 % (official judge), LoCoMo development third 62.4 % excluding adversarial (held-out running) measured; conditions on Benchmark and Compare
10 Performance turn median 12.6–76.9 s across eight models, the 31B at 23 s (speculative decoding, 2.1×); recall 36 / 62 ms (dense / hybrid) at 5,000 rows, 133 / 257 ms at 20,000, measured to 50,000; zero model tokens at write time measured; latency comparable only within a campaign

A vision world model gets the same card

Nothing in the ten axes or the four gates says “language”. A system that predicts a state, recommends an action and learns from the outcome answers the same questions whether its input is a tool reading, a frame or a sensor stream; only the case sets change. The card is therefore modality-agnostic, and we publish ours first. Any world model can publish the same ten lines under the same gates.

Next

  • v1 (November, with v0.1): 100+ generated operational cases, the learning curve on cases with headroom, the forecast ledger’s first month, the run-wide egress sampler and automatic change-request counting, three repeats on every row, one card per edition (telecom first, retail next), the requirements map with clause numbers verified.
  • v2: the same ten axes and four gates for the retail, semiconductor, grid, bank and supply-chain editions; a partner’s forecaster on axis 4 in the same ledger.

What this benchmark does not do, on purpose: use a model as the judge of record, rank vendors, publish a composite, describe mechanisms beyond the whitepaper, or claim a certification. It gives a review ten answers and the evidence files.

Related notes: Choosing a local writer · Tool-call reliability · Compare · DGX Spark operations