Skip to content

Notes

Notes

Notes

Longer pages by the people who build veneta: what we compared, what we tried ourselves, what we measured. Every number carries its conditions.

2026-10-05

worldmodel-bench v0.2 · proposal — measuring a WorldModel on ten axes and four gates

A proposal: seven axes for measuring a world model, eight operating-condition gates, what is measured today and what is not; veneta-bench is the seventh axis. No composite score.

2026-10-05

Separate the agent sessions, connect them with the memory layer — two weeks of notes

Ten agent sessions on one repository for two weeks: a worktree per session, shared folders vs copies, and the memory layer that keeps what a session learned. What worked and what bit us.

2026-10-05

The same 31B, twice as fast — speculative decoding with Google's Gemma 4 drafter

The same 31B decodes about twice as fast with speculative decoding (Google's Gemma 4 drafter): 13.5–13.8 vs 6.3–6.4 tok/s, 12/12 on the memory abilities, turn median 23 s vs 50 s, one vLLM flag.

2026-10-05

Choosing a local writer — eight models through the same loop

Eight writers on the same machine, loop and cases: memory abilities, operational cases, turn time, unsupported figures, leaked names. Who should pick what.

2026-10-05

Two weeks on one DGX Spark — the power cut out, the clock got stuck, and we gave up on the 70B

Two power cuts under a 70B, the clock stuck at 507 MHz and the cold drain that clears it, the decode probe that tells, and the rules we adopted.

2026-10-04

How reliably do local models call tools — 8 models, 9,759 turns

Tool calls per model, counted from 9,759 saved turns: own calls, tools that do not exist, nothing read, unsupported figures, leaked tool names.

2026-10-03

RAG, LLM wiki, and veneta

What differs, and what happens when you use them together

2026-10-03

We tried an LLM wiki ourselves

One scene that worked, one that did not, and the design that came out of it