Notes
Notes
Notes
Longer pages by the people who build veneta: what we compared, what we tried ourselves, what we measured. Every number carries its conditions.
worldmodel-bench v0.2 · proposal — measuring a WorldModel on ten axes and four gates
A proposal: seven axes for measuring a world model, eight operating-condition gates, what is measured today and what is not; veneta-bench is the seventh axis. No composite score.
Separate the agent sessions, connect them with the memory layer — two weeks of notes
Ten agent sessions on one repository for two weeks: a worktree per session, shared folders vs copies, and the memory layer that keeps what a session learned. What worked and what bit us.
The same 31B, twice as fast — speculative decoding with Google's Gemma 4 drafter
The same 31B decodes about twice as fast with speculative decoding (Google's Gemma 4 drafter): 13.5–13.8 vs 6.3–6.4 tok/s, 12/12 on the memory abilities, turn median 23 s vs 50 s, one vLLM flag.
Choosing a local writer — eight models through the same loop
Eight writers on the same machine, loop and cases: memory abilities, operational cases, turn time, unsupported figures, leaked names. Who should pick what.
Two weeks on one DGX Spark — the power cut out, the clock got stuck, and we gave up on the 70B
Two power cuts under a 70B, the clock stuck at 507 MHz and the cold drain that clears it, the decode probe that tells, and the rules we adopted.
How reliably do local models call tools — 8 models, 9,759 turns
Tool calls per model, counted from 9,759 saved turns: own calls, tools that do not exist, nothing read, unsupported figures, leaked tool names.
RAG, LLM wiki, and veneta
What differs, and what happens when you use them together
We tried an LLM wiki ourselves
One scene that worked, one that did not, and the design that came out of it
