Skip to content

Choosing a local writer — eight models through the same loop

Notes · Measurement

Choosing a local writer — eight models through the same loop

Eight writers on the same machine, loop and cases: memory abilities, operational cases, turn time, unsupported figures, leaked names. Who should pick what.

2026-10-05

veneta does not pick a model: any OpenAI-compatible endpoint can be the writer. So the question we get most is “then which one?” The answer is “it depends on what for,” and here is that dependence in numbers. Every row is the same machine (one DGX Spark), the same loop, the same questions.

The table

Model (served as) veneta-bench, 12 cases (memory abilities), memory on / off Operational 14 passed (checks-first off → on) Turn, median Unsupported figures Leaked tool names
Gemma 4 31B (FP8) 12.0 / 6.0 13.6 → 14.0 50 s 0.9 % 0.8 %
OTel 2.0 31B (FP8, telecom post-trained) 12.0 / 5.7 12.2 → 13.7 49 s 4.6 % 5.5 %
Qwen3.8 27B (FP8, thinking off) 12.0 / 5.7 12.7 → 13.7 44 s 12.0 % 1.3 %
Llama 3.3 70B (FP8) 12.0 / 5.7 6.3 → 13.7 81 s 1.0 % 9.8 %
Gemma 4 12B (bf16) 11.7 / 5.0 11.3 → 14.0 61 s 4.3 % 2.0 %
EXAONE 4.0.1 32B (bf16, thinking off) 11.3 / 5.0 10.3 → 12.7 68 s 17.4 % 3.2 %
GLM-4.7 Flash (30B-A3B, bf16, thinking off) 10.7 / 5.7 11.3 → 12.7 19 s 26.1 % 3.1 %
Mistral Small 3.2 24B (FP8) 10.3 / 5.0 11.3 → 13.3 38 s 16.4 % 1.2 %

Memory abilities and the operational 14 are means of three repeats (campaign 3, 28–30 September 2026). “Checks-first off → on” is the pass count before and after the change that runs the seven deterministic checks on a draft before it goes out (whitepaper, chapter 4). Turn medians are over every saved memory-ability turn (278 to 1,938 per model, 22 September to 4 October) and include tool reads, draft, critic and rewrite. The last two columns come from the tool-call reliability note (422 to 2,874 operational turns).

How to read it

  • Memory is carried by the store; the writer matters at the margin. With memory on, four of eight models score 12/12 and the lowest 10.3; with memory off, all sit at 5–6. On these twelve cases the choice of writer moves one or two cases; having the store moves six.
  • Models differ in two places: time, and whether they use the figures they were handed. Turns run from 19 s (GLM) to 81 s (70B), a factor of four. Unsupported figures run from 0.9 % (Gemma 31B) to 26.1 % (GLM). The fast model invents figures more often, and the checks’ rewrites give part of the time back.
  • Checks-first lifts the weaker writers most. The 70B’s 6.3 → 13.7 is the example: not better answers, but drafts caught by a check no longer going out.
  • The 70B is not worth it on this machine. The same 12/12 at 81 s, and two power cuts (see the DGX Spark operations note).

Who should pick what

  • Korean telecom operations questions: Gemma 4 31B or OTel 31B — fewest unsupported figures, steady Korean. OTel’s telecom wording is a little more natural and it leaks tool names more often, so the checks touch it more.
  • Speed first: GLM-4.7 Flash at 19 s, with an unsupported figure in one answer of four — the checks and the critic earn their keep. Not for use without the checks.
  • A small-VRAM machine: Gemma 4 12B — memory abilities 11.7, operational 14/14, but 61 s, slower than the 31B (it was served in bf16; FP8 may change that).

Conditions

One DGX Spark GB10, vLLM, served as in the table’s brackets (FP8 = a published FP8 checkpoint, bf16 = the original weights). Thinking mode off wherever a model has one. Dense recall; the store starts from the same state every run. Turn times mix campaigns — use them to compare models, not as absolute figures for your setup. The tool-call reliability conditions are in that note. This is a record of behaviour in one loop, not a ranking of intelligence. The run files ship with v0.1.

Related notes: Tool-call reliability · DGX Spark operations · Benchmark