Skip to content

How reliably do local models call tools — 8 models, 9,759 turns

Notes · Measurement

How reliably do local models call tools — 8 models, 9,759 turns

Tool calls per model, counted from 9,759 saved turns: own calls, tools that do not exist, nothing read, unsupported figures, leaked tool names.

2026-10-04

A post in a local-LLM community (2026-06) put it well: when you attach a local model to an agent, the real problem is not tokens per second but tool calling, and what is missing is an independent measurement that counts the failure modes separately — the call that never comes, the call to a tool that does not exist, the call that breaks and leaks into the text, the result the model then ignores. Our saved run files hold every call the model made, its result and the final answer for every turn of the last two weeks, so this was counting, not measuring.

What we counted

Operational turns only (fault, optimisation, quantum); small talk and no-data questions call no tool and are left out. Per turn:

  • The model’s own calls — beyond the readings the harness read first and handed over (see conditions). Of those,
    • unknown tool — a name that is not in the catalogue (recallFromMemory, none, …). The loop returns an error result and the answer goes on.
    • state error — a real tool refused for the state it found (e.g. “no open incident”). Mostly not the model’s fault; informational.
  • Nothing read — the turn answered without any tool, harness included.
  • Unsupported figure — a number in the answer that no tool result or memory line carries. This is the “got the result and did not use it, or misread it” axis.
  • Leaked tool name — a tool’s name or call syntax left in the prose. A call that became text.
  • Request error — the request to the model server itself failed.

Results

Model (served as) Turns Own-call turns Own calls Unknown tool State err Nothing read Unsupported figure Leaked Req err
gemma-4-31b (FP8) 2,874 8.7 % 252 0 0 0.0 % 0.9 % 0.8 % 0
otel-31b (FP8) 1,442 0.9 % 17 0 1 0.5 % 4.6 % 5.5 % 1
gemma-4-12B (bf16) 746 5.8 % 51 0 0 0.0 % 4.3 % 2.0 % 0
qwen3.8-27b (FP8) 710 0.0 % 0 — 0 0.0 % 12.0 % 1.3 % 0
exaone-4.0.1-32b (bf16) 632 13.4 % 107 0 2 0.0 % 17.4 % 3.2 % 0
llama-3.3-70b (FP8) 512 95.3 % 656 22 (3.4 %) 45 0.0 % 1.0 % 9.8 % 21
glm-4.7-flash (bf16) 422 20.6 % 120 0 3 0.0 % 26.1 % 3.1 % 0
mistral-small-3.2-24b (FP8) 422 2.1 % 12 0 0 0.0 % 16.4 % 1.2 % 0

How to read it

  • “Nothing read” is near zero for every model because of the harness, not the model. veneta reads the tools first, as the first step of every operational question, and hands the readings to the model. These are rates inside that loop, not the failure rate of a bare model in an agent OS.
  • The own-call rate is a habit, not a score. Some models stop when the readings suffice (gemma 8.7 %, qwen 0 %); one calls more on nearly every turn (llama 95 %). More calls did not make better answers.
  • Only one model called tools that do not exist. llama-3.3-70b did so on 22 of its 656 calls (3.4 %): none, recall, recallLastAction, recallFromMemory. Most are attempts to call memory as a tool; memory is already in the prompt.
  • The axis that separates models is not calling but using the result. Unsupported figures run from 0.9 % (gemma-4-31b) to 26.1 % (glm-4.7-flash). Given the same tool results, models invent or misread numbers at very different rates — which is why the answer is bound to checks.
  • Leaked names appear when the call format fights the model’s “voice”: llama 9.8 % and otel 5.5 % are high, the rest 1–3 %. vLLM’s per-model parser is the first guard; our check catches the remainder and rewrites it.

Conditions

  • One machine (DGX Spark GB10), vLLM, served per model: gemma-4-31b FP8-Dynamic, parser gemma4 / otel-31b FP8, gemma4 / gemma-4-12B bf16, gemma4 / qwen3.8-27b FP8, hermes, thinking off / exaone-4.0.1-32b bf16, hermes, thinking off / llama-3.3-70b FP8-dynamic, llama3_json / glm-4.7-flash bf16, glm47, thinking off / mistral-small-3.2-24b FP8, mistral.
  • Turns come from three campaigns between 2026-09-22 and 10-04 (the operational 14, the 12 veneta-bench cases (the memory-ability benchmark) and their extension sets, the layer-comparison runs) in a mix that differs per model (the run file lists turns per set). The loop changed in between (checks bind the Critic since 09-27), so these are the rates of the final answers our loop returned in that period, re-graded with today’s checks.
  • One or two model calls per turn (a draft, a rewrite when needed), not a deep multi-step agent loop — so the “format breaks as the turn deepens” effect is not in this table.
  • Not a ranking of model intelligence. A record of how each handled tools inside the same loop.

Next

A console panel with these four axes live (this session’s tool-accept rate), and the same count for the public-benchmark writer rows, cloud models included. The counting script (npm run eval:tools) and the run file ship with v0.1, so anyone who runs the same loop on their own model can draw the same table.