Skip to content

The same 31B, twice as fast — speculative decoding with Google's Gemma 4 drafter

Notes · Measurement

The same 31B, twice as fast — speculative decoding with Google's Gemma 4 drafter

The same 31B decodes about twice as fast with speculative decoding (Google's Gemma 4 drafter): 13.5–13.8 vs 6.3–6.4 tok/s, 12/12 on the memory abilities, turn median 23 s vs 50 s, one vLLM flag.

2026-10-05

On one DGX Spark, Gemma 4 31B (FP8) under vLLM decoded at 6.4 tok/s. With the model and its quantization unchanged, adding Google’s published Gemma 4 drafter (google/gemma-4-31B-it-assistant, Apache 2.0, 939 MB) through vLLM’s speculative decoding took it to 13.7 tok/s. Answer quality was checked on the same benchmark.

The numbers

No drafter (night of 4 Oct, idle GPU) Drafter on (morning of 5 Oct, three probes)
Decode probe (about 190 tokens) 6.3 / 6.4 tok/s 13.5 / 13.7 / 13.8 tok/s
SM clock · power during decode 2528 MHz · 36 W 2528 MHz · 36–37 W
veneta-bench, 12 cases (memory abilities, hybrid recall, one run) 12 / 12 · turn median 49.5 s 12 / 12 · turn median 23.1 s

Same clock, same power: decode 2.1× faster, and a whole turn — tool reads and the critic included — also 2.1× faster. All twelve verdicts were identical (the answer text is sampled, so it is not identical to the letter). Speculative decoding has the base model verify the drafter’s tokens, so the output distribution is the base model’s; there is no reason for quality to drop, and the twelve cases confirm once that it did not.

How to turn it on (vLLM 0.27.1, container)

Download the drafter first (in offline mode it must be in the cache), then add one argument to the existing serving command:

--speculative-config '{"method":"mtp","model":"google/gemma-4-31B-it-assistant","num_speculative_tokens":3}'

The log shows Resolved architecture: Gemma4MTPModel and the Gemma4 MTP: draft layer 0–3 mapping when it is attached. About three minutes to load the 32.6 GB model, about six including compile and warm-up; the drafter adds roughly 0.9 GB. vLLM warns that num_speculative_tokens > 1 may lower the acceptance rate; the table above is with 3. Whether 2 or 4 is better was not measured.

Conditions and what we do not know

One DGX Spark GB10, DGX OS 7.2.3, driver 580.173.02, vLLM 0.27.1 (custom image), RedHatAI gemma-4-31B-it-FP8-Dynamic, 16k context, --gpu-memory-utilization 0.55, one request at a time. The probe is a single 200-token generation, so the ratio is for short answers; long generations and many concurrent requests were not measured. Other models (OTel 31B and the rest) cannot reuse this drafter; each needs its own. A community-trained Korean Eagle-3 drafter reported 1.3–1.5× on the same machine; Google’s official drafter did better.

Related notes: DGX Spark operations · Choosing a local writer