Korean decoding gets 54% faster — a draft vocabulary for the MTP head
Notes · Measurement
Korean decoding gets 54% faster — a draft vocabulary for the MTP head
16.9 to 26.1 tok/s on Korean prompts, from the MTP head's own 65k vocabulary file, not a separate drafter
2026-10-07
A different lever from attaching a smaller drafter model (see the same 31B, twice as fast): an MTP head’s own draft vocabulary. Speculative decoding restricts the draft head’s output projection to a frequency-ranked subset of the vocabulary; the target model still verifies over the full vocabulary, so the output is unchanged — only the draft gets cheaper. The subsets shipped today are ranked on English and code, so Korean tokens fall outside them and Korean drafts get rejected, slowing decoding down. MiaAI-Lab already ships language-extended files for seven languages; Korean was missing, so we built one.
| shipped en+code vocabulary | + Korean 65k file | |
|---|---|---|
| Decode, 20 fixed Korean prompts, median of 3 | 16.9 tok/s | 26.1 tok/s (+54%) |
| Draft acceptance | 1.33 accepted/draft | 2.16 accepted/draft |
| Quality (replacement characters, Korean-script fidelity, veneta-bench’s 12 memory-ability cases) | — | all clean, unaffected |
Qwen3.8-Flash-Next (nvidia/Qwen3.8-Flash-Next-NVFP4), one DGX Spark, vLLM (MiaAI-Lab’s kit), MTP k=3. The vocabulary file, the build script and every raw run file are public under Apache-2.0: GitHub · Hugging Face.
Only Qwen3.8-Flash-Next works today. EXAONE 4.5 and GLM 5.3 Flash are confirmed to need the same thing — not a file, an engine patch — and that patch is under review in vLLM (PR #60387). The first measurements we reported for it (Llama 3.1 8B + EAGLE-1, EXAONE 4.5) are withdrawn pending re-measurement: a reviewer pointed out that the patched code path runs only when a configuration flag is on, and our run records do not show the flag was set, so those numbers cannot be attributed to the patch. We will re-measure with the flag explicitly on and off and publish the run files with the result, whichever way it comes out. DeepSeek V4(.1) turned out to be a different question entirely — it uses its own drafting mechanism (“DSpark”), not ours, so it’s parked alongside Gemma 4 as a separate investigation rather than a target for this patch. Engine support is vLLM only today; Ollama, LM Studio and SGLang support is planned. More languages, across Asia, Africa and the West, are planned on the same recipe after Korean.
We are not first. The idea of trimming a drafter’s vocabulary isn’t ours — vLLM issue #58578 (akapug) and PR #59740 (stecasta, “Context Aware Sparse LM Head”, a more sophisticated implementation, measured on the same model and hardware class) came before us. Their numbers (+15.2%/+22.6%) differ from our +54%; we have not yet established why. Our work addresses, in a simpler way, the model families that PR does not cover. Full detail in the repository.
Related: The same 31B, twice as fast · Choosing a local writer
