The idea

A language model writes text token by token, and every step costs a full pass through the network. In speculative decoding a small, fast part guesses several tokens ahead, and the large model checks the proposals in a single pass. If they hold, one pass yields several tokens. If they do not, the model discards everything from the first deviation and continues decoding normally. The method shortens only the time and leaves the output distribution unchanged, apart from rounding effects.

Qwen3.8-27B ships with its own drafting part, a small additional head published with the model. That removes the need for a second model and for matching the two. The vLLM server can use this head, and the only setting left is the depth, the number of tokens the head guesses ahead per step.

The measurement

We measured on 25 September 2026 on an RTX 5090 with 32 GB. The model ran in the four-bit NVFP4 version with a context window of 130,000 tokens, one request at a time and thinking mode off, and we took the median of three runs. The three columns stand for three kinds of task, because guessing succeeds unevenly: reproducing a file with a small change, writing new code from a specification, explaining a subject in prose.

DepthEdit a fileWrite codeProseProposals accepted
none23.723.723.7—
149.948.547.395%
271.868.262.089%
393.986.475.586%
4115.496.983.979%

The table gives tokens per second. At depth 4 the same card writes three and a half to five times as fast. The task in which the model copies most gains most, and for a coding agent that is the normal case, because it returns existing code with small changes. The share of accepted proposals falls with depth from 95 to 79%, which is why the gain flattens. One inconsistency remains: arithmetically, depth 1 can at most double the speed, and 2.1 times was measured. We therefore measured the baseline slightly too low, within the spread that becomes visible further down.

Where memory becomes the limit

The additional head takes about 1.3 GiB and brings the loaded model to 22.15 GiB, and that alone squeezes memory. What remains after the working memory for computation goes to the cache for the context. By default vLLM uses 90% of the card. That left 4.35 GiB for the cache, while a single request of 130,000 tokens needs 4.73, and the server refused to start. Only with a higher share of the card did the full window fit again. At depth 5 and 6 memory was no longer enough for such a request even then.

The same arithmetic applies to CUDA graphs, which vLLM normally uses by default and which the measurement above ran without. With them vLLM records the sequence of computation steps on the card once and then replays it as a whole. The graphs speed up writing and need memory themselves, and together with the head and 130,000 tokens that does not fit on the card. With 100,000 tokens it does:

SettingEdit / code / proseReading, tokens per second
130,000 tokens, without CUDA graphs, depth 497 / 81 / 705,200
100,000 tokens, with CUDA graphs, depth 4228 / 190 / 1595,500
100,000 tokens, with CUDA graphs, depth 5253 / 202 / 159—

This second series dates from the same day. Its first row, at 70 to 97, lies below the 84 to 115 of the table before it although the setting is the same; we list both measurements and do not average them. At 110,000 tokens with graphs, memory was no longer enough for one full request. The choice therefore lies between a window of 130,000 tokens at around 100 tokens per second and 100,000 tokens at around 200. The table does not show what the smaller window costs, because the agent then has to compact its history earlier.

Under load

The measurement above runs with one request, without thinking mode and with the most probable continuation. With the agent in operation we read the server’s counters during a benchmark of five harder tasks. There the acceptance rate of the four guessed tokens fell to 73, 54, 42 and 33%, because the model ran at temperature 1.0 and writes less predictably while reasoning. We then lowered the temperature for local models to 0.6.

In the same run the agent read 2.57 million tokens and wrote 158,000, a ratio of sixteen to one. 74% of what it read came from the cache, because the beginning of the request stayed the same. The cache thus spares just under three quarters of the reading, and a history with a stable beginning pays off alongside faster writing. In a second, simpler benchmark the agent passed all five tasks at depth 4, in 22 to 94 seconds per task.

Limits

We measured the baseline of 23.7 tokens per second without CUDA graphs. How fast the model writes with graphs and without drafting we did not measure, and part of the gain in the second table is due to the graphs. All figures hold for this card, this model in this quantisation and the vLLM version of September 2026. The single-request measurement says nothing about several simultaneous users, who share the cache. The acceptance rates under load come from a single benchmark run. The two series of the same day differ by about a sixth for the same setting without our knowing the cause, and that difference shows how exact the figures are.