Transformer Systems Lab
Speculative decoding
A small draft model proposes tokens and the target model checks them. A faster path counts only if the output still follows the target model and the target model really does less work.
What this step asks
A speed-up claim has two tests: the output is unchanged, and the target model does less work.
Draft tokens can make generation faster. The output keeps the target model's distribution only if the target model checks each draft token and, when it rejects one, samples a replacement from the leftover distribution (Leviathan et al., 2023).
Step 6 of 8
Speculative decoding
- 01Text to one update
- AtlasTokens and position
- AtlasAttention routing
- 02RoPE phase
- 03KV memory
- 04Long-context pressure
- 05Serving and decoding
- 06Speculative decoding
- 07Evaluation and falsification
- 08Capstone systems claim
From the serving step
A faster decoder counts only if the output still follows the target model.
The serving step showed that decode time grows with every output token. This step asks whether a small, fast draft model can cut the number of target-model steps without changing the distribution the target model samples from.
The example
One round of draft and check on the served request
accept xᵢ with probability min(1, p(xᵢ) / q(xᵢ))The draft model q proposes k tokens. The target model p scores all of them in one pass and accepts them in order, each with the probability above. At the first rejection, a replacement is drawn from the residual distribution, normalize(max(0, p − q)); this page calls that step residual repair. Cost is counted in target-model steps.
Predict first
Which intervention makes this served request faster while preserving the target distribution?
Not graded. After you choose, the page computes what each move does at your settings; your first choice is kept and not marked right or wrong.
Try it
Change the draft’s agreement, length and cost, and how the target checks it.
Every setting is an assumed value, not a measurement. After you choose, the page works one round from these values: which tokens are kept, what the round costs, and an estimate for the whole request.
Assumed values: d = 0.82 · k = 4 · c = 0.25 units per drafted token · parallel check · residual repair on (0.18 extra units in a round with a rejection) · 1 unit = one target-model step
The served request
The same request settings as the serving step, used for the whole-request estimate.
Outline of this lesson
q drafts x1...xk; p checks them; resampling a rejection keeps pone served request, one draft-and-check round
faster is not the same as unchanged output
which change helps
draft match rate, k, cost, the target check, resampling
accepted tokens and the target model’s checks
what must hold for speed, and for the target distribution?
try to break the claim that it is faster
Where this step sits
Faster, but what does that prove?
In this worked example speculative decoding reduces the number of target-model passes. The next step asks not “is it faster?” but “which narrower claim about the output survives an evaluation?”
- 04Long-context use
which tokens a model can actually use
previous - 05Serving and decoding
time to first token, time per output token, cache memory, batching
previous - 06Speculative decoding
the target model checks draft tokens and resamples rejections
current - 07Evaluation and falsification
which claim about the faster output can fail
next
Saving your work
Save the result, then take it to evaluation.
Saving keeps your first prediction, the full accounting for the round (its settings, workload, costs and results) and the assumption that rejected tokens are resampled, in this browser only. The evaluation step then asks which claim about the faster output survives.
Continue to Evaluation and falsification