Transformer Systems Lab

Speculative decoding

A small draft model proposes tokens and the target model checks them. A faster path counts only if the output still follows the target model and the target model really does less work.

Step 6 of 8Predict, then checkSaved in this browser

What this step asks

A speed-up claim has two tests: the output is unchanged, and the target model does less work.

Draft tokens can make generation faster. The output keeps the target model's distribution only if the target model checks each draft token and, when it rejects one, samples a replacement from the leftover distribution (Leviathan et al., 2023).

PredictionWhich intervention makes the served request faster while preserving the target distribution?
From the last stepThe serving workload, with its time to first token and time per output token.
What you checkOne worked round: which tokens are accepted, the cost of the same output (R = e / C), whether the output still follows the target model, and a toy estimate for decoding the whole request.
NextTry to break the claim that the output is faster and just as good, on one evaluation slice.

Step 6 of 8

Speculative decoding

Checking saved workNo prediction yet
  1. 01Text to one update
  2. AtlasTokens and position
  3. AtlasAttention routing
  4. 02RoPE phase
  5. 03KV memory
  6. 04Long-context pressure
  7. 05Serving and decoding
  8. 06Speculative decoding
  9. 07Evaluation and falsification
  10. 08Capstone systems claim

From the serving step

A faster decoder counts only if the output still follows the target model.

The serving step showed that decode time grows with every output token. This step asks whether a small, fast draft model can cut the number of target-model steps without changing the distribution the target model samples from.

Serving stepChecking this browser for saved work.
Long-context stepChecking this browser for saved work.
KV memory stepChecking this browser for saved work.

The example

One round of draft and check on the served request

accept xᵢ with probability min(1, p(xᵢ) / q(xᵢ))

The draft model q proposes k tokens. The target model p scores all of them in one pass and accepts them in order, each with the probability above. At the first rejection, a replacement is drawn from the residual distribution, normalize(max(0, p − q)); this page calls that step residual repair. Cost is counted in target-model steps.

Predict first

Which intervention makes this served request faster while preserving the target distribution?

Not graded. After you choose, the page computes what each move does at your settings; your first choice is kept and not marked right or wrong.

Try it

Change the draft’s agreement, length and cost, and how the target checks it.

Every setting is an assumed value, not a measurement. After you choose, the page works one round from these values: which tokens are kept, what the round costs, and an estimate for the whole request.

Draft–target agreement d
Draft length k
Draft cost c per token
Target check
After a rejection

Assumed values: d = 0.82 · k = 4 · c = 0.25 units per drafted token · parallel check · residual repair on (0.18 extra units in a round with a rejection) · 1 unit = one target-model step

The served request

The same request settings as the serving step, used for the whole-request estimate.

Choose an answer first, or open the worked round without choosing.
Outline of this lesson
ExampleOne served request: a draft model proposes, the target model checksq drafts x1...xk; p checks them; resampling a rejection keeps p
Draftcheap qproposes k tokens
Target checkp checks the drafted tokensaccepted tokens advance
Resamplefrom the leftover distributionthe rejected position still follows p
01Example

one served request, one draft-and-check round

02Question

faster is not the same as unchanged output

03Prediction

which change helps

04Controls

draft match rate, k, cost, the target check, resampling

05Result

accepted tokens and the target model’s checks

06What stays true

what must hold for speed, and for the target distribution?

07Next

try to break the claim that it is faster

Where this step sits

Faster, but what does that prove?

In this worked example speculative decoding reduces the number of target-model passes. The next step asks not “is it faster?” but “which narrower claim about the output survives an evaluation?”

  1. 04Long-context use

    which tokens a model can actually use

    previous
  2. 05Serving and decoding

    time to first token, time per output token, cache memory, batching

    previous
  3. 06Speculative decoding

    the target model checks draft tokens and resamples rejections

    current