Transformer Systems Lab

Serving and decoding

A prompt can fit in the context window and still be served too slowly. Prefill, decoding, cache memory, scheduling and sampling each add a cost; this step shows which one dominates for one toy workload.

Step 5 of 8Predict, then checkSaved in this browser

What this step asks

A long context is only useful if it can be served in time.

The earlier steps separated the bytes in the cache from whether a model actually uses the context. Serving adds another test: the same workload has to meet its latency targets while its cache stays in memory and many requests are scheduled together.

PredictionWhich cost decides whether this workload meets its latency target?
From the last stepLong context: a context that fits is not the same as a context the model uses well.
What you checkToy numbers for time to first token, time per output token, wasted cache pages, request latency and sampling.
NextSpeculative decoding: a draft model proposes tokens and the target model checks them, without changing the target's output distribution.

Step 5 of 8

Serving and decoding

Checking saved workNo prediction yet
  1. 01Text to one update
  2. AtlasTokens and position
  3. AtlasAttention routing
  4. 02RoPE phase
  5. 03KV memory
  6. 04Long-context pressure
  7. 05Serving and decoding
  8. 06Speculative decoding
  9. 07Evaluation and falsification
  10. 08Capstone systems claim

From the long-context step

A prompt that fits in the context window still has to be served.

The long-context step separated storing tokens from using them. This step asks what it costs to serve the same long prompt to several users at once: processing the prompt (prefill), producing output tokens one at a time (decode), sharing the hardware between requests (batching) and choosing each token (sampling).

Long-context stepChecking this browser for saved work.
KV memory stepChecking this browser for saved work.
Not covered hereServing settings do not change how well a model uses its context, and sampling settings do not change how the cache is allocated.

The example

One long-prompt workload served to several users

W = (prompt, output, requests, block size, decode mode)

A workload is fixed by five settings that act together: prompt length, output length, number of concurrent requests, KV block size and decoding mode. Two times describe how it feels to a user: TTFT, the time to the first output token, and TPOT, the time per output token after that.

Predict first

When this long prompt is served to several users at once, which pressure becomes the bottleneck first?

Try it

Change the workload and see which cost grows fastest.

Illustrative numbers from a simple formula on this page, not measurements of any server. They show the shape of the trade-offs only.

How the numbers are computed

P = prompt tokens, N = output tokens, B = concurrent requests. Times are in milliseconds; every constant is chosen for illustration.

  • TTFT = 180 + 18 × (P / 1,024) × (1 + 0.14 log₂ B)
  • TPOT = (18 + P·B / 18,000 + N / 72) × mode × load, where mode is 1 for greedy, 1.08 for nucleus and 0.72 for speculative, and load is 1, 1.06 or 1.18 for 1, 8 or 32 requests
  • Request time = TTFT + (N − 1) × TPOT
  • Cache slots: each request reserves whole blocks for P + N tokens; the unused share is the empty part of each request’s last block
  • The bottleneck is the largest of five scores: TTFT / 1,800 ms, TPOT / 95 ms, B(P + N) / 1,500,000 + 1.8 × unused share, B / 28 (+ 0.28 when N = 1,024), and a fixed score for the decoding mode (greedy 0.32, nucleus 0.55, speculative 0.64)
Prompt tokens
Output tokens
Concurrent requests
KV block size (tokens)
Decode mode
Choose a pressure first.
Outline of this lesson
ExampleOne long-prompt generation workloadW = prompt, output, requests, pages, decode_mode
PrefillTTFT: time to first tokenprompt enters once
DecodeTPOT: time per output tokenK/V read every step
Schedulerbatchiteration by iteration
01Example

one long-prompt generation workload

02Question

a context that fits vs a request that is served on time

03Prediction

name the bottleneck first

04Controls

prompt, output, requests, block size, decoding

05Result

time to first token, time per output token, cache memory, sampling

06What stays true

the costs of serving depend on each other

07Next

speculative decoding

Where this step sits

Next: a way to decode faster, and what it must not change.

Speculative decoding comes next. It reduces the target model's work only when draft tokens pass the target model's check, which makes it a good test of what you found here.

  1. 03KV memory

    H_kv and T make cache bytes visible

    previous
  2. 04Long-context use

    Fitting tokens is not reliable retrieval

    previous
  3. 05Serving and decoding

    time to first token, time per output token, cache memory, batching, sampling

    current