Transformer Systems Lab
Serving and decoding
A prompt can fit in the context window and still be served too slowly. Prefill, decoding, cache memory, scheduling and sampling each add a cost; this step shows which one dominates for one toy workload.
What this step asks
A long context is only useful if it can be served in time.
The earlier steps separated the bytes in the cache from whether a model actually uses the context. Serving adds another test: the same workload has to meet its latency targets while its cache stays in memory and many requests are scheduled together.
Step 5 of 8
Serving and decoding
- 01Text to one update
- AtlasTokens and position
- AtlasAttention routing
- 02RoPE phase
- 03KV memory
- 04Long-context pressure
- 05Serving and decoding
- 06Speculative decoding
- 07Evaluation and falsification
- 08Capstone systems claim
From the long-context step
A prompt that fits in the context window still has to be served.
The long-context step separated storing tokens from using them. This step asks what it costs to serve the same long prompt to several users at once: processing the prompt (prefill), producing output tokens one at a time (decode), sharing the hardware between requests (batching) and choosing each token (sampling).
The example
One long-prompt workload served to several users
W = (prompt, output, requests, block size, decode mode)A workload is fixed by five settings that act together: prompt length, output length, number of concurrent requests, KV block size and decoding mode. Two times describe how it feels to a user: TTFT, the time to the first output token, and TPOT, the time per output token after that.
Predict first
When this long prompt is served to several users at once, which pressure becomes the bottleneck first?
Try it
Change the workload and see which cost grows fastest.
Illustrative numbers from a simple formula on this page, not measurements of any server. They show the shape of the trade-offs only.
How the numbers are computed
P = prompt tokens, N = output tokens, B = concurrent requests. Times are in milliseconds; every constant is chosen for illustration.
- TTFT = 180 + 18 × (P / 1,024) × (1 + 0.14 log₂ B)
- TPOT = (18 + P·B / 18,000 + N / 72) × mode × load, where mode is 1 for greedy, 1.08 for nucleus and 0.72 for speculative, and load is 1, 1.06 or 1.18 for 1, 8 or 32 requests
- Request time = TTFT + (N − 1) × TPOT
- Cache slots: each request reserves whole blocks for P + N tokens; the unused share is the empty part of each request’s last block
- The bottleneck is the largest of five scores: TTFT / 1,800 ms, TPOT / 95 ms, B(P + N) / 1,500,000 + 1.8 × unused share, B / 28 (+ 0.28 when N = 1,024), and a fixed score for the decoding mode (greedy 0.32, nucleus 0.55, speculative 0.64)
Outline of this lesson
W = prompt, output, requests, pages, decode_modeone long-prompt generation workload
a context that fits vs a request that is served on time
name the bottleneck first
prompt, output, requests, block size, decoding
time to first token, time per output token, cache memory, sampling
the costs of serving depend on each other
speculative decoding
Where this step sits
Next: a way to decode faster, and what it must not change.
Speculative decoding comes next. It reduces the target model's work only when draft tokens pass the target model's check, which makes it a good test of what you found here.
- 03KV memory
H_kv and T make cache bytes visible
previous - 04Long-context use
Fitting tokens is not reliable retrieval
previous - 05Serving and decoding
time to first token, time per output token, cache memory, batching, sampling
current - 06Speculative decoding
A draft model proposes tokens; the target model checks them
next
Saving your work
Save the result, then take the workload to speculative decoding.
Saving keeps the workload, your first prediction, the table of costs and the source caveats, in this browser only. The next step tries speculative decoding on the same workload.
Continue to Speculative decoding