Transformer Lab
KV memory to serving pressure
Start from the KV-cache memory equation, predict what changes when multi-head attention becomes grouped-query attention, check the numbers, then follow the same relationship into long context and serving.
Mem_KV = B * N_layers * T * H_kv * d_head * 2 * bytesOne memory equation links how attention works to what long contexts and serving cost.
BN_layersTH_kvd_headbytesTransformer Systems Lab
KV memory to long-context pressure
- 01Text to one update
- AtlasTokens and position
- AtlasAttention routing
- 02RoPE phase
- 03KV memory
- 04Long-context pressure
- 05Serving and decoding
- 06Speculative decoding
- 07Evaluation and falsification
- 08Capstone systems claim
Step 1 of 8
From text to one gradient update
Take a toy sentence, choose a tokenization, let the last context token attend to itself and the tokens before it in one causal attention row, then work through the exact softmax cross-entropy update.
Example
One next-token training example
the model learns patternsThe tokenizer here is deliberately tiny. It is not BPE, a unigram language model or real SentencePiece; it shows how token boundaries change the positions the loss is computed at. For the last training position, the context is every token before the last one. The target y is one of three toy output classes; by default it is patterns, the word that ends the sentence. All vectors and weights are hand-set toy numbers.
Your prediction
With the key of model scaled by 2x, which source token receives the largest gradient from the loss on the last token, through its value and the residual path?
Controls
Change one setting, then predict again.
Coarse: fewer pieces, so fewer next-token positions to train on.
Companion to step 1
One token, one attention-gradient step
The smallest training step that still contains attention: one query from the last token, two keys, two values, one target and a scalar loss.
Example
How much gradient each value receivesd_k = 1; dL/dv_j = (y_hat - y) * alpha_jThis step tests whether the update to a value follows the size of the value, the distance to the target, or the attention weight that carried the value into the output.
A common confusion
Values are not attention weights.The raw value can be large while its gradient is small: its gradient is scaled by the attention weight the softmax gave it.
Then
Change the keys, then ask again.Changing k_0/k_1 changes alpha, which changes the value gradients.
Your prediction
Which value scalar receives the larger gradient magnitude?
Controls
Change one setting, then predict again.
q_1 and the key gap change the attention weights. The values and the target change the shared error. The learning rate changes the size of the update, not the gradient.
fixed scalar slice1.000[-0.300, 0.300][0, 0] final-token row[0.200, 1.200]shown after you answershown after you answershown after you answerForward pass
s = q_1 * kalpha = softmax(s)y_hat = alpha_0 * v_0 + alpha_1 * v_1L = 0.5 * (y_hat - y)^2
Step 2 of 8
RoPE: rotate the query and key by position
The same kind of query–key score as step 1, now with positions. Which part of the score survives when both positions change?
Example
One query and one key in two dimensions, at one frequency
(R_m q)^T (R_n k)R_m rotates the query q by m·ω, its position m times the frequency ω; R_n rotates the key k by n·ω. The vectors q and k stay fixed; only the positions and the frequency change, and each rotation is small enough to check by hand.
Your prediction
Shift both positions by +3: query 2 → 5, key 0 → 3. What happens to the query–key score?
Controls
Change a setting, then predict again.
Medium (default): 0.42 rad per position, enough movement to see each change clearly.
Sources and the notebook link appear with the result, after you lock in a guess or choose Just show me, so they cannot give the answer away.
Attention head sharing
How are the heads shared?
Thirty-two query heads share stored keys and values in some pattern. The pattern and the result are shown after you answer.
- Head diagram
- Equation
- Code
- Calculated result
The head count and memory are shown after you lock in a guess or choose Just show me.