Transformer Lab

KV memory to serving pressure

Start from the KV-cache memory equation, predict what changes when multi-head attention becomes grouped-query attention, check the numbers, then follow the same relationship into long context and serving.

saved in this browserprediction firstKV memorylong context and serving
The equation this path followsKV cache memory equationMem_KV = B * N_layers * T * H_kv * d_head * 2 * bytes

One memory equation links how attention works to what long contexts and serving cost.

BN_layersTH_kvd_headbytes
See it in the efficient-attention notebook
QuestionWhich term changes when full MHA becomes GQA?
PredictionChoose before seeing the totals: T, H_kv, bytes or layers?
ResultThe memory for MHA, GQA and a longer context.
What stays the sameWith the model’s shape fixed, KV memory is proportional to T × H_kv.

Transformer Systems Lab

KV memory to long-context pressure

Checking saved workReading your saved answers in this browser
  1. 01Text to one update
  2. AtlasTokens and position
  3. AtlasAttention routing
  4. 02RoPE phase
  5. 03KV memory
  6. 04Long-context pressure
  7. 05Serving and decoding
  8. 06Speculative decoding
  9. 07Evaluation and falsification
  10. 08Capstone systems claim

Step 1 of 8

From text to one gradient update

Take a toy sentence, choose a tokenization, let the last context token attend to itself and the tokens before it in one causal attention row, then work through the exact softmax cross-entropy update.

Your answerPrediction neededYour first answer is kept even if you change it later.
Checking saved lesson state

Example

One next-token training example

the model learns patterns

The tokenizer here is deliberately tiny. It is not BPE, a unigram language model or real SentencePiece; it shows how token boundaries change the positions the loss is computed at. For the last training position, the context is every token before the last one. The target y is one of three toy output classes; by default it is patterns, the word that ends the sentence. All vectors and weights are hand-set toy numbers.

0the1model2learns3patterns
4 toy tokens3 next-token positionsh=[0.83, 0.76, 0.31]query from position 2

Your prediction

With the key of model scaled by 2x, which source token receives the largest gradient from the loss on the last token, through its value and the residual path?

First answerNo prediction yetChanging your answer later does not replace this one.

Controls

Change one setting, then predict again.

toy tokenizer
target y
learning rate η
scale of the model key
attention temperature τ
causal mask
residual path

Coarse: fewer pieces, so fewer next-token positions to train on.

Companion to step 1

One token, one attention-gradient step

The smallest training step that still contains attention: one query from the last token, two keys, two values, one target and a scalar loss.

Your workChecking saved workPrediction and controls persist in this browser.

Example

How much gradient each value receivesd_k = 1; dL/dv_j = (y_hat - y) * alpha_j

This step tests whether the update to a value follows the size of the value, the distance to the target, or the attention weight that carried the value into the output.

A common confusion

Values are not attention weights.

The raw value can be large while its gradient is small: its gradient is scaled by the attention weight the softmax gave it.

Then

Change the keys, then ask again.

Changing k_0/k_1 changes alpha, which changes the value gradients.

Your prediction

Which value scalar receives the larger gradient magnitude?

First answerNo prediction yetChanging your answer later does not replace this one.

Controls

Change one setting, then predict again.

q_1 and the key gap change the attention weights. The values and the target change the shared error. The learning rate changes the size of the update, not the gradient.

q_1: 1.00
k_1 - k_0: 0.60
v_0: 0.20
v_1: 1.20
target y: 0.85
learning rate: 0.30
QuantityShapeCurrent value
d_k[1]fixed scalar slice
q_1[1]1.000
k[2][-0.300, 0.300]
m[2][0, 0] final-token row
v[2][0.200, 1.200]
alpha[2]shown after you answer
y_hat[1]shown after you answer
L[1]shown after you answer
token 0token 1
score_0-0.300
score_10.300
alpha_0hidden
alpha_1hidden

Forward pass

  1. s = q_1 * k
  2. alpha = softmax(s)
  3. y_hat = alpha_0 * v_0 + alpha_1 * v_1
  4. L = 0.5 * (y_hat - y)^2

Step 2 of 8

RoPE: rotate the query and key by position

The same kind of query–key score as step 1, now with positions. Which part of the score survives when both positions change?

Your answerPrediction neededYour first answer is kept even if you change it later.

Example

One query and one key in two dimensions, at one frequency

(R_m q)^T (R_n k)

R_m rotates the query q by m·ω, its position m times the frequency ω; R_n rotates the key k by n·ω. The vectors q and k stay fixed; only the positions and the frequency change, and each rotation is small enough to check by hand.

q[0.72, 0.32]k[0.58, 0.45]ω0.42
Open the RoPE notebook

Your prediction

Shift both positions by +3: query 2 → 5, key 0 → 3. What happens to the query–key score?

First answerNo prediction yetChanging your answer later does not replace this one.

Controls

Change a setting, then predict again.

query position m
key position n
equal shift
frequency ω

Medium (default): 0.42 rad per position, enough movement to see each change clearly.

Your guess

Thirty-two query heads stay active while the model changes how they share memory. Which stored quantity becomes smaller?

Back to step 2: RoPE

A wrong guess still gives you something concrete to compare. Your first guess is saved in this browser only, and later answers do not replace it. “Just show me” saves nothing.

Sources and the notebook link appear with the result, after you lock in a guess or choose Just show me, so they cannot give the answer away.

Attention head sharing

How are the heads shared?

Shown after you answer

Thirty-two query heads share stored keys and values in some pattern. The pattern and the result are shown after you answer.

  1. Head diagram
  2. Equation
  3. Code
  4. Calculated result

The head count and memory are shown after you lock in a guess or choose Just show me.