Step 3 of 8

Grouped-query attention changes what the model stores.

Predict what shrinks when query heads share keys and values, then check it with a diagram, the equation, code and a calculation.

Predict, then checkAnswers saved in this browserCalculated from the formula

Grouped-query attention and KV memory

When query heads share cached keys and values, which part of the cache shrinks, and what does the calculation leave out?

equationKV cache memory equation

Equation referenceequation:attention-transformers/efficient-attention#math-object-2
Checking saved workReading your saved answers in this browser
Your guess

Thirty-two query heads stay active while the model changes how they share memory. Which stored quantity becomes smaller?

Back to step 2: RoPE

A wrong guess still gives you something concrete to compare. Your first guess is saved in this browser only, and later answers do not replace it. “Just show me” saves nothing.

Sources and the notebook link appear with the result, after you lock in a guess or choose Just show me, so they cannot give the answer away.

Intuition

Picture the heads before you see the result.

Attention head sharing

How are the heads shared?

Shown after you answer

Thirty-two query heads share stored keys and values in some pattern. The pattern and the result are shown after you answer.

  1. Head diagram
  2. Equation
  3. Code
  4. Calculated result

The head count and memory are shown after you lock in a guess or choose Just show me.