Step 3 of 8
Grouped-query attention changes what the model stores.
Predict what shrinks when query heads share keys and values, then check it with a diagram, the equation, code and a calculation.
Grouped-query attention and KV memory
When query heads share cached keys and values, which part of the cache shrinks, and what does the calculation leave out?
equationKV cache memory equation
Equation reference
equation:attention-transformers/efficient-attention#math-object-2Checking saved workReading your saved answers in this browser
Sources and the notebook link appear with the result, after you lock in a guess or choose Just show me, so they cannot give the answer away.
Intuition
Picture the heads before you see the result.
Attention head sharing
How are the heads shared?
Thirty-two query heads share stored keys and values in some pattern. The pattern and the result are shown after you answer.
- Head diagram
- Equation
- Code
- Calculated result
The head count and memory are shown after you lock in a guess or choose Just show me.