Canonical learning room
Grouped-query attention changes what the model stores.
Make one sealed prediction, then connect the same variables across intuition, tensor shape, equation, code, and a local memory calculation.
Canonical learning room
Grouped-query attention and KV memory
When query heads share cached keys and values, which stored axis changes—and what stays outside this calculation?
Make one prediction, then test it against one deterministic comparison.
Need the source context? Lock in a guess or choose Just show me. The registered notebook appears with the calculation so it cannot contaminate the pre-reveal choice.
Intuition
Trace reuse before seeing the measured direction.
Attention head sharing
One mechanism, still sealed
Thirty-two query heads enter a sealed head-sharing comparison. The changed quantity and calculated result are not shown.
- Head diagram
- Equation
- Code
- Calculated result
The evaluated count and memory stay sealed until you lock in a guess or choose Just show me.