Selected connection: Embeddings to Attention: Comparison produces scores
Comparison produces scores
Comparison
Embeddings supply coordinates. Attention compares a query with each key using a dot product, then divides by the square root of the key dimension. The result is one compatibility score per key, not yet a share of the output.
One concrete example
Let the query be [1] and keys A, B, C be [ln 2], [0], [0]. Their dot products are [ln 2, 0, 0]. Here the key dimension is 1, so division by its square root leaves those scores unchanged. ln 2 is approximately 0.693.
All three connections use this same small example, with no masking or dropout.
What stays true. For fixed queries and keys, changing a value does not change these scores.
Where this connection stops
A dot product depends on magnitudes as well as alignment. A large score is not proof that a position is relevant or correct.
Try it yourself
In the attention notebook, choose Reset example, then Scores. Move Query q, x from 2 to 1; leave Query q, y at 1 and keep keys and values fixed. Which scores change, which stays the same, and how are the unchanged values mixed afterward?
Investigate AttentionPrerequisites and related concepts
Prerequisites for Attention
Taken from the Foundations index, not from the connection you selected. Open one only if you need to review it.
Related concepts (17)
Each link keeps the direction and type given in the Foundations index. A related concept is not a prerequisite.
- Attention to MoE · same trick: Softmax routingOpen MoE
- Attention to Diffusion · same trick: Cross-attentionOpen Diffusion
- Attention to Efficient Attention · breaks when: O(n²) memoryOpen Efficient Attention
- Attention to RoPE · invented to fix: Position encodingOpen RoPE
- Attention to Efficient Attention · invented to fix: KV cache memoryOpen Efficient Attention
- MoE to Attention · analogy: Sparse selectionOpen MoE
- Attention to Embeddings · same trick: Dot-product logitsOpen Embeddings
- Attention to Decoding & Sampling · same trick: Softmax selectionOpen Decoding & Sampling
- Attention to State Space Models · duality: Kernel ↔ RecurrenceOpen State Space Models
- Attention to Long Context · breaks when: Attention dilutionOpen Long Context
- Attention to Speculative Decoding · invented to fix: Autoregressive latencyOpen Speculative Decoding
- Attention to LLM Serving · invented to fix: Production servingOpen LLM Serving
- Attention to Multimodal Pretraining · same trick: Cross-attention fusionOpen Multimodal Pretraining
- Circuits to Attention · same trick: Copy via attentionOpen Circuits
- State Space Models to Attention · invented to fix: Linear complexityOpen State Space Models
- Multimodal Pretraining to Attention · same trick: Shared attentionOpen Multimodal Pretraining
- Efficient Attention to Attention · invented to fix: Attention costsOpen Efficient Attention