Concept connections

Understand a connection before following it.

Choose one connection around attention. Read how it works and a small worked example, then try it yourself in a notebook.

Why do these ideas connect?

Start with attention: comparing a query with keys makes scores, normalizing the scores makes weights, and a weighted sum mixes the values. Select a connection to read about it here; only the links lead to other pages.

Four concepts and three connections between them. A connection explains how one idea is used in another; it is not a list of prerequisites or a suggested order of study.

Choose a connection

Selected connection: Embeddings to Attention: Comparison produces scores

Comparison produces scores

Comparison

Embeddings supply coordinates. Attention compares a query with each key using a dot product, then divides by the square root of the key dimension. The result is one compatibility score per key, not yet a share of the output.

One concrete example

Let the query be [1] and keys A, B, C be [ln 2], [0], [0]. Their dot products are [ln 2, 0, 0]. Here the key dimension is 1, so division by its square root leaves those scores unchanged. ln 2 is approximately 0.693.

All three connections use this same small example, with no masking or dropout.

What stays true. For fixed queries and keys, changing a value does not change these scores.

Where this connection stops

A dot product depends on magnitudes as well as alignment. A large score is not proof that a position is relevant or correct.

Try it yourself

In the attention notebook, choose Reset example, then Scores. Move Query q, x from 2 to 1; leave Query q, y at 1 and keep keys and values fixed. Which scores change, which stays the same, and how are the unchanged values mixed afterward?

Investigate Attention
Prerequisites and related concepts

Prerequisites for Attention

Taken from the Foundations index, not from the connection you selected. Open one only if you need to review it.

Related concepts (17)

Each link keeps the direction and type given in the Foundations index. A related concept is not a prerequisite.

Attention concept connectionsEmbeddings to Attention: Comparison produces scores. Attention to Decoding & Sampling: Normalization turns scores into shares. Attention to MoE: Weighted aggregation mixes values. Selected: Embeddings to Attention: Comparison produces scores. Use the concept buttons and the list of connections to select; the map does not require dragging.Embeddings to Attention: Comparison produces scores1Attention to Decoding & Sampling: Normalization turns scores into shares2Attention to MoE: Weighted aggregation mixes values3EmbeddingsAttentionDecoding & SamplingMoE
The same three connections as a diagram. Numbers match the list; each arrow points from the first concept named to the second. Where a box sits carries no meaning.
Learning routes for research questions

Learning routes

Choose a question and see what to learn next.

Each question below has a prepared route: its prerequisites, the equations and experiments it involves, and why each idea leads to the next. If you came from the paper map, the route you saved there stays visible.

Answer

I know transformers and Adam. What do I need for Mamba-2?

First see why long contexts are costly for attention (the KV cache grows with every token), then see how a fixed-size recurrent state avoids that cost.

Search this routeChecking whether this route is saved in this browser.
Selected itemrecurrence, attention, or control theory?h_t = A(x_t)h_{t-1} + B(x_t)x_t

First see why long contexts are costly for attention (the KV cache grows with every token), then see how a fixed-size recurrent state avoids that cost.

questionexampleconnectionnext step
Go to the discussion questions
Part of the routeSteps 1-3 of 7, around Efficient Attention
Full routeShow all 7 steps
QuestionI know transformers and Adam. What do I need for Mamba-2?
Short answerFirst see why long contexts are costly for attention (the KV cache grows with every token), then see how a fixed-size recurrent state avoids that cost.
On this route4 connections and 4 papers, equations, experiments and questions.
First connectionAttention as a weighted average of values comes before any method that saves its memory.
ReviewWhat should you review first?

Efficient Attention: why the KV cache is costly

Open Efficient Attention

Route finder

Find a path from what you know to a new topic.

Choose the concepts you already know and a topic you want to reach. The finder searches the concept map for the shortest weighted path between them and names the first concept you would need to learn.

Known concepts

Topic to reach

6 conceptsLearn Efficient Attention next
path length (weighted)5.50
Preview only; changing your choices does not replace your saved route.
01Attentionweighted copying02Efficient Attentionmemory pressure03Long Contextstress regime
04SSM Hybridsfixed-state sequence modelsplanned
05Parallel Scantrainable recurrenceplanned
06State-Space DualityMamba-style bridgeplanned
prerequisiteAttention -> Efficient Attention

You need Q/K/V weighted copying before cache and memory optimizations are meaningful.

invented to fixEfficient Attention -> Long Context

Long context exposes the cost of storing and repeatedly reading all prior keys and values.

same pressureLong Context -> SSM Hybrids

Fixed-state sequence models become attractive when KV memory grows with sequence length.

implementation dependencySSM Hybrids -> Parallel Scan

Recurrent-looking models need parallel sequence computation to train at scale.

paper-specific bridgeParallel Scan -> State-Space Duality

Modern Mamba-style papers often connect recurrence, convolution, and attention-like views.

Connections

Why each idea leads to the next

prerequisiteAttention → Efficient Attention

Attention as a weighted average of values comes before any method that saves its memory.

invented to fixEfficient Attention → Long Context

Long contexts expose the cost of storing and reading all prior keys and values.

same pressureLong Context → SSM Hybrids

Fixed-size states are attractive because KV memory grows with sequence length.

implementation dependencySSM Hybrids → Parallel Scan

Training recurrent models at scale needs a way to compute the whole sequence in parallel.

On this route

Papers, equations, experiments and questions

paperMamba-2 / state-space duality

paper to map

equationh_t = A(x_t)h_{t-1} + B(x_t)x_t

not yet explained on this site

toy experimentfixed state vs KV memory simulator

planned

claimrecurrence, attention, or control theory?

open question

Discussion

Questions to discuss

Pick a paper, equation, experiment, claim or misconception. It is saved in this browser as the focus of this route.
paperpaper to map

Mamba-2 / state-space duality

Starting question

What should we check about Mamba-2 / state-space duality before treating it as understood?

SourcesNo cited source is attached; check it before relying on it.Notes cannot be saved for this item yet
Points of view on this item

Fixed prompts generated from this item. They are not comments from people or independent reviews.

Learner’s first stepAsk what would make "Mamba-2 / state-space duality" feel predictable rather than familiar.
Assumption

No source is attached yet; treat claims about sources as provisional.

Source-checking summary

Turn the paper into cited passages, equations, prerequisites and one runnable check.

Proposed experiment

Choose one claim from the paper and map it to the smallest concept, equation, or toy demo that could test it.

Next action

The paper contribution is mapped to a concrete mechanism

Evidence3 checks
ObservationChecking for a saved observation
ActionCannot be saved yet
PromptLearner prompt ready to copy
01ObservationChecking what this browser has saved
02EvidenceChecking for a saved observation
03SourcesNo cited source is attached; check it before relying on it.
04Next stepSave one next action
Local action draftDraft unavailableNotes cannot be saved for this item yet
Local action draft

Notes cannot be saved for this item yet.

No local draft saved.
Evidence to inspect
  • Paper metadata, abstract claim, and any pasted source spans
  • Related concepts, equations and prerequisites
  • Which claims have been checked against the source and which have not
What would resolve this
  • The paper contribution is mapped to a concrete mechanism
  • Unverified author, date, benchmark, and novelty claims are separated from learning claims
  • The next concept or lab action is specific enough to resume later