The Lab
Work the mechanism instead of only reading the explanation.
Investigations rather than lessons. Some steps ask for your prediction before they show the calculation; others ask you to adjust a figure, complete a calculation, compare cases or explain a result. Start with one gradient update worked by hand; every other investigation is listed below it. Each page says what, if anything, it keeps in this browser.
Start here: from text to one gradient update
A language model learns by predicting the next token and adjusting its weights according to how wrong the prediction was. Work one such update by hand first, on a string short enough that every number can be checked. Then let a small transformer repeat it hundreds of times in your browser, and read back the record of what it did.
All three run on your device. The trainer keeps nothing after you close the tab; download its run record to keep it.
Investigations
Each begins from one question about how a model behaves. Some run on real benchmark features, some on small models you train, some on exact worked examples; each says which.
Your own questions, and what you keep from these, stay in this browser and are listed in Your work.
Investigations inside the notebooks
Seven investigations sit inside a notebook, a Lab step or Research, next to the material they test. Each works one exact example small enough to check by hand, and none of them saves anything.
Transformer systems
Eight steps, meant to be worked in order. They start at a single gradient update and end with you stating a claim about a serving tradeoff and the evidence that would settle it. Each step can be opened on its own, but later steps reuse the quantities built in earlier ones.
Several steps ask for a prediction before they show the calculation, so the result is compared with an answer you gave in advance rather than one settled on after seeing the numbers. The KV memory step also offers Just show me, which opens the calculation without recording or saving anything.
What an investigation looks like
One formula to explore directly, separate from the eight steps. The numbers with a dashed underline in the sentences below are controls: drag one sideways, or focus it and use the arrow keys, and every result updates. Nothing here is saved, and there is no score.
Before you drag the context length, predict two things: what doubling the context does to the cache, and whether sharing key/value heads saves a larger fraction of it at longer context. Then drag it and compare.
A transformer serving one sequence keeps a key and a value for every token it has already read, at every layer. Take a model of whose width of 4,096 is split into of 128 dimensions each, serving in . If every query head keeps its own key and value, the cache for this one sequence is 2.15 GB.
Now let each key/value head serve . The model still runs 32 query heads, but it stores 8 key/value heads per layer, and the cache falls to 0.54 GB, 75% less.
8 key/value heads per layer. This is grouped-query attention.
2 bytes per stored value.
Mem_KV = B × N_layers × T × H_kv × d_head × 2 × bytes
= 1 × 32 × 4,096 × 8 × 128 × 2 × 2
= 536,870,912 bytes ≈ 0.54 GB
B = 1 sequence
H_kv = query heads / group size (32 / 4 = 8)
d_head = model width / query heads (4,096 / 32 = 128)
2 = one key and one value per token
bytes = bytes per stored value (fp16: 2)
GB = 10⁹ bytesGuided investigations
Longer routes that link notebooks and the steps above, for when you want a whole topic rather than one mechanism. Opening a route records no progress.
Single-sitting checks
Small and complete. Each takes your prediction before it shows the result to compare it with, and each says what, if anything, it keeps in this browser.
Or bring your own question
Write the question you are trying to answer. Nothing here answers it: it opens in the research workbench, where you can turn it into a test and decide whether to keep it.