The Lab

Work the mechanism instead of only reading the explanation.

Investigations rather than lessons. Some steps ask for your prediction before they show the calculation; others ask you to adjust a figure, complete a calculation, compare cases or explain a result. Start with one gradient update worked by hand; every other investigation is listed below it. Each page says what, if anything, it keeps in this browser.

Runs
In your browser
Saved
On some pages, in this browser only
Account
None needed
Scoring
None

Start here: from text to one gradient update

A language model learns by predicting the next token and adjusting its weights according to how wrong the prediction was. Work one such update by hand first, on a string short enough that every number can be checked. Then let a small transformer repeat it hundreds of times in your browser, and read back the record of what it did.

What runs where

All three run on your device. The trainer keeps nothing after you close the tab; download its run record to keep it.

Investigations

Each begins from one question about how a model behaves. Some run on real benchmark features, some on small models you train, some on exact worked examples; each says which.

Your own questions, and what you keep from these, stay in this browser and are listed in Your work.

Learn the ideaWaterbirds · 1 of 3Why an average can hide a failing group, worked out exactly on a small two-feature model (about 8 minutes).Train a small model on itWaterbirds · 2 of 3A synthetic three-feature task: you train a small network here, then look inside one checkpoint.Is this bird classifier ready to ship?Waterbirds · 3 of 3A classifier that is right on average fails nearly half of one group of birds. Find out why, and decide what would have to change before it ships. Real benchmark features (Waterbirds); the classifier’s last layer is refit in your browser.What should count as better retrieval?InvestigationFix a way of scoring two retrieval methods, then apply it unchanged to a changed case and see whether the verdict survives. Two supplied synthetic cases, calculated exactly in your browser; no retrieval system is run.What will this token copy?InvestigationPredict which value a query will draw on in one attention calculation, then check the weights that decide it. One worked attention example with exact numbers, in the attention lesson.Moving the camera or moving the object?InvestigationFollow the same four landmarks in world coordinates, camera coordinates and a pinhole image, then see the image that results when the camera transform is applied in the wrong direction. Nothing is saved.Learning with an AI assistantPrototypePredict whether one gradient step lowers a loss, check the calculation, then take the same example to an assistant you already use. Hints are prepared in advance; no AI model is connected. Your investigations are saved in this browser and listed in Your work.

Investigations inside the notebooks

Seven investigations sit inside a notebook, a Lab step or Research, next to the material they test. Each works one exact example small enough to check by hand, and none of them saves anything.

What weighting did this data decision create?In the first transformer stepKeep, drop or down-weight one of three training records, one an exact mirror of another, and see which objective the next gradient step follows.Four tasks, sixteen rows: which signs may reverse?In the evaluation stepA variant compared with a baseline on four tasks, with four paired trials on each: what the rows fix, and which reversals a null hypothesis allows.What does one scored group change?In the RLHF notebookFour scored episodes become relative weights, one gradient step moves the logits, and the clipped objective scores the result against the same records.What did the recorded preferences identify?In the DPO notebookFour preference records about one prompt fix one update, yet policies that fit them equally well can treat a response the records never compared very differently.Which error do more sampling steps reduce?In the flow-matching notebookOn a coupling where every quantity has a closed form: the labels a field averages, an error in the field itself, and the error Euler steps add. More steps reduce only the last.Does this observation justify deleting a connection?In Circuit discoveryOn a hand-built graph, replace one named input and recompute everything downstream, to tell a value that carries information from one a later computation uses.What finished, and does it still answer your question?In ResearchOne exact trace of an assistant asked to save a comparison, under three delivery schedules fixed in advance: what was said, what the service did, and which answer still fits.

Transformer systems

Eight steps, meant to be worked in order. They start at a single gradient update and end with you stating a claim about a serving tradeoff and the evidence that would settle it. Each step can be opened on its own, but later steps reuse the quantities built in earlier ones.

Why prediction first

Several steps ask for a prediction before they show the calculation, so the result is compared with an answer you gave in advance rather than one settled on after seeing the numbers. The KV memory step also offers Just show me, which opens the calculation without recording or saving anything.

What an investigation looks like

One formula to explore directly, separate from the eight steps. The numbers with a dashed underline in the sentences below are controls: drag one sideways, or focus it and use the arrow keys, and every result updates. Nothing here is saved, and there is no score.

Try this first

Before you drag the context length, predict two things: what doubling the context does to the cache, and whether sharing key/value heads saves a larger fraction of it at longer context. Then drag it and compare.

A transformer serving one sequence keeps a key and a value for every token it has already read, at every layer. Take a model of whose width of 4,096 is split into of 128 dimensions each, serving in . If every query head keeps its own key and value, the cache for this one sequence is 2.15 GB.

Now let each key/value head serve . The model still runs 32 query heads, but it stores 8 key/value heads per layer, and the cache falls to 0.54 GB, 75% less.

010203040506070032K64K96K128KContext length (tokens)GB
4:1

8 key/value heads per layer. This is grouped-query attention.

32
fp16

2 bytes per stored value.

One KV head per query headGrouped, 4:1One KV head for all queries
Cache size grows in a straight line with context. With layers, head width and precision fixed, the slope is set by how many key/value heads are kept. Sharing lowers the slope but leaves the line straight, so the fraction saved is the same at every context length (75% at 4:1, for example); only the number of gigabytes saved grows with context.
Mem_KV = B × N_layers × T × H_kv × d_head × 2 × bytes
       = 1 × 32 × 4,096 × 8 × 128 × 2 × 2
       = 536,870,912 bytes ≈ 0.54 GB

  B       = 1 sequence
  H_kv    = query heads / group size   (32 / 4 = 8)
  d_head  = model width / query heads  (4,096 / 32 = 128)
  2       = one key and one value per token
  bytes   = bytes per stored value     (fp16: 2)
  GB      = 10⁹ bytes
The arithmetic, so the numbers above can be checked rather than trusted. The controls start at the reference case that the KV-cache reference derives step by step: 32 layers, 32 query heads of 128 dimensions, 4,096 tokens in fp16, grouped four to one.

Guided investigations

Longer routes that link notebooks and the steps above, for when you want a whole topic rather than one mechanism. Opening a route records no progress.

Single-sitting checks

Small and complete. Each takes your prediction before it shows the result to compare it with, and each says what, if anything, it keeps in this browser.

Or bring your own question

Write the question you are trying to answer. Nothing here answers it: it opens in the research workbench, where you can turn it into a test and decide whether to keep it.

No AI model is connected, so nothing answers here yet. Your question opens word for word in the framing workbench, where you say what would count as an answer and the smallest test that could settle it. If you keep it there, it stays in this browser.