Transformer Systems Lab

Long context

A token can be stored in memory and still be unavailable to attention. Work one eight-token example by hand: what is stored, which query–key pairs a mask allows, and how far stacked layers reach.

Step 4 of 8Worked exampleNext: serving

Stored, visible, reachable: one eight-token example

Checking this browser for earlier answers0 answered · 0 shown · 5 open · checking the route

1The worked setup: eight tokens

Eight tokens sit at positions 0 to 7, so T = 8. Follow one piece of information, whatever token 2 contributed, and ask whether it can influence the newest position, 7. No language model runs on this page; every number below is counted from these definitions.

  1. 0token
  2. 1token
  3. 2token 2
  4. 3token
  5. 4token
  6. 5token
  7. 6token
  8. 7query
Token
One position in the sequence. In every diagram here, one axis position is one token.
Query
What a position uses to ask which positions are relevant. We follow the query at position 7.
Key
What each position offers to be matched. Query i and key j form one query–key pair.
Value
What a position contributes to the output when its pair is allowed and weighted.
Cache
During generation, earlier tokens’ keys and values are stored so they need not be recomputed. A token is “in memory” when its key and value are stored.
Mask
The rule that decides which pairs are allowed at all. A disallowed pair gets weight zero, whatever the key holds.
Layer
Attention repeats in stacked layers. A position’s output from one layer is its input to the next.

For token 2 to affect position 7, three things must hold:

  1. Stored: token 2’s key and value are in the cache (section 2).
  2. Visible: the mask allows the pair in one layer (section 3).
  3. Reachable: a chain of allowed pairs across layers connects it to position 7 (section 4).

None of the three says a model uses the information well. That is measured, not counted (section 7).

2What it costs to keep a token stored

From the KV memory step (optional)

Checking this browser for earlier KV memory work. Open KV memory

A layer keeps one key and one value per KV head for every stored token. With these givens, count one token’s entry, then multiply by the number of tokens.

B1 sequence
N_layers2 layers
H_kv2 KV heads per layer
d_head4 numbers per head
bytes2 per number (fp16/bf16)
2one key and one value

Mem_KV = B · N_layers · T · H_kv · d_head · 2 · bytes

One token (T = 1): 1 × 2 × 1 × 2 × 4 × 2 × 2 = 64 bytes.
Eight tokens: 8 × 64 = 512 bytes.

// Same factors as Mem_KV = B · N_layers · T · H_kv · d_head · 2 · bytes
const kvBytes = ({ B, layers, T, Hkv, dHead, bytes }) =>
  B * layers * T * Hkv * dHead * 2 * bytes // 2: one key and one value

kvBytes({ B: 1, layers: 2, T: 8, Hkv: 2, dHead: 4, bytes: 2 }) // 512

This is nominal payload for the stored numbers. Real caches add block padding and metadata, and none of this is a time.

T = 88 × 64 B = 512 bytes

The T = 32 cache appears after you answer or choose Show me.

Try it

Hold every other given fixed and let the sequence grow to T = 32 tokens. How many bytes do the stored keys and values occupy?

Givens
  • B = 1 sequence
  • N_layers = 2
  • H_kv = 2 KV heads
  • d_head = 4 numbers per head
  • 2 bytes per number
  • T grows from 8 to 32 tokens

Your first answer is kept. “Show me” records that the answer was shown without one.

3Which query–key pairs one layer allows

Draw every pair as a cell: row i is the query at position i, column j is the key at position j. The 8 × 8 grid has 64 cells, one per query–key pair.

A causal mask lets query i read keys 0 through i, never a later position. Row i has i + 1 allowed cells: 1 + 2 + … + 8 = 36. The other 28 cells pair a query with a later key.

Row 7 is query 7. Under the causal mask it can see every earlier token, including token 2. A window also removes keys that are too far back.

The W = 3 grid opens after you answer or choose Show me below.

Rows: query position i. Columns: key position j. One token per position; one query–key pair per cell.1 + 2 + 3 + 4 + 5 + 6 + 7 + 8 = 36 allowed of 64 cells allowed later key (causal)

Try it

Keep T = 8, but now each query may read only itself and the two positions just before it (W = 3, counting the query itself). How many query–key pairs are allowed?

Givens
  • T = 8 tokens, positions 0–7
  • Causal: a query never reads a later key
  • Window W = 3: query i reads keys j with 0 ≤ i − j ≤ 2

Your first answer is kept. “Show me” records that the answer was shown without one.

4How far stacked layers can carry information

With W = 3, one layer lets query 7 read positions 5, 6 and 7: those pairs are directly visible. But position 5 computed its own output in the layer before by reading positions 3, 4 and 5.

After 1 layer, inputs 5–7 can reach position 7. After 2 layers, inputs 3–7 can: 7 reads 5, which had read 3. A token is reachable when some chain of allowed pairs across layers connects it to the query.

Step through the layers. Each row down is one more layer; lines are allowed pairs of the same window as the grid.

Three layers opens after you answer or choose Show me below.

inputlayer 1layer 2layer 301234567
After 1 layer with W = 3, inputs 5–7 can reach position 7. Token 2: not yet reachable. has a path to position 7 token 2

Try it

With W = 3 in every layer, query 7 directly reads positions 5, 6 and 7. How many stacked window layers are needed before information from token 2 can first reach position 7?

Givens
  • W = 3 in every layer, counting the query itself
  • Each layer’s output at position i is position i’s input to the next layer
  • After 1 layer, inputs 5–7 can reach position 7
  • After 2 layers, inputs 3–7 can reach position 7
Choose one

Your first answer is kept. “Show me” records that the answer was shown without one.

5Check a new case

This case changes the window and the depth. The worked diagrams above stay available; the diagram for this exact case appears after you answer or choose Show me.

New case

A new case. Same eight tokens and query at position 7, but the window is W = 4 (the query and the three positions before it) and the model has 2 layers. Which input positions can reach position 7?

Givens
  • T = 8, query at position 7
  • W = 4, counting the query: query i reads keys i − 3 … i
  • 2 stacked layers
Choose one

Your first answer is kept. “Show me” records that the answer was shown without one.

6An assumption sheet, not a measurement

Sections 2–5 counted things. Retrieval quality, whether a model finds and uses the right token among many, cannot be counted from a mask. Evaluations such as Lost in the Middle (Liu et al., 2024) measured it for particular models and tasks and reported that accuracy depended on where the relevant passage sat. This sheet does not reproduce those measurements. It subtracts four terms chosen by this page’s authors, so you can inspect the shape of a claim an evaluation would have to test. Changing a setting recomputes the sheet; it does not run a test.

Nominal T
Evidence position
Task
Distractors
0.94base (chosen)
−0.06context term, T = 32,768
−0.10task term, aggregation
−0.13distractor term, high
−0.22position term, middle
=0.43index for the middle row
Each row uses its own position term
Evidence positionPosition termIndex
Beginning0.0659%
Middle0.2243%
End0.0857%
  • Context term grows with nominal T to stand for the claim that longer inputs are harder. Chosen, not fitted.
  • Task term: aggregation and two-hop need more than one span; lookup needs one.
  • Distractor term: similar-looking passages compete with the evidence.
  • Position term: the middle is penalized most, echoing the reported pattern; the sizes are ours.
  • Limits: the index is clamped to 0.16–0.94; no offered setting reaches either limit.
  • Because the terms add, changing one control shifts every row by the same amount. The sheet cannot show interactions a real evaluation might find.

Read the sheet

At the sheet’s default setting (T = 32,768, aggregation, high distractors) it gives 43% when the evidence is in the middle and 59% when it is at the beginning. What produces that difference?

Givens
  • Index = 0.94 − context term − task term − distractor term − position term
  • Default terms: context 0.06, task 0.10, distractors 0.13
  • Position terms: beginning 0.06, middle 0.22, end 0.08
Choose one

Your first answer is kept. “Show me” records that the answer was shown without one.

Five pressures a long prompt raises, and where each is examined

  • KV memory pressureThe cache still grows as T grows, even after GQA narrows H_kv.Counted in section 2; the KV memory step works the same formula at model scale.
  • RoPE / range pressureThe model may be asked to handle positions beyond the range it was trained on.Not computable here. How a model scores positions beyond its training range must be measured; methods such as position interpolation and YaRN rescale RoPE for longer inputs (see the reference). The RoPE step shows the phase geometry.
  • Retrieval quality pressureEvidence inside the window may still be hard to use under position and distractors.Stated as assumptions in section 6; section 7 describes the measurement it would need.
  • Paging / compression pressureThe cache can fit only after allocator, paging, eviction, or compression tradeoffs.Changes how the cache is laid out or shrunk, not which pairs a mask allows; the serving step works allocation.
  • Serving latency pressureA request can fit in memory while prefill, cache reads, batching, or scheduling bind.Section 8 defines TTFT and TPOT; the serving step works them.

7What this page cannot tell you

  • Whether a trained model uses token 2 when it is reachable, or is swamped by other tokens.
  • How accuracy changes with position, length and distractors for a real model.
  • How long any of this takes. No number on this page is a timing.

Illustrative test question · not run

Question. For one fixed model checkpoint, does exact-match accuracy drop when the answer passage moves from the beginning to the middle of a 32,768-token context?

Design. The same question–document items in both conditions; only the answer passage’s position changes. Distractor documents, prompt template and decoding settings stay fixed. Compare paired per-item outcomes, report the difference with an interval, and decide in advance what size of drop would matter.

Nothing on this page runs this test. It needs a model and a dataset; this lesson is not evidence for either outcome.

8Before the next step: serving terms

Serving
Running a trained model for many requests on shared hardware.
Prefill
Processing all T prompt tokens once, which writes their keys and values into the cache.
TTFT
Time to first token: from a request’s arrival to its first output token, including waiting and prefill.
Decode
Producing output tokens one at a time; each new token’s query reads the cached keys and values.
TPOT
Time per output token: the average time between later output tokens during decode.

Save a summary to this browser’s route

Answers given in the lesson are saved to this browser’s lesson record as you give them, when the status above says so. Saving here also adds a summary to your Transformer Lab route so later steps can refer to it: first answers, hints and reference opened first, answers shown, and the assumption-sheet setting with all three positions. A save records what happened here; it is not a measure of what you learned. If you use the optional account features, the copy they receive leaves out your answers, help and outcomes.

Reference

Formulas, conventions, code, sources and larger interactive figures, open at any time. They can be used to work any question above, so opening the reference is recorded against each question still open, the same way as a hint.

Where this step sits

From stored and reachable tokens to a serving workload.

The lab continues with serving and decoding. State-space hybrids, which replace some attention layers with layers that keep a fixed-size state instead of a growing cache, are a related concept page, not part of this lab.

  1. 03KV memory

    The cache formula at model scale, with a prediction and a new case

    previous
  2. 04Long-context use

    Storage, allowed pairs and layered reach on one eight-token example

    current
  3. AtlasRelated concept

    State-space hybrids: some attention layers replaced by layers with a fixed-size state

    related