Key and value tensor bytes for the equation shown, in decimal GB. Every number on this page is exact arithmetic you can redo by hand.
A question first, an equation second
The 4x KV-cache challenge
Two predictions, both checkable by hand. Everything here is exact arithmetic; nothing is a benchmark.
One question, before anything is revealed
Hold batch, layers, sequence length, head dimension, and precision fixed. Query heads stay at 32. Cached key/value heads fall from 32 to 8. What is the exact KV-cache ratio?
- Batch
- 1
- Layers
- 32
- Sequence length
- 8,192
- Head dimension
- 128
- Precision
- fp16/bf16
What this page is and is not
Nothing on this page is a benchmark. No latency, throughput, allocator overhead, paged-cache metadata, or model quality was measured to produce any figure above.
In this browser, under one key, and nowhere else. No account, no email, no analytics beyond whatever the rest of the site already does. Clearing site data removes them.