A question first, an equation second

The 4x KV-cache challenge

Two predictions, both checkable by hand. Everything here is exact arithmetic; nothing is a benchmark.

prediction firstbrowser-local answersno account

One question, before anything is revealed

Hold batch, layers, sequence length, head dimension, and precision fixed. Query heads stay at 32. Cached key/value heads fall from 32 to 8. What is the exact KV-cache ratio?

Batch
1
Layers
32
Sequence length
8,192
Head dimension
128
Precision
fp16/bf16
Your prediction

The MHA cache is how many times the size of the new grouped-query cache?

x smaller

A wrong number is more useful than no number: it tells you which factor you were tracking. The answer stays in this browser only, and nothing is sent anywhere.

What this page is and is not

Calculated here

Key and value tensor bytes for the equation shown, in decimal GB. Every number on this page is exact arithmetic you can redo by hand.

Not measured here

Nothing on this page is a benchmark. No latency, throughput, allocator overhead, paged-cache metadata, or model quality was measured to produce any figure above.

Where your answers live

In this browser, under one key, and nowhere else. No account, no email, no analytics beyond whatever the rest of the site already does. Clearing site data removes them.