Transformer Systems Lab
Evaluation: try to break the claim
Take the speculative-decoding result into a test of one claim. Choose the slice you expect to fail first, see the overall and per-slice scores, then save the narrower claim. All numbers are toy values.
What this step asks
You understand a claim when you know what would break it.
The earlier steps set limits on memory, context, serving and speculative decoding. This step turns them into one scoped claim with a stated metric, slice, threshold and uncertainty, and the next counterexample to try.
Step 7 of 8
Evaluation and falsification
- 01Text to one update
- AtlasTokens and position
- AtlasAttention routing
- 02RoPE phase
- 03KV memory
- 04Long-context pressure
- 05Serving and decoding
- 06Speculative decoding
- 07Evaluation and falsification
- 08Capstone systems claim
From the speculative decoding step
A speed-up says nothing about quality until a test tries to break the claim.
The speculative decoding step separated two questions: does the output follow the target model, and does a round save work? This step turns a claim about the faster system into a test that could fail. The test names a task, a subset of test cases (a slice), a metric with a pass threshold, an uncertainty interval and the origin of the test data.
The claim under test
Faster served long-context behavior is acceptable
A claim can be tested only when it names the system under test, the slice, the metric and its threshold, where the test data came from, and which result would count against it.
Predict first
Which part of this claim will fail first?
Try it
Set up a test that could break the claim.
Illustrative scores from a fixed rule on this page; no model is run. Watch how the verdict changes with the slice, the perturbation, a threshold or the origin of the test data.
Outline of this lesson
claim + scenario + slice + metric + threshold + provenance + counterexampleone scoped claim
a benchmark score is not proof
which part fails first
slice, metric, threshold, provenance
overall score, slice score, interval, counterexample
a claim holds only within the scope it states
build one systems claim you can defend
Investigation · tasks and trials
Four tasks, sixteen rows: which signs may reverse?
A variant is compared with a baseline on four tasks, with four paired trials on each. The rows and their task labels fix the observed difference. How surprising that difference would be under a null hypothesis is a separate question: it depends on which units the null lets reverse independently, and the rows alone cannot show that.
Illustrative record, not a measured model or benchmark result. The controls change only this view; nothing is saved or sent.
Five distinctions used below
- Pair
- Each row pairs a baseline and a variant score from the same task and trial, and d = variant − baseline is taken inside the pair. Reordering or relabelling rows moves whole pairs; it cannot re-pair them.
- Target and record
- The target is what the claim is about: here the equal-task mean difference on these four fixed tasks. The rows are what was recorded. Neither says how other tasks would behave.
- Task and trial
- A task is a distinct problem; a trial is one attempt at it. More trials describe a task more fully. They are not more tasks.
- Averaging unit
- The equal-task mean T weights each task 1/J; the row mean weights each task by its number of rows. Equal counts, as here, guarantee that the two agree for any scores; with unequal counts the weights differ, although the values may happen to coincide.
- Null transformation
- Reversing a sign swaps the two method labels inside a unit, so d becomes −d. It is a hypothetical relabelling used to build a reference distribution, not a new run.
The record: 16 rows in 4 tasks
Each row is one (task, trial) pair; higher is better and d = variant − baseline. A sign of −1 swaps the two method labels for that row only. Baseline mean 0.25, variant mean 0.75, difference 0.5.
Task A
as recorded| Trial | Baseline → variant | d | Sign | Sign · d |
|---|---|---|---|---|
| 1 | 0 → 1 | +1 | +1 | |
| 2 | 0 → 1 | +1 | +1 | |
| 3 | 0 → 1 | +1 | +1 | |
| 4 | 0 → 1 | +1 | +1 | |
| Task mean mA | 1 | |||
Task B
as recorded| Trial | Baseline → variant | d | Sign | Sign · d |
|---|---|---|---|---|
| 1 | 0 → 1 | +1 | +1 | |
| 2 | 0 → 1 | +1 | +1 | |
| 3 | 0 → 1 | +1 | +1 | |
| 4 | 0 → 1 | +1 | +1 | |
| Task mean mB | 1 | |||
Task C
as recorded| Trial | Baseline → variant | d | Sign | Sign · d |
|---|---|---|---|---|
| 1 | 0 → 1 | +1 | +1 | |
| 2 | 0 → 1 | +1 | +1 | |
| 3 | 0 → 1 | +1 | +1 | |
| 4 | 0 → 1 | +1 | +1 | |
| Task mean mC | 1 | |||
Task D
as recorded| Trial | Baseline → variant | d | Sign | Sign · d |
|---|---|---|---|---|
| 1 | 1 → 0 | −1 | −1 | |
| 2 | 1 → 0 | −1 | −1 | |
| 3 | 1 → 0 | −1 | −1 | |
| 4 | 1 → 0 | −1 | −1 | |
| Task mean mD | −1 | |||
T(g·d) = (1/4) (1 + 1 + 1 − 1) = 0.5, at least the observed 0.5, and a whole-task pattern: the justified count includes it in the tail.
- One of the 16 whole-task patterns?
- Yes: the identity, nothing reversed.
- One of the 2 to the power 16 single-row patterns?
- Yes: every sign pattern is.
Count every pattern in each set
Each tile is one whole-task pattern; its small grid shows which tasks it reverses (dark), and tiles giving the same T share a column. Bars count single-row patterns by T; height is proportional to the count, and the smallest counts are drawn at a minimum visible height. Shaded: T ≥ 0.5, the inclusive upper tail. The outlined tile or bar is the pattern applied above.
Whole tasks reverse: 16 patterns (2 to the power 4)
Justified by the stated null
T ≥ 0.5 in 5 of 16: 4 ties (the identity among them) and 1 greater. Count 5/16 = 0.3125. The five values of T are not equally likely: they hold 1, 4, 6, 4, 1 patterns.
Single rows reverse: 65,536 patterns (2 to the power 16)
Hypothetical: this design does not supply it
T ≥ 0.5 in 2,517 of 65,536: 1,820 ties (the identity among them) and 697 greater. Count 2,517/65,536 ≈ 0.0384.
What this design justifies
The stated null lets whole tasks reverse. 5 of its 16 equally likely patterns give T ≥ 0.5, so the count is 5/16 = 0.3125. With 4 tasks no record can give less than 1/16. The single-row count, 2,517 of 65,536, would need a design stating that trials inside a task reverse separately; this one does not.
Neither count is the probability that the variant is better, a confidence interval, a bound on a useful effect or a release criterion.
Both counts as tables
| Reversed tasks | T | In the tail (T ≥ 0.5) |
|---|---|---|
| none (identity) | 0.5 | yes |
| A | 0 | no |
| B | 0 | no |
| A, B | −0.5 | no |
| C | 0 | no |
| A, C | −0.5 | no |
| B, C | −0.5 | no |
| A, B, C | −1 | no |
| D | 1 | yes |
| A, D | 0.5 | yes |
| B, D | 0.5 | yes |
| A, B, D | 0 | no |
| C, D | 0.5 | yes |
| A, C, D | 0 | no |
| B, C, D | 0 | no |
| A, B, C, D | −0.5 | no |
| T | Patterns | In the tail (T ≥ 0.5) |
|---|---|---|
| −1 | 1 | no |
| −0.875 | 16 | no |
| −0.75 | 120 | no |
| −0.625 | 560 | no |
| −0.5 | 1,820 | no |
| −0.375 | 4,368 | no |
| −0.25 | 8,008 | no |
| −0.125 | 11,440 | no |
| 0 | 12,870 | no |
| 0.125 | 11,440 | no |
| 0.25 | 8,008 | no |
| 0.375 | 4,368 | no |
| 0.5 | 1,820 | yes |
| 0.625 | 560 | yes |
| 0.75 | 120 | yes |
| 0.875 | 16 | yes |
| 1 | 1 | yes |
Why the two counts differ
Both counts use the same statistic and the same observed value, 0.5. They differ in the set of patterns the null hypothesis treats as equally likely. If whole tasks may reverse there are 16, and 5 give T ≥ 0.5: the record itself, three ties that reverse D together with one of A, B or C, and the reversal of D alone, which gives 1. If every row could reverse on its own there would be 65,536; splitting tasks produces many values near zero that the whole-task null never produces, so the observed value looks rarer: 2,517 of 65,536.
Neither count is chosen by which is smaller, and the two are not ordered in general. Here the whole-task count is larger; with two tasks whose differences are (2, −1) and (2, −1) it is smaller, 1/4 against 3/8. The design decides. This one says the tasks were fixed and whole tasks may reverse; it says nothing that lets trials inside a task reverse separately. The identical signs within each task could not establish that, and neither could the shared task labels.
What goes wrong otherwise can be counted. Suppose the whole-task null holds and the record is equally likely to be any of its 16 whole-task reversals. In 5 of them the single-row count is at most 2,517 of 65,536, so a rule that rejects when the counted fraction is 0.05 or less would reject in 5 of these 16 equally likely records, although the null holds for every one of them. The whole-task count never falls below 1/16, so the same rule never rejects with four tasks. (The 0.05 is used only to show the miscalibration; this page sets no threshold.)
Copies and fresh trials
Choose Rows copied. Every recorded row appears twice, the copy under a new trial id. The task means, T and the whole-task count are unchanged, because a copy reverses with its task. Counted as if all 32 rows could reverse independently, the tail becomes 15,033,173 of 2 to the power 32 (≈ 0.00350), about 11 times smaller, from no new observation. A copy cannot reverse apart from its source; in every possible outcome it equals its source. A fresh trial is different: a new run on the same task can show how variable that task is. It is still another trial of the same task, not another task.
The claim, revised
On these four fixed tasks, the variant’s equal-task mean difference is 0.5: three tasks favour it on every trial and one favours the baseline on every trial. Under the stated whole-task null, 5 of 16 reversals give a value at least this large. That is not evidence against the null, and it is not evidence of no difference or of equivalence: with four tasks, no record can give less than 1/16.
Which evidence comes next depends on the target. If the claim stays about these four tasks, more tasks would change the set that T averages rather than inform it. Fresh trials on the same four, meaning new runs rather than copies and collected under a stated repeated-trial design, can estimate each task’s expected difference and its variability more closely. They add no tasks, and the whole-task count still cannot fall below 1/16.
If the claim is meant to reach other tasks, that is a different target. First say which tasks it covers and how new tasks will be chosen from them. Tasks chosen under that design, not more trials of these four, are then the evidence, and the task remains the unit that reverses. Copies add nothing to either question.
Definitions
- dt,i = vt,i − bt,i
- The paired difference for task t and trial i: variant minus baseline, higher is better. Dt is task t’s list of differences.
- mt = (1/nt) Σi dt,i
- Task t’s mean difference over its nt rows (here nt = 4).
- T(d) = (1/J) Σt mt
- The target, the equal-task mean over J = 4 tasks: (1/4)(1 + 1 + 1 − 1) = 0.5. Equal nt guarantee that the row mean (1/N) Σ d equals it for any scores. Unequal counts weight the tasks differently, though the two values can still coincide: two tasks with 1 and 2 rows and every d = 1 give 1 for both.
- gεd = (ε1D1, …, εJDJ), ε ∈ {−1, +1}J
- A whole-task reversal: each sign multiplies a task’s entire list. These form a group of 2J elements, the identity included. Single-row reversals give each of the N rows its own sign: 2N elements. A reordering of tasks is in neither group; it leaves T unchanged and reverses nothing.
- p = #{ε : T(gεd) ≥ T(d)} / 2J
- The inclusive upper tail: every group element counts once, the identity and every tie included, even when several give the same value.
When the count is valid. The whole-task null says d and gεd have the same joint distribution for every ε. Under a null of that kind, a test that rejects when p ≤ α rejects with probability at most α. Ties and the discreteness of the count can make it conservative, so exact counting does not make the rejection rate equal α. Independent task lists, each distributed like its own reversal, are sufficient. They need not share a distribution, and the rows inside a task may depend on each other. Independence is not necessary either. Two weaker statements do not suffice. A zero mean is not symmetry: a difference of +1/2 with probability 2/3 and −1 otherwise has mean zero and is not symmetric. Symmetry of each task separately is not joint invariance: if all tasks share one fair sign, each task is symmetric, but reversing one task alone produces data that never occur.
The validity result is from Hemerik and Goeman (2018; arXiv:1411.7565v3, §2.1 Definition 1 and Theorem 1, and §2.2) and Koning and Hemerik (2024; arXiv:2202.00967v3, §2: Example 1 gives the sign-flipping group, their test is defined with the inclusive ≥ tail, and Theorem 1 bounds its size). Applying it to the whole-task group, and every number on this page, is this page’s own derivation and calculation, not a result reported by either paper.
Background, if needed: Probability basics and Random variables.
The same record in Python, counted exactly
from collections import Counter
from fractions import Fraction
# One paired row per (task, trial): (baseline, variant), higher is better.
record = {(task, trial): (0, 1) for task in "ABC" for trial in "1234"}
record |= {("D", trial): (1, 0) for trial in "1234"}
def differences_by_task(record):
"""d = variant - baseline, grouped by task id only, never by similar scores."""
tasks = {}
for (task, _trial), (baseline, variant) in record.items():
tasks.setdefault(task, []).append(variant - baseline)
return tasks
def parts(tasks, unit):
"""Each reversible unit's share of T = (1/J) * (sum of the task means)."""
J = len(tasks)
if unit == "task":
return [Fraction(sum(d), len(d)) / J for d in tasks.values()]
return [Fraction(x, len(d)) / J for d in tasks.values() for x in d]
def reference_counts(parts):
"""How many sign patterns give each value of T; 2 ** len(parts) patterns in all."""
counts = Counter({Fraction(0): 1})
for c in parts:
step = Counter()
for total, k in counts.items():
step[total + c] += k
step[total - c] += k
counts = step
return counts
def upper_tail(parts):
"""Inclusive: patterns with T(g d) >= T(d), the identity and every tie counted."""
counts, observed = reference_counts(parts), sum(parts)
at_least = sum(k for value, k in counts.items() if value >= observed)
return at_least, counts[observed], 2 ** len(parts)
def show(label, tail):
at_least, ties, total = tail
print(f"{label}: {at_least} of {total} patterns, {ties} of them ties")
tasks = differences_by_task(record)
T = sum(parts(tasks, "task"))
print("T =", T)
print({str(v): k for v, k in sorted(reference_counts(parts(tasks, "task")).items())})
show("whole tasks", upper_tail(parts(tasks, "task")))
show("single rows", upper_tail(parts(tasks, "row")))
# Copy every row under a new trial id: no new run, so no new information.
copied = record | {(task, trial + " copy"): pair for (task, trial), pair in record.items()}
copied_tasks = differences_by_task(copied)
print("copied: T =", sum(parts(copied_tasks, "task")))
show("copied, whole tasks", upper_tail(parts(copied_tasks, "task")))
show("copied, rows as if independent", upper_tail(parts(copied_tasks, "row")))
assert T == Fraction(1, 2) == sum(parts(copied_tasks, "task"))
assert upper_tail(parts(tasks, "task")) == upper_tail(parts(copied_tasks, "task")) == (5, 4, 16)
assert upper_tail(parts(tasks, "row")) == (2517, 1820, 2 ** 16)
assert upper_tail(parts(copied_tasks, "row"))[::2] == (15033173, 2 ** 32)
Standard library only. It builds the same sixteen pairs, groups them by task id, and counts both sets with exact fractions; its assertions check the counts shown in the figure. The page’s own figure uses a separate integer implementation.
Source reading · historical method
How τ-bench v1 kept tasks and trials apart
Two sources, each at a fixed version, read for their method: §3 of the paper (Yao et al., arXiv:2406.12045v1, 17 June 2024) and tau_bench/run.py at commit 59a200c. They are separate sources; the commit is not claimed to be the code of the paper’s first version. The README at that commit says the tasks are outdated and points to a newer benchmark, so this is a historical method, not a recommendation to run it.
- What it retains
- Each result stores its task id and trial number (lines 82–95), inside a loop over trials (58–64).
- What it averages
- A trial succeeds when its reward is within a small tolerance of 1, and n is the number of distinct trial numbers in the whole run (180–186). Successes c are counted per task id, turned into comb(c, k) / comb(n, k), and averaged over tasks with equal weight (188–199).
- What it does not justify
- The shared n assumes every task has the same complete, unduplicated set of trials. Reading the result as the paper’s pass^k, the chance that k future repetitions all succeed, also assumes the trials of a task are independent and identically distributed given the task: sufficient for that reading, not required for every joint-success target. The paper notes that its end-state reward can pass a run that broke a policy. It supplies no sign-reversal null and no test of this kind.
One task, four trials with outcomes (1, 1, 1, 0), k = 2. The six pairs of trials:
- {1, 2}: both succeed
- {1, 3}: both succeed
- {2, 3}: both succeed
- {1, 4}: one succeeds
- {2, 4}: one succeeds
- {3, 4}: one succeeds
All succeed in 3 of 6 pairs, comb(3, 2) / comb(4, 2) = 1/2; at least one succeeds in 6 of 6, which is 1. As a count over this record’s pairs, neither number needs independence. The source keeps task and trial identity; the sign-reversal null above is a separate assumption, not something it supplies.
Optional practice cued
Conditions: the figure and explanation stay in view, and the answer appears as soon as you choose. Choices last only while this page is open; nothing is recorded.
Under the stated whole-task null, does the count include this transformation of the record?
Reverse task C and task D, each as a whole.
Reverse trials 3 and 4 of task A, and nothing else.
Swap the positions of task B and task D in the list.
Where this step sits
After evaluation, a headline becomes a claim with stated limits.
You now have enough to say what held, what failed, what is still uncertain, and what the final systems claim may say.
- 05Serving and decoding
time to first token, time per output token, cache reads, sampling
previous - 06Speculative decoding
speed-up from a draft model, and whether the output is unchanged
previous - 07Evaluation and falsification
claim, slice, metric, counterexample
current - 08Capstone systems claim
assumptions, interventions, evidence, what would refute the claim
next
Saving your work
Save the narrower claim, then take it to the capstone.
Saving keeps the claim, your first prediction, the part that failed, the slice, the counterexample and the uncertainty, in this browser only. The capstone step starts from this claim.
Continue to Capstone systems claim