Verify · under lens

Your eval went up. Did your change do that?

Does the paired difference survive item resampling, repeat across matched reruns, and remain interpretable under the design you actually used?

Browser-local · no upload · no account · no saved record · no receipt

Claim under test

The variant improved the selected metric on these submitted evaluation units.

Turn

Turn the instrument

Paste paired row contributions from one pre-committed comparison.

Which direction means better?
Does each row represent one independent evaluation unit?
Why this matters

Repeated samples from the same prompt, user, or document are clustered rather than independent. Treating every repetition as a new item can make the interval look much tighter than the evidence warrants; aggregate by the independent unit or use a cluster-aware method instead.

Data and pairing

Use one score per evaluation item, in the same item order for both systems: 1/0 for correctness, or a continuous row contribution whose arithmetic mean is the metric you intend to compare.

Two summary metrics do not reveal which items changed. This calculation resamples paired rows. Repeated generations from one prompt are not independent rows; aggregate them or use a cluster-aware method.

Method settings · 10,000 resamples
10,000

More resamples refine this deterministic estimate; they do not add evidence. The browser stops above ten million item draws.

Read

Read the submitted evidence

The visual is descriptive. The marks keep design and repeatability separate.

No reading yet

Declare the metric direction, paste matching rows, mark the independent-unit boundary, and analyze.

Test

Test repeatability

At least 3 matched baseline/variant reruns are required before the reading can call the signal repeated.

No matched reruns yet. A positive row calculation remains an item-level signal only.

Mark

Mark the design boundaries

0 declared yes · 0 no · 0 unsure · 6 unanswered. These are your declarations; the browser cannot verify them.

Are you certain no evaluation item appeared in training or fine-tuning data?
Why this matters

Contamination is the dominant failure in open-weight evaluation, and decontaminating a benchmark alters the benchmark, so the usual remedy does not restore it. Nothing in a score vector reveals overlap.

Have you checked that the input alone cannot predict the label?
Why this matters

A gain can be real, reproducible, and still measure something else. In this project’s own Atlas-64 program a five-point improvement survived independent recomputation and then collapsed: the prompt alone recovered the hidden mode 95.9% of the time.

Is this the only configuration you are reporting, rather than the best of many?
Why this matters

Trying many configurations and reporting only the winner selects for noise. Resampling the winning rows does not correct that selection; this reading assumes one pre-committed comparison.

Were baseline and variant evaluated with identical prompts, decoding, and scoring?
Why this matters

A changed prompt template or sampling temperature between runs makes the comparison measure the harness rather than the model.

Is the reported metric exactly the arithmetic mean of these row values?
Why this matters

This harness averages row-level contributions. Per-example perplexities, corpus F1 or BLEU, percentiles, and unequal-token-weight losses do not generally average back to the reported metric; use the additive contribution and independent unit your estimand requires, or a method designed for that aggregate.

Read

Bound the reading

Design, item evidence, and reruns remain separate all the way to the conclusion.

Choose the metric direction and analyze the submitted rows to get a reading.

What this reading cannot support

This harness does not run your model, see your data, or check your training set. It tests one pre-committed comparison on one evaluation set under one harness.

Even a favorable reading is not proof of causality, capability on other distributions, publication readiness, or production impact.

Browser-local Reloading the page discards every value.