Verify · under lens

Your eval went up. What does this paired comparison support?

Declare what the rows represent, set a meaningful threshold, inspect the method-specific item analysis, and then test repeatability across independent matched runs.

Browser-local calculation · no application upload or save action · no account · no receipt

Claim under test

The variant's submitted arithmetic mean was better in the declared direction by more than the predeclared useful threshold.

Turn

Turn the instrument

Paste paired additive row contributions from one predeclared comparison.

Which direction means better?
What do these evaluation units represent?
Data, pairing, and supported methods

Put each stable case ID, baseline value, and variant value on one row. Tabs, commas, semicolons, or spaces are accepted. A header and the observed row status are optional.

Both finite scores are required. Missing, invalid, duplicate, or errored rows are rejected rather than dropped, imputed, or realigned. Exact ties are retained.

Exact 0/1 pairs use McNemar plus Newcombe method 10 for an IID target. Other additive row values use a paired Student-t interval only when its declared assumptions are established.

Clustered, weighted, repeated-generation, fold-dependent, sequential, and nonadditive metric designs are deliberately unsupported.

Read

Read the item analysis

The calculation is conditional on the scope and threshold you declared.

No reading yet

Declare direction, exact scope, metric unit, and a positive useful threshold; then paste keyed paired rows and analyze.

Test

Test repeatability

Eligibility first requires at least 5 independent matched run pairs from one declared rerun-generating process. Conditional on eligibility, the one-sided 95% exact lower bound must exceed 50%. The minimum-count rule and statistical bound are separate.

Is the matched-run probability model established?

Missing, invalid, duplicate, or errored run pairs are rejected rather than dropped or realigned. Exact threshold equality counts as not above threshold. The estimand is p = P(direction-adjusted run-pair difference > useful threshold) under the declared IID matched-run process.

No matched reruns yet. An item analysis cannot establish run-to-run repeatability.

Mark

Mark the design boundaries

0 yes · 0 no · 0 unsure · 7 unanswered. These are context declarations; the browser cannot verify them.

Was evaluation overlap with training or fine-tuning data ruled out?
Why this matters

Contamination can inflate an evaluation result. Pasted scores cannot reveal overlap or verify a decontamination procedure.

Did you check that the input alone cannot predict the label?
Why this matters

A reproducible gain can still measure input-only label leakage instead of the intended mechanism. The browser cannot perform this check.

Before viewing any baseline–variant difference, were the target, metric and direction, threshold, analysis method, comparison, run-pair plan and count, stopping rule, and configuration fixed?
Why this matters

Selecting a winner after searching comparisons changes the interpretation. This small tool does not correct multiple comparisons, optional stopping, or winner selection.

Were the target scope and useful-effect threshold chosen before inspecting this result?
Why this matters

Answer yes, no, or unsure. This is an unverified declaration, not proof of preregistration or precommitment. A no or unsure answer withholds a threshold-clearing reading.

Were baseline and variant evaluated with identical prompts, decoding, and scoring?
Why this matters

If the harness also changed, the submitted difference describes a combined model-plus-harness intervention rather than the model change alone.

Does each row represent one genuinely independent evaluation unit?
Why this matters

Repeated prompts, users, documents, generations, folds, or time-adjacent observations can be clustered. This tool does not implement cluster-aware or hierarchical methods.

Is the target metric exactly the arithmetic mean of these row values?
Why this matters

Corpus BLEU, pooled F1, AUC, perplexity, ratios, percentiles, unequal weights, and other nonadditive metrics require a metric-specific estimator.

Read

Bound the reading

Design declarations, item analysis, and matched-run repeatability remain separate all the way to the conclusion.

Complete the instrument and analyze the submitted rows to get a reading.

What this reading cannot support

This harness does not run your model, inspect your data, detect contamination or leakage, correct selection, or validate your declarations.

A bounded result is conditional on this submitted comparison. It does not establish causality, validity on other distributions, publication readiness, or production impact.

Reloading resets this application state. This version exposes no upload, save, export, publish, or receipt action. The automated privacy check covers only its documented Chromium flow and API set.