Fieldwork ask → train → look inside → change one thing, on a synthetic three-feature task

No run yet.

Type a question. Recognized settings are underlined; every other word is kept as a note and configures nothing.

What the question box understands

Recipes ordinary60 paired60 original120 longer150 logistic60. Settings steps=60 or 60 steps (1–500), seed=17 (init seed), lr=0.03 (0.001–0.1), ρ=0.95 or rho=0.95 (multiples of 0.005), words paired ordinary mlp logistic. Conflicting or out-of-range values are rejected. All other words are kept as your note and change nothing. Questions and open questions are limited to 400 characters.

Comparing 400 examples × 120 steps with 800 examples × 60 steps matches how many times examples are seen in all; it does not match the number of optimizer steps or the CPU time.

Decision field p(y=1) over a × b at one binary cue slice (−1 or +1) · label: same sign, a·b > 0

no run yet

No model yet. After a run, the field shades a finite grid: each cell is coloured by one forward pass of the selected checkpoint at the cell centre, with the cue fixed at the chosen slice (−1 or +1).

+1
Points

reliance band: no run

Reliance band: the cells of the supported grid (a finite set) where the predicted class changes when the cue goes from −1 to +1. No such cells does not mean the cue has no effect on the probability, and a band of any size is not a complete causal explanation.

Arrows: nearest example · Enter: dissect · F: flip cue · [ ]: step · Space: replay · Home/End: first/last checkpoint

Field and points as a table

Specimen

select a point

Three numbers, not a picture. Shading marks the gap |v| < 0.2, where no examples are generated. A dashed ring (or *) marks the example’s stored value before you moved it.

Move this example · the model stays fixed — — —

Generated examples have |a| ≥ 0.2 and |b| ≥ 0.2, so a and b snap out of the central gap (−0.2, 0.2). The cue is only −1 or +1. Values outside that support are not evaluated.

Contributions as a table

Runs over training

step —

solid: IID accuracy · dashed: reversed-cue accuracy · dotted: band cells / supported grid cells (class disagreement between cue −1 and +1)

The next run that would tell these apart

    400 × 120 vs 800 × 60 matches how many times examples are seen, not optimizer steps or CPU time.

    Run histories as a table

    Keep this investigation and come back to it

    This browser keeps one saved copy for this site address, and only when you choose to save. Work opened at another address (another port) is kept apart: export its JSON there, then choose Import JSON here.

    Checking the saved copy…

    Copy the saved JSON

    These are the exact stored bytes. Its hashes stay unchecked until rebuilding it here succeeds.

    Copy text you have not added yet

    This text is not a finished investigation you can reopen. Add your note and pin your next question before saving; unfinished runs stay in this tab only. Export JSON keeps finished work.

    Rebuilding reruns the same code and fixed seeds: it checks the arithmetic, not the finding, and is not an independent experiment. Once you have looked at the diagnostic rows, they are evidence you used to choose between runs, not a final unseen test. Flipping the cue keeps a, b and the captured weights fixed; moving a or b recomputes the label from a·b. Fieldwork tabs take turns to save under one shared lock, check that the saved copy is exactly what they last saw, and read it back after writing. Anything else writing to this browser’s storage is outside that arrangement, and a copied hash proves nothing by itself.