Research

Turn one question into a test someone else can inspect.

Learned or varied attention? Separate the calculation you can explain from the question you still need to test. Start with a worked framing, then open a separate evidence-free local draft.Frame a question from attention Available here: page-memory drafting and manual JSON export, not study execution or independent review.Inspect the current research state

Scientific state
Mapped · evidence none
Executable work
P0 feasibility only
Independent review
None
Participant study
Not authorized
Publication
Not authorized

Current conditional objective

SBJ-001 is the sole conditional scientific objective, not an admitted study. Its only executable work is the model-free SBJ-001-P0 feasibility preflight; no protocol is locked, no participant study is executing, and no scientific evidence or result exists.

Mapped · P0 feasibility only · as of 2026-08-16

Source-bound judgment · SBJ-001

Among predeclared, domain-qualified reviewers making bounded promote / hold / reject decisions on independently adjudicated AI-generated ML or safety evidence, and among case families where a matched full derangement is constructible, what is the causal decision cost of complete claim-to-source misbinding when mechanically valid claim-level exact-source-span pointers are compared with the same pointers reassigned by a fixed-point-free, label-blind, within-case full derangement, while report assertions, artifact and source bytes, navigation shell, and review budget are fixed?

Current claim boundary: The objective is conditional on case families where a matched full derangement is constructible. It cannot claim a generic provenance-interface benefit, citation-validity result, behavioral-verification result, expert-use effect, or platform efficacy.

P0 failure boundary: The P0 must return KILL or REVISE if the direct-prior, exact-source, rights, treatment-invariance, ethics, reviewer-route, endpoint, power, or resource gates cannot be closed without acquiring study materials or involving participants.

Scientific state: mapped / none / none / none; no study is executing.

Exact next gate: Complete only the model-free SBJ-001-P0 feasibility record and obtain a disposition from a named human who did not design or implement the preflight. Even ELIGIBLE-FOR-SEPARATE-PROTOCOL-ADMISSION would not admit a study.

How this record fits the project

Continuous Function is an independent research project for questions that can be made explicit, challenged, and tested. The public record separates a possible direction from an admitted study, a protocol from a run, and a source from empirical evidence.

Research thesis: What is the causal decision cost of complete claim-to-source misbinding when valid claim-level exact-source-span pointers are replaced by a matched full derangement while the report, artifact and source bytes, navigation, and review budget stay fixed?

What open means here: Open here means inspectable and criticizable: questions, assumptions, falsifiers, evidence states, and next gates are visible. It does not mean public model execution, automatic research intake, or that every suggested idea becomes active work.

Learn → Question → Assumptions → Local draft

What question does your attention change leave open?

Separate the mechanism you can explain from the claim you still need to test.

Worked example · deterministic toy calculation

Hold one attention row fixed: query q = [1], keys kA = kB = [0], and scalar values vA = 2 and vB = 8. The key dimension is 1, so both scaled dot-product scores are 0. Change only which sources are allowed.

One row, two access conditions
Allowed sourcesWeights (A, B)Output
A and B(0.5, 0.5)5
A only; B masked(1, 0)2
Check why the output changes

Each allowed score contributes exp(0) = 1 to normalization. With both sources, each weight is 1 / (1 + 1) = 0.5, giving 0.5 × 2 + 0.5 × 8 = 5.

Exclude B before normalization: A gets weight 1 and B gets exactly 0, giving 1 × 2 + 0 × 8 = 2. Setting B’s already-zero score to zero would not mask it. At least one source must remain allowed.

Toy result, not research evidence. The mixture changed from 5 to 2. No model was run, and this does not show that an answer became better.

Question · still open
For one fixed model and task, does masking a distractor improve answer accuracy without harming cases that need that source?
Assumptions to state
Define a distractor before inspecting answers. Fix the model, inputs, mask location, and decoding; name the task and the limits of the claim.
Discriminating test · proposed, not run
Plan matched masked/unmasked cases and a control where B is needed. Predeclare accuracy, a useful-change threshold, stopping and invalidity rules. Worse or inconclusive outcomes must remain possible; a lower toy output is not higher accuracy.

Continue in the existing local workbench

Start blank, or inspect the existing Evaluation and Falsification toy starter. That starter is a different worked plan, not this attention example. Neither link transfers your question, values, mask, or lesson state.

Write the question in the workbench; edits live only in that page’s memory. Export JSON before leaving or refreshing. Drafting runs no study, creates no review request, locks no protocol, admits no evidence, approves no claim, and publishes nothing. Human review happens elsewhere; model advice is advisory, never independent named-human review.

One question · five honest capability phases

Research is a chain of decisions, not a run button.

Keep the question visible while the instrument changes. Every stage leaves an inspectable artifact and ends at a human gate; negative, inconclusive, and invalid outcomes stay in the record.

02 · Available here

Commit

What comparison, held-fixed conditions, measurement, falsifier, and decision rule will you state before seeing an outcome?

Draft the study
Artifact you leave
A versioned local study draft and deterministic JSON representation.
Human gate
Approve the draft for independent methods criticism. This site cannot approve or lock it.
Later bounded help
Completeness checks, alternative controls, and protocol criticism may help; approval and lock remain human decisions.
Next handoff
Export the draft for independent methods review and separately authorized execution.
Capability boundary
The browser-local workbench drafts and format-checks the packet. It uploads nothing and creates no evidence.
Available here Partial instrument External today Future capability

Not live: accounts, canonical ledgers, task queues, model runners, agent orchestration, protocol locks, or publication connectors.

Protocol starters · evidence none

Choose the shape of the first question.

Each starter is editable, browser-local, and deliberately empty of evidence. It gives you a stronger first draft, not an admitted study.

Build from an empty question

Blank falsifiable study

Name one comparison, one held-fixed boundary, one falsifier, and the smallest test that could change a decision.

Protocol starter · Evidence none · No run performed

Use this starter
Behavioral alignment

Behavior change, not wording

Compare action-grounded behavior against matched language-only and sham controls while protecting collateral utility.

Protocol starter · Evidence none · No run performed

Use this starter
Multi-principal safety

Permissioned handoff integrity

Test bilateral permissions and content-addressed lineage against metadata-only logging, then measure bad-artifact acceptance and useful completion.

Protocol starter · Evidence none · No run performed

Use this starter
Reproducibility

Reproduce a claimed effect

Specify an exact independent reproduction against the released environment, primary output, tolerance, and decision rule.

Protocol starter · Evidence none · No run performed

Use this starter

How to read the state

Work state, evidence disposition, and validity scope are separate. Completing a study does not make its hypothesis true; a clean contradiction or inconclusive result is as visible as support. Sources are provenance, never evidence.

FieldValue
Source-mapped · not admittedA topic and exact admission gate are recorded. No bounded study, protocol, run, or result exists.
Protocol draftA bounded draft record exists. Current authorization and sequencing must be stated separately; no model call or result claim is allowed.
Preregistered · lockedA named version is frozen and dated before eligible evidence collection. Locking does not imply execution or support.
Evidence availableActual artifacts receive a bounded supported, contradicted, inconclusive, or invalid disposition. A source map never qualifies.

Source-mapped, not admitted

These are possible future directions, not active programs. Each needs the stated human admission decision before protocol work begins.

Behavioral Specification and Human ControlSource-mapped · not admittedExplicit specifications, adaptive-data policies, and human-AI interfaces under a named context shift. Missing decision: Select one specification, context-shift population, comparator, machine-behavior outcome, and human-control decision outcome that can be studied offline.
Multi-Agent Systems SafetySource-mapped · proposal notePopulation-level coordination, failure, attribution, and control in realistic multi-principal agent networks. Missing decision: Identify an existing safe, reproducible testbed or trace corpus with distinct principals, system-level units, a non-single-agent threat model, and an external-validity map.
Public Evidence and Technical GovernanceSource-mapped · not admittedEvidence that could change one named risk-management, audit, governance, or public decision. Missing decision: Name the decision-maker, decision, uncertainty to reduce, lawful fixed corpus, analysis method, conflicts, and causal-versus-descriptive boundary.

Methods and prior work

The records below provide methods depth and learning context. They are not additional active programs in the current research agenda.

Atlas-64

Attention serving systems

Methods track · no benchmark claim

A public learning and methods track for cache, batching, and serving tradeoffs. It does not claim production superiority or a completed institutional evaluation.

FieldValue
IdentifierATTN-SERVING
Work stateConcept and protocol development
Evidence dispositionNone at program level; no model result is published here.
Exact next gateAttach a governed experiment and admissible evidence before promoting any performance claim.