Refusal language versus policy action
AE-001Source records explain where a question was encountered. They do not name, sponsor, govern, or empirically support Continuous Function research.
Under one fixed small-model family and a harmless synthetic policy environment, does action-grounded fine-tuning reduce prohibited structured actions more than surface-language fine-tuning, without producing blanket refusal or materially degrading allowed-task success?
Selected for protocol development. No protocol is locked and no run is authorized.
Current claim boundary
Only a synthetic study question and candidate design are admitted. A machine-readable action is generated output, not an executed action. AE-001 cannot establish deception, underlying goals, honesty, real-world safety, tool behavior, or generality beyond any model and task populations later locked.
What would count against it
The surface-language arm is non-inferior to the action-grounded arm on prohibited-action rate within the locked margin across held-out task families, and the message/action dissociation is smaller than the locked useful-effect margin. Blanket refusal or capability collapse cannot count as support.
No empirical evidence available.
- Work state
- Protocol-formingOne bounded question has been admitted for protocol design. Its measures, exclusions, analysis, and stopping rules are still being fixed. Nothing has been preregistered or run; no empirical evidence is available.
- Evidence disposition
- noneNo empirical evidence available.
- Validity scope
- noneNo empirical validity scope has been earned.
Latest update · 2026-08-07 · Question admissionThree independent design reviews returned REFINE. Their shared corrections selected one question, separated source provenance from evidence, and blocked execution until a complete protocol is independently approved and locked.
01Human authority and execution boundary
Human authority
The research lead retains authority over the question, constructs, assumptions, predictions, evidence standards, exclusions, stopping, interpretation, publication, and disposition.
Agent role and execution
None in the current state. Future automation may execute only a frozen procedure and may not change the question, thresholds, exclusions, stopping rule, evidence admission, interpretation, or publication.
No model, evaluator, training, fine-tuning, or experimental call is authorized. Continuous Function does not observe or verify runs; any future execution must occur outside the static site under a separately locked and authorized protocol.
02Alternatives considered
AE-Q2hypothesis-readyCan a pretraining-time module selectively suppress and later restore a named capability under a fixed recovery threat model?
Candidate question · not admitted under AE-001. Pretraining, checkpoint rights, recovery attacks, compute, and misuse review remain separate gates.
AE-Q3too-immatureWhen do model-generated reports add validated information about an independently known internal variable?
Candidate question · not admitted under AE-001. The reference state, independent target, calibration rule, and anti-anthropomorphic boundary are not fixed.
AE-Q4hypothesis-readyHow stable and controllable is steering resistance across interventions and task changes?
Candidate question · not admitted under AE-001. The correction mechanism, perturbation population, matched ablations, and transfer criterion are not fixed.
03Source provenance — not evidence
These records explain where the question was encountered. They do not sponsor, govern, or empirically support this study.
AE-S1scientific preprintRethinking harmless refusals when fine-tuning foundation models
- Organization
- arXiv author manuscript
- Accessed
- 2026-08-07
- Motivated
- Motivates a narrower test of whether surface-language and later output channels differ under distinct fine-tuning strategies.
- Does not establish
- It does not validate AE-001’s synthetic construct, endpoint, margins, or claim. Its deception terminology is not imported.
- Rights use
- Deep link and bibliographic metadata only; no paper text or assets copied.
AE-S2foundation research portfolioAI Alignment Foundation research portfolio
- Organization
- AI Alignment Foundation
- Accessed
- 2026-08-07
- Motivated
- Records where behavior versus surface compliance and the surrounding alignment questions were encountered.
- Does not establish
- Portfolio summaries are not evidence, do not establish replication, and do not set Continuous Function’s agenda.
- Rights use
- Deep link and short factual topic labels only.
AE-S3foundation project summaryReason-Based Deception: How Refusal Training Creates a False Sense of Safety
- Organization
- AI Alignment Foundation
- Accessed
- 2026-08-07
- Motivated
- Topic-discovery provenance for the first bounded question.
- Does not establish
- The page’s promotional conclusions are not adopted as Continuous Function claims or evidence.
- Rights use
- Deep link and title only.
04Candidate hypothesis and decision rule
AE-H1protocol-readyOn held-out policy/task combinations, surface-language fine-tuning will increase policy-conforming refusal language over sham control, while action-grounded fine-tuning will reduce prohibited structured actions over surface-language fine-tuning by a prespecified useful margin and preserve allowed-task success within a prespecified collateral-degradation margin.
- Estimand
- The paired, run-level difference in prohibited-action rate between action-grounded and surface-language arms across independently trained matched seed blocks.
- Primary outcome
- Prohibited-action rate in the machine-readable action channel.
- Decision rule
- Not fixed. Numeric useful-effect, non-inferiority, utility, uncertainty, and invalidation boundaries must be justified and locked before any call.
- Falsifier
- The surface-language arm is non-inferior to the action-grounded arm on prohibited-action rate within the locked margin across held-out task families, and the message/action dissociation is smaller than the locked useful-effect margin. Blanket refusal or capability collapse cannot count as support.
- Claim boundary
- At most, the exact training recipes can be compared on the locked synthetic policy environment, model family, held-out task families, and scorer.
- Multiplicity family
- One primary outcome and one primary contrast; secondary family unresolved.
- Locked at
- Not locked
05Protocol draft and governed objects
- research questionadmitted for protocol designAE-001 selected question
AE-Q1 - protocol draftunlocked · execution blockedAE-001 protocol v0.1 draft
AE-001-P0.1
AE-001-P0.1draftMatched three-arm synthetic action-selection study
One surface-language arm, one action-grounded arm, and one sham-control arm begin from an exact shared checkpoint with matched adaptation budgets. Confirmatory items are held out by rule family and task template. The design remains incomplete and cannot authorize execution.
- Version / lock
- 0.1-draft · not locked
- Protocol hash
- Not locked
- Review
- refine
- Assignment unit
- One training run within a matched seed block.
- Analysis unit
- Independently trained seed blocks paired across arms.
- Arms
- Surface-language supervision without action-target supervision; Action-grounded structured-action supervision; Matched benign sham control
- Randomization
- Matched seed blocks are proposed; exact seed generation, data order, and randomization procedure are unresolved.
- Independent runs
- Unresolved. The final count must follow a locked run-level precision target; prompts are repeated measures, not replications.
- Model hashes
- Not fixed
- Data hashes
- Not fixed
- Harness hash
- Not fixed
- Primary outcome
- Prohibited-action rate in the structured action channel.
- Primary contrast
- Action-grounded minus surface-language fine-tuning.
- Decision margins
- Unresolved: useful action effect, language manipulation, collateral utility, and non-inferiority margins.
- Analysis
- Unresolved: paired run-level estimator, uncertainty interval, missing/failed-run treatment, and sensitivity analyses.
- Multiplicity
- One primary contrast is fixed in principle; the secondary-family correction remains unresolved.
- Stopping
- Unresolved. Confirmatory effects may not be inspected to adapt sample size or replace failed runs.
- Contamination
- Hold out rule families and task templates; unknown pretraining contamination remains unknown. Generator and split hashes are unresolved.
- Pilot separation
- Pilot, validation, and confirmatory generators, prompts, and artifacts must use distinct frozen namespaces and hashes.
- Deviations
- None recorded; protocol is not locked
- Model calls
- Not authorized
06Evidence ledger
No empirical evidence available.
07Decisions and blocked claims
Recorded decisions
AE-D1· 2026-08-07Admit AE-Q1 as the sole question under protocol development.It has the closest fit to paired comparisons, a low-compute synthetic design, an independent action-channel endpoint, and an explicit failure boundary. The decision admits protocol labor only.
Project admission decision, narrowed through independent design review.AE-D2· 2026-08-07Keep AE-Q2, AE-Q3, and AE-Q4 unadmitted under AE-001.They require different constructs, units, interventions, threat models, and resource gates. A four-question portfolio would conceal that difference.
Research-governance adjudication.
Blocked claims
- AE-001 has been preregistered or executed.
The protocol is an unlocked draft and no model call is authorized.
Release gate: A complete protocol must pass independent methods review and receive a public hash and lock date. - Refusal language differs from policy action in a model.
No empirical evidence exists.
Release gate: Admissible artifacts must close under the locked protocol and receive a bounded disposition. - AE-001 measures deception, intent, or real-world safety.
Those constructs are outside the synthetic output-channel design.
Release gate: Not releasable from AE-001, regardless of outcome.
08Publication boundary and change log
Source records explain where questions were encountered. They are not evidence, sponsors, endorsements, or agenda authorities.
Publication ceiling: evidence none; validity none.
Allowed to say
- AE-001 is one selected question under protocol development.
- The protocol is not locked, no execution is authorized, and no empirical evidence exists.
- The other scholarship-derived questions and agenda families remain source-mapped and unadmitted.
Withheld
- Any result, effect, benchmark, alignment, deception, safety, or generalization claim.
- Any claim that Continuous Function currently runs agents or experiments.
- Any sponsor, endorsement, grant-award, or access implication.
Change log
- 2026-08-07
Created the protocol-forming admission record after three design reviews converged on one active question, three state axes, run-level units, negative-result parity, and a fail-closed execution gate.