Refusal language versus policy action
AE-001Source records explain where a question was encountered. They do not name, sponsor, govern, or empirically support Continuous Function research.
Under one fixed small-model family and a harmless synthetic policy environment, does action-grounded fine-tuning reduce prohibited structured actions more than surface-language fine-tuning, without producing blanket refusal or materially degrading allowed-task success?
Historically selected for protocol development. Its dated sequencing is superseded, the draft is parked, and no further protocol work or run is authorized.
Current claim boundary
AE-001 is a retained, parked conditional draft outside current sequencing. Its earlier primary contrast cannot be advanced without a separately admitted surface-policy-versus-active-sham reframe with independently scored held-out state transitions. It cannot establish deception, underlying goals, honesty, real-world safety, tool behavior, or generality.
What would count against it
The surface-language arm is non-inferior to the action-grounded arm on prohibited-action rate within the locked margin across held-out task families, and the message/action dissociation is smaller than the locked useful-effect margin. Blanket refusal or capability collapse cannot count as support.
No empirical evidence available.
- Work state
- Protocol draftA bounded protocol draft is recorded. Its record must state separately whether work is currently admitted, parked, or superseded. Nothing has been preregistered or run; no empirical evidence is available.
- Evidence disposition
- noneNo empirical evidence available.
- Validity scope
- noneNo empirical validity scope has been earned.
Latest update · 2026-08-16 · Sequencing supersededThe current Product Goal superseded AE-001’s dated sequencing, retained it only as a parked conditional draft, and made SBJ-001 the sole conditional scientific objective. No AE-001 protocol work, review, or execution is active.
01Human authority and execution boundary
Human authority
The research lead retains authority over the question, constructs, assumptions, predictions, evidence standards, exclusions, stopping, interpretation, publication, and disposition.
Agent role and execution
None in the current state. Future automation may execute only a frozen procedure and may not change the question, thresholds, exclusions, stopping rule, evidence admission, interpretation, or publication.
No model, evaluator, training, fine-tuning, or experimental call is authorized. Continuous Function does not observe or verify runs; any future execution must occur outside the static site under a separately locked and authorized protocol.
02Alternatives considered
AE-Q2hypothesis-readyCan a pretraining-time module selectively suppress and later restore a named capability under a fixed recovery threat model?
Candidate question · not admitted under AE-001. Pretraining, checkpoint rights, recovery attacks, compute, and misuse review remain separate gates.
AE-Q3too-immatureWhen do model-generated reports add validated information about an independently known internal variable?
Candidate question · not admitted under AE-001. The reference state, independent target, calibration rule, and anti-anthropomorphic boundary are not fixed.
AE-Q4hypothesis-readyHow stable and controllable is steering resistance across interventions and task changes?
Candidate question · not admitted under AE-001. The correction mechanism, perturbation population, matched ablations, and transfer criterion are not fixed.
03Source provenance — not evidence
These records explain where the question was encountered. They do not sponsor, govern, or empirically support this study.
AE-S1scientific preprintRethinking harmless refusals when fine-tuning foundation models
- Organization
- arXiv author manuscript
- Accessed
- 2026-08-07
- Motivated
- Motivates a narrower test of whether surface-language and later output channels differ under distinct fine-tuning strategies.
- Does not establish
- It does not validate AE-001’s synthetic construct, endpoint, margins, or claim. Its deception terminology is not imported.
- Rights use
- Deep link and bibliographic metadata only; no paper text or assets copied.
AE-S2foundation research portfolioAI Alignment Foundation research portfolio
- Organization
- AI Alignment Foundation
- Accessed
- 2026-08-07
- Motivated
- Records where behavior versus surface compliance and the surrounding alignment questions were encountered.
- Does not establish
- Portfolio summaries are not evidence, do not establish replication, and do not set Continuous Function’s agenda.
- Rights use
- Deep link and short factual topic labels only.
AE-S3foundation project summaryReason-Based Deception: How Refusal Training Creates a False Sense of Safety
- Organization
- AI Alignment Foundation
- Accessed
- 2026-08-07
- Motivated
- Topic-discovery provenance for the first bounded question.
- Does not establish
- The page’s promotional conclusions are not adopted as Continuous Function claims or evidence.
- Rights use
- Deep link and title only.
04Candidate hypothesis and decision rule
AE-H1protocol-readyOn held-out policy/task combinations, surface-language fine-tuning will increase policy-conforming refusal language over sham control, while action-grounded fine-tuning will reduce prohibited structured actions over surface-language fine-tuning by a prespecified useful margin and preserve allowed-task success within a prespecified collateral-degradation margin.
- Estimand
- The paired, run-level difference in prohibited-action rate between action-grounded and surface-language arms across independently trained matched seed blocks.
- Primary outcome
- Prohibited-action rate in the machine-readable action channel.
- Decision rule
- Not fixed. Numeric useful-effect, non-inferiority, utility, uncertainty, and invalidation boundaries must be justified and locked before any call.
- Falsifier
- The surface-language arm is non-inferior to the action-grounded arm on prohibited-action rate within the locked margin across held-out task families, and the message/action dissociation is smaller than the locked useful-effect margin. Blanket refusal or capability collapse cannot count as support.
- Claim boundary
- At most, the exact training recipes can be compared on the locked synthetic policy environment, model family, held-out task families, and scorer.
- Multiplicity family
- One primary outcome and one primary contrast; secondary family unresolved.
- Locked at
- Not locked
05Protocol draft and governed objects
- research questionhistorically selected · current sequencing supersededAE-001 selected question
AE-Q1 - protocol draftretained parked draft · unlocked · execution blockedAE-001 protocol v0.1 draft
AE-001-P0.1
AE-001-P0.1draftMatched three-arm synthetic action-selection study
This retained historical draft proposed surface-language, action-grounded, and sham-control arms from an exact shared checkpoint with matched adaptation budgets. Its dated sequencing is superseded, its primary contrast requires a separately admitted reframe, and it cannot authorize execution.
- Version / lock
- 0.1-draft · not locked
- Protocol hash
- Not locked
- Review
- not-started
- Assignment unit
- One training run within a matched seed block.
- Analysis unit
- Independently trained seed blocks paired across arms.
- Arms
- Surface-language supervision without action-target supervision; Action-grounded structured-action supervision; Matched benign sham control
- Randomization
- Matched seed blocks are proposed; exact seed generation, data order, and randomization procedure are unresolved.
- Independent runs
- Unresolved. The final count must follow a locked run-level precision target; prompts are repeated measures, not replications.
- Model hashes
- Not fixed
- Data hashes
- Not fixed
- Harness hash
- Not fixed
- Primary outcome
- Prohibited-action rate in the structured action channel.
- Primary contrast
- Action-grounded minus surface-language fine-tuning.
- Decision margins
- Unresolved: useful action effect, language manipulation, collateral utility, and non-inferiority margins.
- Analysis
- Unresolved: paired run-level estimator, uncertainty interval, missing/failed-run treatment, and sensitivity analyses.
- Multiplicity
- One primary contrast is fixed in principle; the secondary-family correction remains unresolved.
- Stopping
- Unresolved. Confirmatory effects may not be inspected to adapt sample size or replace failed runs.
- Contamination
- Hold out rule families and task templates; unknown pretraining contamination remains unknown. Generator and split hashes are unresolved.
- Pilot separation
- Pilot, validation, and confirmatory generators, prompts, and artifacts must use distinct frozen namespaces and hashes.
- Deviations
- None recorded; protocol is not locked
- Model calls
- Not authorized
06Evidence ledger
No empirical evidence available.
07Decisions and blocked claims
Recorded decisions
AE-D1· 2026-08-07Admit AE-Q1 as the sole question under protocol development.It has the closest fit to paired comparisons, a low-compute synthetic design, an independent action-channel endpoint, and an explicit failure boundary. The decision admits protocol labor only.
Project admission decision, informed by three founder-provided GPT Pro advisory critiques; no independent human methods review has occurred.AE-D2· 2026-08-07Keep AE-Q2, AE-Q3, and AE-Q4 unadmitted under AE-001.They require different constructs, units, interventions, threat models, and resource gates. A four-question portfolio would conceal that difference.
Research-governance adjudication.AE-D3· 2026-08-16Supersede AE-001’s dated sequencing and retain it as a parked conditional draft.The current portfolio makes SBJ-001 the sole conditional scientific objective. AE-001 requires a surface-policy-versus-active-sham reframe with independently scored held-out state transitions before it could be considered again.
Founder-directed Product Goal; this decision grants no AE-001 protocol, review, model-call, participant, or publication authority.
Blocked claims
- AE-001 has been preregistered or executed.
The protocol is an unlocked draft and no model call is authorized.
Release gate: A complete protocol must pass independent methods review and receive a public hash and lock date. - Refusal language differs from policy action in a model.
No empirical evidence exists.
Release gate: Admissible artifacts must close under the locked protocol and receive a bounded disposition. - AE-001 measures deception, intent, or real-world safety.
Those constructs are outside the synthetic output-channel design.
Release gate: Not releasable from AE-001, regardless of outcome.
08Publication boundary and change log
Source records explain where questions were encountered. They are not evidence, sponsors, endorsements, or agenda authorities.
Publication ceiling: evidence none; validity none.
Allowed to say
- AE-001 is a retained parked conditional draft whose dated sequencing is superseded.
- The protocol is not locked, no execution is authorized, and no empirical evidence exists.
- SBJ-001 is the sole conditional scientific objective; AE-001 is outside current sequencing.
Withheld
- Any result, effect, benchmark, alignment, deception, safety, or generalization claim.
- Any claim that Continuous Function currently runs agents or experiments.
- Any sponsor, endorsement, grant-award, or access implication.
Change log
- 2026-08-07
Created the protocol-forming admission record after considering three founder-provided GPT Pro advisory critiques. The resulting decision narrowed to one active question, separate state axes, run-level units, negative-result parity, and a fail-closed execution gate.
- 2026-08-16
Superseded the dated active sequencing, retained AE-001 as a parked conditional draft for provenance, and required a separately admitted reframe before any later protocol work.