Research program · Atlas-64

Four results, three of which meant something other than they appeared to.

A program that set out to combine what several open models know into one small model. It has produced a working method for telling real effects from artifacts, a sharply narrowed question, and — so far — no model.

Evidence state
Proposal · no model result
Compute
Frozen. No paid spend authorised.
Landscape cutoff
2026-07-18

The result that did not survive its own audit

The most useful thing this program has produced is not a gain. It is a gain that was real, reproducible, and still wrong about its own cause.

A run reported roughly a five-point improvement from a verifier that chose between candidate answers. We retained the full archive, rehashed every member, and recomputed the metric independently. The arithmetic held: the gain was there, and it reproduced.

Then a separate check asked a question the original design had not: how much of the answer could be predicted from the prompt alone, before any candidate was scored? The answer was 95.9%. The verifier was not verifying. The inputs carried a latent mode that the prompt already revealed, and the measured gain was that mode being recovered — input-side routing wearing a verifier’s clothes.

Narrowed — reproduced, reinterpreted

Evidence for input-side latent routing under this fixture. Not evidence for general answer verification, and not evidence of a transformer-scale effect.

We kept the original interpretation on the record rather than replacing it silently, along with an earlier archive that was retrieved, found to be the wrong bundle, and explicitly rejected. Both are in the record.

The question we started from

The premise was an intuition, offered as one: model weights encode information about a world that exists independently of any model, so several open checkpoints ought to hold complementary pieces of the same thing. If the good pieces could be identified and combined, a small model might carry capability out of proportion to its size.

Stated literally, that premise is false, and most of the work below is the process of finding out exactly how. Stated as an intuition about concepts rather than about coordinates, a defensible version survives. Getting from one to the other is the actual result so far.

What we ran

Four runs, all at toy scale, all with retained artifacts. Scale is the standing caveat: none of these involved loading, training, or evaluating a language model, and none of them should be read as transferring to one.

Runs and what each one licenses
RunWhat it testedObservedWhat it licenses
E1Averaging weights that share a coordinate system+0.66 ppA gain small enough to be indistinguishable from noise at this scale.
E1cThe same average after deliberately permuting one model−16 ppThe control behaved as permutation symmetry predicts. It validated the harness; it is not a finding.
E2Cross-fitted competence selection vs. uniform distillation+17–20 ppSelecting which specialist to compile beats compiling all of them — under routing made artificially easy.
E3Independent recomputation of a reported verifier gain+5 ppReproducible arithmetic, but the effect is input-side routing, not answer verification. See §1.

The one result that still stands, and its caveat

E2 is the run worth building on. Rather than averaging weights, it selected which specialist’s behaviour to compile into a single student of the same size, and beat uniform distillation by 17–20 points across five held-out seeds.

The caveat is severe and structural, not incidental: the fixture exposes an explicit domain bit, which makes deciding which specialist to trust far easier than it would ever be in practice. The run therefore demonstrates that a competence-selection mechanism can work when routing is easy. It says nothing yet about whether routing can be learned when it is hard, which is the entire difficulty.

Supported — at toy scale, under easy routing

Selecting which specialist behaviour to compile outperforms compiling all of them uniformly, at fixed student size, across five held-out seeds.

Why the literal premise fails

Think of a concept as a city and each model as a different map projection of it. The city is real. The coordinates are not shared. Two maps can agree completely about the city and disagree about every number used to describe it.

This is not a metaphor about difficulty; it is the specific reason coordinate-space methods fail. A concept is rarely stored in one neuron, embedding, or weight. It is distributed, contextual, and represented differently at different depths. Two models can implement identical behaviour after arbitrary rotation, permutation, or rescaling of their internal axes. So:

  • Matching one embedding does not establish that two models share a concept.
  • Matching one output can be coincidence.
  • Pairwise similarity is not transitive, so it cannot be chained into an alignment.
  • Token distributions from different tokenizers cannot be compared as strings.
   real concept A
        │
   ┌────┴────┐
   ▼         ▼
model B    model C          different coordinates,
   │         │              same underlying concept
   └────┬────┘
        ▼
 learned shared space        alignment defined by behaviour,
        │                    not by weight similarity
        ▼
 compact student model
The semantic pivot. Concepts are not averaged across models; they are re-expressed through a shared space that is defined by behaviour, then compiled into one student.

The working pipeline that replaced the original premise runs: semantic concept → tokenizer-specific prompt → model-generated solution → objective verifier → selected supervision → compact student. That is concept-mediated compilation. It is not weight merging, and it does not assume weights are universal concept coordinates.

What is blocked, and why

The program has verified sources, frozen protocols, an audited evaluator, and a passing tokenizer admission. It has not loaded a model. That is a deliberate ordering, not a stall: rights disposition for every planned action on every artifact is still unresolved, and compute is frozen at zero authorised spend until it is.

Standing boundary

No model has been selected, constructed, trained, inferred, or benchmarked. Nothing on this page is a claim about model quality.

The next gate, as the program states it

Preregister the next Qwen private structural experiment before reading another weight value. The primary and official-parent libraries remain independently verified at 266,806,166,050 combined bytes. The general official-parent source packet still allows 0 of 16 actions and human legal review remains incomplete, while a separate Qwen-only engineering receipt admits only private storage and private weight statistics, 2 of 8 actions. Under that narrow exception, the 26-shard/1,045-tensor header inventory, one 1,073,741,824-byte tensor pilot, its 1,304-gate same-host independent implementation replay, and a 92-tensor/164-range broader structural-precision screen are complete. The broader screen processed 529,563,648 BF16 bytes across 13 shards after 12 of 12 pre-source tests, peaked at 258,720,248 coalition bytes, and caused zero swap growth. Its initial no-output operational failure and sole preflighted retry are recorded. Every metric, hypothesis disposition, retained-candidate identity, method ranking, and numeric delta remains withheld; no model was constructed or loaded, no inference ran, and no derivative was written. Next freeze a matched-global-storage comparison of compressed scale metadata and role-aware precision/group allocation against uniform controls, with exact bit accounting and no derivative write. Exact project-quantization rights and functional fidelity, target runtime admission, clean execution commit, canonical non-benchmark fixtures, prompt adapters, native model-memory receipts, local inference, training, evaluation, redistribution, and publication remain separate gates.

The next experiment

The parent ambition stays on the page because it is what the program is for. But the question that can actually be answered next is narrower, and falsifiable:

Preregistered question

Can verified competence-guided behaviour compilation produce a better compact student than uniform distillation, at fixed model size and evaluation budget, when the routing signal is not handed to the selector?

The falsifier is direct: remove the explicit domain bit that made E2 easy. If the selector’s advantage over uniform distillation disappears once routing must be inferred, the mechanism does not survive, and the parent ambition needs a different route. That is a CPU-scale experiment on synthetic data with a separate verifier, and it is the cheapest test that would change what we do next.

The record

Everything above is downstream of the projection below: verified acquisition receipts, structural evidence, rights disposition, admission gates, and the claim boundary the program publishes about itself. It is reproduced in full, and it is the part to read if you intend to disagree with anything.

Program
FieldValue
IdentifierATLAS-64
Statusproposal only
Computefrozen
MachineMac mini Mac16,11 · Apple M4 Pro · 64 GiB unified
Peak memory gate52 GiB (12 GiB reserved for the OS)
Paid spend cap$0
External exportnot authorised
Result stateno model result
Acquisition — bytes and identity verified, nothing interpreted
FieldValue
Candidate files verified91 / 91
Candidate bytes157.4 GiB
Parent files verified59 / 59
Parent bytes91.0 GiB
Combined against ceiling248.5 GiB / 256.0 GiB
Identity checkslfs-sha256, git-blob-sha1
Model instantiatedno
Inference performedno
Bounded coordinate fidelity — official parent against community candidate
MetricValue
Coordinates compared6,022,656
Ranges read90 across 6 shards
NRMSE0.0117
Cosine similarity0.999931
Max absolute error0.00479
Norm ratio0.999919
Model instantiatedno
Functional fidelity provenno
Gates — what is frozen, what is unresolved
GateStatus
Rights disposition0 allowed of 64; 64 unresolved
Human legal reviewincomplete
Evaluator6 implemented, 6 pass and 38 rejection self-tests; used model data: no
Fixture generator30 / 30 self-tests; network used: no
Model loadnot authorised
Terminal evaluation240 items designed across 6 strata; 0 runs executed
Receipt validator4 valid bundles accepted, 17 mutated bundles rejected; 0 real bundles validated
Independent reproductionnot started
Claim ladder
RungNameStatusPromotion requires
0phase0_identity_and_feasibilityspecified_blockedPin the host, storage root, model and quantization revisions, backends, dependencies, fixture bundle, measurement method, output schema, budget, and explicit acquisition/inference approval; then reproduce the non-benchmark smoke run.
1local_baseline_characterizedblockedOne candidate passes local correctness, memory, stability, and cost gates under the fixed harness.
2agent_system_gain_without_trainingblockedA preregistered ablation attributes a sealed development gain to tools, memory, verifier, or search without benchmark leakage.
3residual_expert_gainblockedA rights-cleared, repository-disjoint expert improves its declared held-out family while passing retention, safety, and cost vetoes.
4dated_k3_suite_winblockedFreeze a date, exact K3 endpoint, complete candidate cohort, sealed suite, common harness, and matched total budget before exposing outputs; pass simultaneous inference and independent review.
5independent_replicationblockedObtain a separately approved implementation and reproduction on an independently controlled host and sealed suite.

What may be said today

  • AllowedAtlas-64 is an owner-mandated, proposal-only local small-model research program.
  • AllowedThe dated landscape contains a four-candidate admission cohort and six falsifiable mechanism hypotheses: five original first-pass questions and one documented pre-execution coverage addition.
  • AllowedThe observed target machine has 64 GiB unified memory; E0 reserves 12 GiB and caps candidate peak use at 52 GiB.
  • AllowedThe manifest-bound eight-artifact primary library completed an independent 91-file source-identity re-read covering 169,052,193,857 bytes; inference, training, benchmark evaluation, and paid spend remain blocked.
  • AllowedA separate 96 GiB giant-weight observatory completed a 124-tensor exact-range structural pilot; it did not reconstruct or run a giant checkpoint or produce behavioral evidence.
  • AllowedPinned public Qwen3.6 quantizer metadata yielded a component-level mixed-precision profile and a blocked matched-budget ablation; no tensor values or model behavior were loaded.
  • AllowedA preregistered 11.49 MiB official-parent range pilot compared 6,022,656 Qwen3.6 coordinates without model execution; the bounded sample is strongly consistent with the pinned candidate parent but does not prove complete lineage or behavior.
  • AllowedFour source-only artifact-to-runtime lanes, six exact smoke evaluators, and a deterministic fixture generator are frozen. The model-free evaluator bundle passes six known-pass and 38 adversarial cases, including malformed UTF-8 rejection over exact base64-carried raw bytes, with byte-identical repeated receipts and a restricted AST patch runner. The reserved-namespace fixture generator passes 30/30 checks and 6/6 evaluator interoperability checks while canonical seed 640018, its public/hidden payload, authorization record, and external fixture root remain absent; no candidate artifact, model output, dependency resolution, runtime install, prompt adapter, model load, or inference has occurred.
  • AllowedThe four frozen runtime commits are now independently source-verified as 4,679 tracked files and 265,119,346 bytes across clean detached MLX-LM, MLX-VLM, llama.cpp, and GPT-OSS trees. A 20-file exact-source audit finds the MLX conversion/MTP/APC/lock boundaries, llama.cpp Gemma 4/M4 Pro and explicit cache-disable paths, and GPT-OSS native-token support. It also identifies default llama.cpp prompt caching, GPT-OSS online FetchContent build inputs, no dependency lock, and a Metal sample-API/example mismatch. It does not prove bounded conversion memory, target lock validity, install/build success, cache or Harmony conformance, model compatibility, or behavior.
  • AllowedA hash-bound native macOS memory sampler now passes five model-free structural tests at 20 Hz, including a Metal shared-buffer staircase and a controlled shared-page helper coalition. The implementation run observed about 1.02x private-anonymous, 2.03x shared-coalition, and 1.06x Metal-shared-buffer footprint-to-declared ratios, confirming conservative shared-page double counting. Full quiescent-host, private-Metal, MLX, realistic helper-tree, clean-commit, and independent calibration remains incomplete, so no candidate has a memory result or pass.
  • AllowedThe claim-bearing acquisition verifier independently re-read all 91 files, checked LFS SHA-256 and canonical Git-blob SHA-1 identities, rejected partials, symlinks, and undeclared files, reproduced 169,052,193,857 bytes, and passed its separate receipt validator.
  • AllowedAn eight-artifact rights plan separates private storage/inference, derivative training, private/public statistics, original/derivative redistribution, and commercial use; all model-facing and public actions remain unresolved.
  • AllowedA contamination-limited terminal evaluation design freezes 240 objective sealed items across six equal strata, one matched total-inference envelope, simultaneous uncertainty controls, and independent reproduction; the cohort, suite payload, K3 use, execution, and claim remain blocked.
  • AllowedA 97,753,972,193-byte exact official-parent follow-up for Qwen3.6 and Devstral passed a separate independent 59-file source-identity reread after the primary verification gate. Combined independently verified acquisition is 266,806,166,050 bytes below the existing 256 GiB ceiling. A Qwen-only engineering disposition admits private storage and weight statistics: a receipt-bound header inventory, one private 1 GiB tensor-range structural pilot, a separately hash-bound same-host C++20 replay, and a preregistered 92-tensor/164-range broader structural-precision screen completed without constructing or executing a model. The replay passed 8 of 8 pre-source self-tests and all 1,304 frozen comparison gates; the broader screen passed 12 of 12 pre-source tests, processed 529,563,648 bytes across 13 shards, peaked below 259 MB, and caused zero swap growth. An initial no-output commit failure is separately recorded; the only retry proved external write access before source access. These are private structural screens, not independent-host replication or model fidelity; all metric values, hypothesis dispositions, candidate identities, rankings, Devstral analysis, derivative weights, public statistics, and redistribution remain blocked.
  • AllowedA post-pilot full-depth census of 235 giant router matrices is preregistered and its primary-acquisition prerequisite has passed, but its separate execution action remains unauthorized; no new census tensor has been downloaded or analyzed.

What may not

  • BlockedAtlas-64 exists as a trained model or runnable agent package.
  • BlockedAny proposed candidate fits, trains, or performs acceptably on this Mac.
  • BlockedA community quantization preserves the official model's quality.
  • BlockedKimi K3 provider-reported benchmark numbers are directly comparable to a local candidate.
  • BlockedTeacher outputs are clean, licensed training data without item-level review.
  • BlockedResidual experts, routing, memory, verifiers, or test-time search improve a sealed task.
  • BlockedAtlas-64 beats Kimi K3, is frontier, is the world's best local model, or is independently replicated.

Candidate cohort

Hypotheses and their retirement rules

sparse-backbone-deterministic-aciuntestedIntervention. Pair the locally admitted sparse backbone with one deterministic localization -> repair -> validation agent-computer interface and revision-bound repository retrieval. Threshold. >=5 percentage-point held-out repository-task gain Retired if. Retire if the gain is below threshold on repository- and time-disjoint development tasks or if any resource, contamination, retention, or severe safety veto fails.
phase-routed-residual-expertsuntestedIntervention. Train common-base QLoRA or DoRA experts for localization/planning, implementation, and testing/review, then select one expert per declared phase without modifying the backbone identity. Threshold. Oracle phase routing shows >=4 percentage-point attainable gain; the learned router captures >=75% of that gain and beats the monolithic adapter by >=3 points. Retired if. Retire if a parameter-matched monolithic adapter is non-inferior, if the router cannot recover 75% of oracle routing, or if retention/rights fail.
verifier-filtered-trajectory-distillationuntestedIntervention. Distill rights-cleared teacher proposals, execution failures, verifier feedback, and successful repairs using hidden mutation/property tests. Threshold. >=3 percentage-point gain over equal-token success-only distillation on repository- and time-disjoint held-out tasks. Retired if. Retire if failures plus repair traces do not beat equal-token success-only data or if any provenance, leakage, rights, or safety gate fails.
adaptive-verification-budgetinguntestedIntervention. Choose N from {1, 2, 4, 8} under a fixed total token and wall-time budget, execute candidates, and use a calibrated verifier to stop or continue. Threshold. >=10% more verified successes per million generated tokens and >=4 percentage-point absolute task gain over single-sample inference. Retired if. Retire if the gain disappears under matched total budget, calibration fails, or verifier-selected outputs exploit visible tests.
revision-bound-structured-memoryuntestedIntervention. Construct a revision-bound repository/source graph plus a provenance-safe failure memory and retrieve only location-bearing records. Threshold. Non-inferior within 2 percentage points while reducing input tokens by >=25% and tool calls by >=20%. Retired if. Retire if non-inferiority fails, source-location precision drops below threshold, or memory introduces leakage/staleness failures.
verifier-guided-search-to-weight-compilationuntestedIntervention. After H4 independently establishes a positive verifier-guided search gain, convert rights-cleared selected and rejected development candidates, verifier deltas, and successful repairs into a common-base residual expert, then evaluate that expert with search disabled. Threshold. With search disabled, retain >=80% of H4's absolute verified-search gain, beat the equal-update-token success-only adapter by >=2 percentage points, and reduce generated tokens and verifier calls by >=50% versus H4 search. Retired if. Retire if the retained-gain ratio is below 0.80, the advantage over the matched success-only adapter is below 2 points, inference cost is not cut by half, or any separation, rights, retention, or safety gate fails.

Generated 2026-07-18 from 137 hashed source documents. Schema cf-atlas-64-public-summary-v1.

Where to push back

The most useful review is not agreement. Four things here are worth attacking: whether E2’s mechanism survives without the domain bit; whether the 95.9% narrowing is itself measured correctly; whether the semantic-pivot framing rules out approaches it shouldn’t; and whether any sentence on this page claims more than its evidence allows.

Other programs · Investigations you can run in the browser