Research program · Atlas-64
Four results, three of which meant something other than they appeared to.
A program that set out to combine what several open models know into one small model. It has produced a working method for telling real effects from artifacts, a sharply narrowed question, and — so far — no model.
The result that did not survive its own audit
The most useful thing this program has produced is not a gain. It is a gain that was real, reproducible, and still wrong about its own cause.
A run reported roughly a five-point improvement from a verifier that chose between candidate answers. We retained the full archive, rehashed every member, and recomputed the metric independently. The arithmetic held: the gain was there, and it reproduced.
Then a separate check asked a question the original design had not: how much of the answer could be predicted from the prompt alone, before any candidate was scored? The answer was 95.9%. The verifier was not verifying. The inputs carried a latent mode that the prompt already revealed, and the measured gain was that mode being recovered — input-side routing wearing a verifier’s clothes.
Evidence for input-side latent routing under this fixture. Not evidence for general answer verification, and not evidence of a transformer-scale effect.
We kept the original interpretation on the record rather than replacing it silently, along with an earlier archive that was retrieved, found to be the wrong bundle, and explicitly rejected. Both are in the record.
The question we started from
The premise was an intuition, offered as one: model weights encode information about a world that exists independently of any model, so several open checkpoints ought to hold complementary pieces of the same thing. If the good pieces could be identified and combined, a small model might carry capability out of proportion to its size.
Stated literally, that premise is false, and most of the work below is the process of finding out exactly how. Stated as an intuition about concepts rather than about coordinates, a defensible version survives. Getting from one to the other is the actual result so far.
What we ran
Four runs, all at toy scale, all with retained artifacts. Scale is the standing caveat: none of these involved loading, training, or evaluating a language model, and none of them should be read as transferring to one.
| Run | What it tested | Observed | What it licenses |
|---|---|---|---|
| E1 | Averaging weights that share a coordinate system | +0.66 pp | A gain small enough to be indistinguishable from noise at this scale. |
| E1c | The same average after deliberately permuting one model | −16 pp | The control behaved as permutation symmetry predicts. It validated the harness; it is not a finding. |
| E2 | Cross-fitted competence selection vs. uniform distillation | +17–20 pp | Selecting which specialist to compile beats compiling all of them — under routing made artificially easy. |
| E3 | Independent recomputation of a reported verifier gain | +5 pp | Reproducible arithmetic, but the effect is input-side routing, not answer verification. See §1. |
The one result that still stands, and its caveat
E2 is the run worth building on. Rather than averaging weights, it selected which specialist’s behaviour to compile into a single student of the same size, and beat uniform distillation by 17–20 points across five held-out seeds.
The caveat is severe and structural, not incidental: the fixture exposes an explicit domain bit, which makes deciding which specialist to trust far easier than it would ever be in practice. The run therefore demonstrates that a competence-selection mechanism can work when routing is easy. It says nothing yet about whether routing can be learned when it is hard, which is the entire difficulty.
Selecting which specialist behaviour to compile outperforms compiling all of them uniformly, at fixed student size, across five held-out seeds.
Why the literal premise fails
Think of a concept as a city and each model as a different map projection of it. The city is real. The coordinates are not shared. Two maps can agree completely about the city and disagree about every number used to describe it.
This is not a metaphor about difficulty; it is the specific reason coordinate-space methods fail. A concept is rarely stored in one neuron, embedding, or weight. It is distributed, contextual, and represented differently at different depths. Two models can implement identical behaviour after arbitrary rotation, permutation, or rescaling of their internal axes. So:
- Matching one embedding does not establish that two models share a concept.
- Matching one output can be coincidence.
- Pairwise similarity is not transitive, so it cannot be chained into an alignment.
- Token distributions from different tokenizers cannot be compared as strings.
real concept A
│
┌────┴────┐
▼ ▼
model B model C different coordinates,
│ │ same underlying concept
└────┬────┘
▼
learned shared space alignment defined by behaviour,
│ not by weight similarity
▼
compact student modelThe working pipeline that replaced the original premise runs: semantic concept → tokenizer-specific prompt → model-generated solution → objective verifier → selected supervision → compact student. That is concept-mediated compilation. It is not weight merging, and it does not assume weights are universal concept coordinates.
What is blocked, and why
The program has verified sources, frozen protocols, an audited evaluator, and a passing tokenizer admission. It has not loaded a model. That is a deliberate ordering, not a stall: rights disposition for every planned action on every artifact is still unresolved, and compute is frozen at zero authorised spend until it is.
No model has been selected, constructed, trained, inferred, or benchmarked. Nothing on this page is a claim about model quality.
The next gate, as the program states it
Preregister the next Qwen private structural experiment before reading another weight value. The primary and official-parent libraries remain independently verified at 266,806,166,050 combined bytes. The general official-parent source packet still allows 0 of 16 actions and human legal review remains incomplete, while a separate Qwen-only engineering receipt admits only private storage and private weight statistics, 2 of 8 actions. Under that narrow exception, the 26-shard/1,045-tensor header inventory, one 1,073,741,824-byte tensor pilot, its 1,304-gate same-host independent implementation replay, and a 92-tensor/164-range broader structural-precision screen are complete. The broader screen processed 529,563,648 BF16 bytes across 13 shards after 12 of 12 pre-source tests, peaked at 258,720,248 coalition bytes, and caused zero swap growth. Its initial no-output operational failure and sole preflighted retry are recorded. Every metric, hypothesis disposition, retained-candidate identity, method ranking, and numeric delta remains withheld; no model was constructed or loaded, no inference ran, and no derivative was written. Next freeze a matched-global-storage comparison of compressed scale metadata and role-aware precision/group allocation against uniform controls, with exact bit accounting and no derivative write. Exact project-quantization rights and functional fidelity, target runtime admission, clean execution commit, canonical non-benchmark fixtures, prompt adapters, native model-memory receipts, local inference, training, evaluation, redistribution, and publication remain separate gates.
The next experiment
The parent ambition stays on the page because it is what the program is for. But the question that can actually be answered next is narrower, and falsifiable:
Can verified competence-guided behaviour compilation produce a better compact student than uniform distillation, at fixed model size and evaluation budget, when the routing signal is not handed to the selector?
The falsifier is direct: remove the explicit domain bit that made E2 easy. If the selector’s advantage over uniform distillation disappears once routing must be inferred, the mechanism does not survive, and the parent ambition needs a different route. That is a CPU-scale experiment on synthetic data with a separate verifier, and it is the cheapest test that would change what we do next.
The record
Everything above is downstream of the projection below: verified acquisition receipts, structural evidence, rights disposition, admission gates, and the claim boundary the program publishes about itself. It is reproduced in full, and it is the part to read if you intend to disagree with anything.
| Field | Value |
|---|---|
| Identifier | ATLAS-64 |
| Status | proposal only |
| Compute | frozen |
| Machine | Mac mini Mac16,11 · Apple M4 Pro · 64 GiB unified |
| Peak memory gate | 52 GiB (12 GiB reserved for the OS) |
| Paid spend cap | $0 |
| External export | not authorised |
| Result state | no model result |
| Field | Value |
|---|---|
| Candidate files verified | 91 / 91 |
| Candidate bytes | 157.4 GiB |
| Parent files verified | 59 / 59 |
| Parent bytes | 91.0 GiB |
| Combined against ceiling | 248.5 GiB / 256.0 GiB |
| Identity checks | lfs-sha256, git-blob-sha1 |
| Model instantiated | no |
| Inference performed | no |
| Metric | Value |
|---|---|
| Coordinates compared | 6,022,656 |
| Ranges read | 90 across 6 shards |
| NRMSE | 0.0117 |
| Cosine similarity | 0.999931 |
| Max absolute error | 0.00479 |
| Norm ratio | 0.999919 |
| Model instantiated | no |
| Functional fidelity proven | no |
| Gate | Status |
|---|---|
| Rights disposition | 0 allowed of 64; 64 unresolved |
| Human legal review | incomplete |
| Evaluator | 6 implemented, 6 pass and 38 rejection self-tests; used model data: no |
| Fixture generator | 30 / 30 self-tests; network used: no |
| Model load | not authorised |
| Terminal evaluation | 240 items designed across 6 strata; 0 runs executed |
| Receipt validator | 4 valid bundles accepted, 17 mutated bundles rejected; 0 real bundles validated |
| Independent reproduction | not started |
| Rung | Name | Status | Promotion requires |
|---|---|---|---|
| 0 | phase0_identity_and_feasibility | specified_blocked | Pin the host, storage root, model and quantization revisions, backends, dependencies, fixture bundle, measurement method, output schema, budget, and explicit acquisition/inference approval; then reproduce the non-benchmark smoke run. |
| 1 | local_baseline_characterized | blocked | One candidate passes local correctness, memory, stability, and cost gates under the fixed harness. |
| 2 | agent_system_gain_without_training | blocked | A preregistered ablation attributes a sealed development gain to tools, memory, verifier, or search without benchmark leakage. |
| 3 | residual_expert_gain | blocked | A rights-cleared, repository-disjoint expert improves its declared held-out family while passing retention, safety, and cost vetoes. |
| 4 | dated_k3_suite_win | blocked | Freeze a date, exact K3 endpoint, complete candidate cohort, sealed suite, common harness, and matched total budget before exposing outputs; pass simultaneous inference and independent review. |
| 5 | independent_replication | blocked | Obtain a separately approved implementation and reproduction on an independently controlled host and sealed suite. |
What may be said today
- AllowedAtlas-64 is an owner-mandated, proposal-only local small-model research program.
- AllowedThe dated landscape contains a four-candidate admission cohort and six falsifiable mechanism hypotheses: five original first-pass questions and one documented pre-execution coverage addition.
- AllowedThe observed target machine has 64 GiB unified memory; E0 reserves 12 GiB and caps candidate peak use at 52 GiB.
- AllowedThe manifest-bound eight-artifact primary library completed an independent 91-file source-identity re-read covering 169,052,193,857 bytes; inference, training, benchmark evaluation, and paid spend remain blocked.
- AllowedA separate 96 GiB giant-weight observatory completed a 124-tensor exact-range structural pilot; it did not reconstruct or run a giant checkpoint or produce behavioral evidence.
- AllowedPinned public Qwen3.6 quantizer metadata yielded a component-level mixed-precision profile and a blocked matched-budget ablation; no tensor values or model behavior were loaded.
- AllowedA preregistered 11.49 MiB official-parent range pilot compared 6,022,656 Qwen3.6 coordinates without model execution; the bounded sample is strongly consistent with the pinned candidate parent but does not prove complete lineage or behavior.
- AllowedFour source-only artifact-to-runtime lanes, six exact smoke evaluators, and a deterministic fixture generator are frozen. The model-free evaluator bundle passes six known-pass and 38 adversarial cases, including malformed UTF-8 rejection over exact base64-carried raw bytes, with byte-identical repeated receipts and a restricted AST patch runner. The reserved-namespace fixture generator passes 30/30 checks and 6/6 evaluator interoperability checks while canonical seed 640018, its public/hidden payload, authorization record, and external fixture root remain absent; no candidate artifact, model output, dependency resolution, runtime install, prompt adapter, model load, or inference has occurred.
- AllowedThe four frozen runtime commits are now independently source-verified as 4,679 tracked files and 265,119,346 bytes across clean detached MLX-LM, MLX-VLM, llama.cpp, and GPT-OSS trees. A 20-file exact-source audit finds the MLX conversion/MTP/APC/lock boundaries, llama.cpp Gemma 4/M4 Pro and explicit cache-disable paths, and GPT-OSS native-token support. It also identifies default llama.cpp prompt caching, GPT-OSS online FetchContent build inputs, no dependency lock, and a Metal sample-API/example mismatch. It does not prove bounded conversion memory, target lock validity, install/build success, cache or Harmony conformance, model compatibility, or behavior.
- AllowedA hash-bound native macOS memory sampler now passes five model-free structural tests at 20 Hz, including a Metal shared-buffer staircase and a controlled shared-page helper coalition. The implementation run observed about 1.02x private-anonymous, 2.03x shared-coalition, and 1.06x Metal-shared-buffer footprint-to-declared ratios, confirming conservative shared-page double counting. Full quiescent-host, private-Metal, MLX, realistic helper-tree, clean-commit, and independent calibration remains incomplete, so no candidate has a memory result or pass.
- AllowedThe claim-bearing acquisition verifier independently re-read all 91 files, checked LFS SHA-256 and canonical Git-blob SHA-1 identities, rejected partials, symlinks, and undeclared files, reproduced 169,052,193,857 bytes, and passed its separate receipt validator.
- AllowedAn eight-artifact rights plan separates private storage/inference, derivative training, private/public statistics, original/derivative redistribution, and commercial use; all model-facing and public actions remain unresolved.
- AllowedA contamination-limited terminal evaluation design freezes 240 objective sealed items across six equal strata, one matched total-inference envelope, simultaneous uncertainty controls, and independent reproduction; the cohort, suite payload, K3 use, execution, and claim remain blocked.
- AllowedA 97,753,972,193-byte exact official-parent follow-up for Qwen3.6 and Devstral passed a separate independent 59-file source-identity reread after the primary verification gate. Combined independently verified acquisition is 266,806,166,050 bytes below the existing 256 GiB ceiling. A Qwen-only engineering disposition admits private storage and weight statistics: a receipt-bound header inventory, one private 1 GiB tensor-range structural pilot, a separately hash-bound same-host C++20 replay, and a preregistered 92-tensor/164-range broader structural-precision screen completed without constructing or executing a model. The replay passed 8 of 8 pre-source self-tests and all 1,304 frozen comparison gates; the broader screen passed 12 of 12 pre-source tests, processed 529,563,648 bytes across 13 shards, peaked below 259 MB, and caused zero swap growth. An initial no-output commit failure is separately recorded; the only retry proved external write access before source access. These are private structural screens, not independent-host replication or model fidelity; all metric values, hypothesis dispositions, candidate identities, rankings, Devstral analysis, derivative weights, public statistics, and redistribution remain blocked.
- AllowedA post-pilot full-depth census of 235 giant router matrices is preregistered and its primary-acquisition prerequisite has passed, but its separate execution action remains unauthorized; no new census tensor has been downloaded or analyzed.
What may not
- BlockedAtlas-64 exists as a trained model or runnable agent package.
- BlockedAny proposed candidate fits, trains, or performs acceptably on this Mac.
- BlockedA community quantization preserves the official model's quality.
- BlockedKimi K3 provider-reported benchmark numbers are directly comparable to a local candidate.
- BlockedTeacher outputs are clean, licensed training data without item-level review.
- BlockedResidual experts, routing, memory, verifiers, or test-time search improve a sealed task.
- BlockedAtlas-64 beats Kimi K3, is frontier, is the world's best local model, or is independently replicated.
Candidate cohort
Hypotheses and their retirement rules
Generated 2026-07-18 from 137 hashed source documents. Schema cf-atlas-64-public-summary-v1.
Where to push back
The most useful review is not agreement. Four things here are worth attacking: whether E2’s mechanism survives without the domain bit; whether the 95.9% narrowing is itself measured correctly; whether the semantic-pivot framing rules out approaches it shouldn’t; and whether any sentence on this page claims more than its evidence allows.