Our north star

The human layer for AI: people and AI, unbroken.

However capable AI becomes, and even as it begins to help build AI, the people working with it should still be able to see what it is doing, understand it, question it and steer it in time. It starts with machine-learning research. This page maps that direction onto what exists today, verb by verb.

Status
A direction, not a product
Set
30 September 2026
Accounts
None
Repository
Private

The record of who predicted what, who corrected what, and what held up.

The part of the direction to build first. Today that record holds your own guesses, corrections and what held up on this site, kept in your browser; across every person, provider and service is the direction, not yet built.

The direction

Picture one workspace per question, shared by you and an AI partner. You hold what must not move: a test set sealed even from the AI, a limit on what may be traded away. The partner is free to explore everything else. Its proposals arrive in pencil; only you ink a plan, and ink means authorised, not true. The work runs while you are away, within the limits you set, and in the morning the same canvas shows what changed, what held and the one decision that needs you.

Claims are narrowed to their evidence before they leave, and travel as packages that someone else can replay, fork and extend. As AI begins to improve its own methods, those changes pass the same gate: tried on fresh tasks, judged by a checker the improver cannot change, and adopted only where a person agrees.

None of that exists yet as a product. One loop of it runs today, small and real, in the workspace.

The map: six verbs, and where each exists today

The newest film is built from six verbs. Each link below opens the step where a first piece of the verb exists now, and says what answers you there. Everything runs in your browser; nothing calls an AI model.

  1. Show: see the failure, image by image

    What answers: Computed in your browser from the benchmark’s numbers.

    Today: 600 validation photos as anonymous dots (no photographs are shipped), the trade-off every option makes, and which of them each option fixed or broke, all computed in your browser.

    Not built: 3D views and teaching scenes you direct by voice and pen.

  2. Ask: where does it show that?

    What answers: Written in advance; in a draft claim, computed from your record.

    Today: tap a word in a draft claim and the scripted partner answers from the results; there a claim is built only from phrases generated from the numbers, which you choose among; your own words go in a note beside it, never in it, and are not checked. A package from elsewhere shows its claim as written.

    Not built: asking an AI partner out loud; no model is called on this site.

  3. Guess: commit before the answer arrives

    What answers: Your guess, against what is then measured; the partner’s bet is written in advance.

    Today: place your guess on the trade-off, with a reason, before anything runs; it resolves against the measured result beside the partner’s scripted bet.

    Not built: forecast records over many calls, for you and for an AI partner.

  4. Hold: what must not move, including a sealed test

    What answers: Enforced in your browser, not advised: the workspace’s gate, budget and one look; retrieval’s frozen settings.

    Today: hold the average within a limit, the test images sealed and the compute under a time budget. Holds are remembered in pencil and take effect once you ink: the gate judges every candidate option against your limit, going past its verdict needs your acknowledgement and is disclosed, the refit loop stops at your budget, and a checker opens the test file once per workspace, after recording the look in this browser.

    Not built: holds on real compute, data and agents beyond this page, and a checker on separate infrastructure whose files an agent cannot read.

  5. Mark: their words, not a summary

    What answers: From the source: the sentences are quoted from the paper and cited.

    Today: the paper’s own sentences, quoted and cited; mark the word you doubt and it travels with your work, as a pencil note: “not necessarily in our model: test it”.

    Not built: a paper read aloud on a walk, with your marks waiting at the desk.

  6. Ink: authorise, then publish something others can build on

    What answers: Your ink, bound to the plan’s digest; the runs are then computed in your browser.

    Today: only your ink runs a plan, bound by its digest; your claim, the plan history, every completed run, this workspace’s looks and every removal go into a package that anyone can replay run by run or fork into their own pencil.

    Not built: a shared record across people and services, and a commons where another lab adds a test to your evaluation.

The same idea also runs in a smaller form: train a small model on a synthetic three-feature task and look inside one checkpoint. Lessons that teach the mechanisms underneath start with attention.

What answers you, and where a live model would plug in

Every line a partner says carries one of these stamps, always visible beside it:

written in advance
the partner’s scripted lines: the first exchange, its bets, its proposals in pencil.
computed from your record
answers worked out from what you did here: your guesses, runs, holds and marks.
from the source
a paper’s own sentences, quoted and cited, never summarised.
model · provider · cost
shown only when a live model answers. None is connected, so no line on this site carries it.

A live model would plug in where you ask: the questions you tap, and a question you bring in your own words. Its proposals would arrive in pencil with the fourth stamp, and only your ink would make them yours. The connection for it exists and is switched off. How much a partner may do is a dial you set on your desk; it starts at propose, and nothing above propose is built.

The vision films

Short films describe the direction. Each says so on screen: they show an imagined day, not software that exists. The newest grounds its one benchmark rerun in a real, measured result and labels everything else as illustrative. They are where this is going, not what this site is.

What We Can Build On4:24Vision filmA researcher and an AI partner turn one model failure into a narrowed claim that another lab builds on: limits that become enforced, a test sealed even from the AI, and a measured rerun of a published result.
The North Star3:35Vision filmWhy the thread between people and AI breaks, and a layer above every model, agent, tool and cloud where it holds.
The North Star, research edition4:41Vision filmOne machine-learning researcher’s imagined day as AI begins to help build AI: the morning page, the services beneath the layer, and a person at every gate.
Watching

To watch the films, email founder@continuousfunction.ai.

How it is being built

The first real instance of the direction is the work on this site. Much of the implementation is written with AI coding agents working under my direction; I decide what to build and what ships, and agent-written code has to pass the same checks as any other change.

Next is the core beneath the workspace: one record of what was proposed, approved, run and shared, enforcing holds outside the page. It would be tried first on synthetic data, then with outside teams on their own work. It is planned, not built.

Open is the direction. Today the source repository is private.

In more detail

What exists, what does not, and the status of every research question are set out in How Continuous Function works.

The open question

Everything above is one small loop: a guess before the run, a hold that is enforced, an ink that means authorised, not true, and a record of what held up. Whether that loop can stay this clear when the work is larger, faster and partly done by AI is not known. It is the question this site exists to work on.

When AI begins to help build AI, can the people responsible still see each change while it is small enough to steer?