Learning and research in modern AI
Understand how modern AI works, precisely enough to question it.
Continuous Function teaches the mathematics and mechanisms of machine learning, from the dot product to open research questions. Every idea is shown four ways at once — as intuition, as mathematics, as code and as a working instrument — and the four always agree.
One idea, four views
Scaled dot-product attention is the operation at the centre of every transformer. Choose the word the model is reading: the sentence, the equation, the code and the figure change together, because they are one calculation.
A query can read only itself and earlier tokens.
Intuition
Reading sat, this query gives the most weight to cat (77.8%), followed by the (18.9%) and sat (3.2%). Future token on (0.0%) is masked.
Each weight tells how much of a token’s value enters the output. Select a query to change the mixture.
Mathematics
q is the query, k a key, v a value; each has two coordinates. The visible scores s become weights a; future scores are set to −∞. The output is o.
Code
NumPy · visible rows 1–3
import numpy as np
q = np.array([1.5,0.5])
K = np.array([[1,0],[2,1],[-1,1]])
V = np.array([[-1,1],[1,2],[2,-1]])
s = K @ q / np.sqrt(2)
a = np.exp(s - s.max()); a /= a.sum()
o = a @ Va ≈ (18.9%, 77.8%, 3.2%)
o ≈ (0.654, 1.714)
Instrument
Hover or focus a token row to trace it through the other views. The arrow is the weighted mixture of the value points.
Query sat, position 3. the: 18.9%; cat: 77.8%; sat: 3.2%; on: 0.0%, masked. Output o = (0.654, 1.714).
Every notebook here is built this way. Begin from the intuition or from the equation; change an input and see which numbers move. Read the attention notebook
The curriculum
Seven parts and 51 notebooks lead from mathematical foundations to the methods used to align and inspect today's models. Read them in order, or begin wherever your question does.
- IMathematical foundationsVectors, derivatives and probability: the language the rest is written in.13 notebooks
- IILearning and optimizationHow a model improves: descending a loss surface, one step at a time.4 notebooks
- IIIThe transformerFrom text to tokens to attention: the architecture behind language models.5 notebooks
- IVGenerative modellingLearning a distribution well enough to draw new samples from it.5 notebooks
- VScale and generalizationWhy larger models trained on more data behave as they do.4 notebooks
- VIEfficient training and inferenceMaking models cheaper to run, and knowing exactly what that changes.13 notebooks
- VIIAlignment, reasoning and interpretabilityShaping behaviour with human preferences, and checking what was learned.7 notebooks
- VIIIResearchWhere the curriculum ends, the open questions begin.Open questions
From understanding to evidence
Understanding shows in what you can do with it. The lab poses questions about real models and data that you work through in your browser: predict, change one thing, compare, and decide what the result supports.
- InvestigationIs this bird classifier ready to ship?A bird classifier is right 97% of the time, yet wrong on nearly half of one group. Work out why on real benchmark features, and decide what would have to change before it ships.
- ExperimentTrain a small model, and catch it taking a shortcutFit a small network in your browser on a task where an easy cue agrees with the label during training and reverses at test. Look inside, change one setting, and compare.
- InstrumentDoes an improvement survive a paired comparison?Read paired evaluation scores under stated assumptions: the size of the effect, its uncertainty, and the next test that would settle it.
Why this exists
AI systems are growing more capable and have begun to help build their successors. However far that goes, the people who work with them should still be able to see what they are doing, understand it, question it and steer it in time.
That is the human layer for AI, and it begins with understanding. Continuous Function starts as a place to learn the mechanisms precisely and to test claims about them, and is growing toward research carried out by people and AI together, with people deciding.