Learning and research in modern AI
Understand how modern AI works, precisely enough to question it.
Continuous Function teaches the mathematics and mechanisms of machine learning, from the dot product to open research questions. Each notebook sets out one idea as intuition, mathematics, code and a figure you can change; the example below shows all four for one calculation.
One idea, four views
Scaled dot-product attention is the operation at the centre of the transformer architecture. Choose the query word: the sentence, the equation, the code and the figure change together, because they show one calculation.
A query can read only itself and earlier tokens.
Intuition
Reading sat, this query gives the most weight to cat (77.8%), followed by the (18.9%) and sat (3.2%). Future token on (0.0%) is masked.
Each weight tells how much of a token’s value enters the output. Select a query to change the mixture.
Mathematics
q is the query, k a key, v a value; each has d = 2 coordinates, so the scores are divided by √2. The softmax turns the scores s into weights a; scores of later tokens are set to −∞, so their weights are exactly zero. The output is o.
Code
NumPy · tokens 1–3, the ones this query can read
import numpy as np
q = np.array([1.5,0.5])
K = np.array([[1,0],[2,1],[-1,1]])
V = np.array([[-1,1],[1,2],[2,-1]])
s = K @ q / np.sqrt(2)
a = np.exp(s - s.max()); a /= a.sum()
o = a @ Va ≈ (18.9%, 77.8%, 3.2%)
o ≈ (0.654, 1.714)
Figure
Hover or focus a token row to trace it through the other views. The arrow is the weighted mixture of the value points.
Query sat, position 3. the: 18.9%; cat: 77.8%; sat: 3.2%; on: 0.0%, masked. Output o = (0.654, 1.714).
The notebooks follow the same plan at greater length: intuition, mathematics, code and a figure you can change. Begin from the intuition or from the equation. Read the attention notebook
The curriculum
Seven parts and 51 notebooks lead from mathematical foundations to the methods used to align and inspect today's models, and the course ends at open research questions. Read the parts in order, or begin wherever your question does.
- IMathematical foundationsVectors, derivatives and probability: the language the rest is written in.13 notebooks
- IILearning and optimizationHow a model improves: descending a loss surface, one step at a time.4 notebooks
- IIIThe transformerFrom text to tokens to attention: the architecture behind language models.5 notebooks
- IVGenerative modellingLearning a distribution well enough to draw new samples from it.5 notebooks
- VScale and generalizationWhy larger models trained on more data behave as they do.4 notebooks
- VIEfficient training and inferenceMaking models cheaper to run, and knowing exactly what that changes.13 notebooks
- VIIAlignment, reasoning and interpretabilityShaping behaviour with human preferences, and checking what was learned.7 notebooks
- ResearchQuestions not yet settled, each written down with what would count against it before any study is run.Open questions
From understanding to evidence
The lab poses questions that you work through in your browser, some on real benchmark data and some on small models you train yourself: predict, change one thing, compare, and decide what the result supports.
- InvestigationIs this bird classifier ready to ship?A bird classifier is right 97% of the time on photos mixed like its training data, yet wrong on nearly half of one group. Work out why on real benchmark features, and decide what would have to change before it ships.
- ExperimentTrain a small model, and catch it taking a shortcutFit a small network in your browser on a task where an easy cue agrees with the label during training and reverses at test. Look inside, change one setting, and compare.
- ToolDoes an improvement survive a paired comparison?Paste paired scores from your own evaluation and read what they support: the size of the difference, whether it repeats across reruns, and the design questions no statistic can settle.
A prototype for learning alongside an AI assistant: predict whether one gradient step lowers a loss, check the calculation, then take the same example to an assistant you already use. Its hints are prepared in advance; no AI model is connected here. Open the prototype
Why this exists
People who build, evaluate or rely on AI systems need to be able to check what those systems do: follow the calculation behind a claim, test it on a small case, and see where the result stops holding.
Continuous Function is built to help with that. The curriculum teaches the mechanisms of machine learning precisely, and the lab gives practice in testing claims about them. The longer aim is research in which people work with AI tools and make the decisions themselves; the vision page describes that direction and how much of it exists today.