Learning and research in modern AI

Understand how modern AI works, precisely enough to question it.

Continuous Function teaches the mathematics and mechanisms of machine learning, from the dot product to open research questions. Each notebook sets out one idea as intuition, mathematics, code and a figure you can change; the example below shows all four for one calculation.

One idea, four views

Scaled dot-product attention is the operation at the centre of the transformer architecture. Choose the query word: the sentence, the equation, the code and the figure change together, because they show one calculation.

Query token

A query can read only itself and earlier tokens.

Intuition

Reading sat, this query gives the most weight to cat (77.8%), followed by the (18.9%) and sat (3.2%). Future token on (0.0%) is masked.

Each weight tells how much of a token’s value enters the output. Select a query to change the mixture.

Mathematics

q is the query, k a key, v a value; each has d = 2 coordinates, so the scores are divided by √2. The softmax turns the scores s into weights a; scores of later tokens are set to −∞, so their weights are exactly zero. The output is o.

sj=q⋅kj2,a=softmax⁡(s)s_j=\frac{q\cdot k_j}{\sqrt{2}},\quad a=\operatorname{softmax}(s)
s≈(s\approx(1.0611.061,2.4752.475,−0.707-0.707,−∞-\infty))
a≈(a\approx(18.9%18.9\%,77.8%77.8\%,3.2%3.2\%,0.0%0.0\%))
o=∑j=13ajvj≈(0.654, 1.714)o=\sum_{j=1}^{3}a_jv_j\approx(0.654,\,1.714)

Code

NumPy · tokens 1–3, the ones this query can read

import numpy as np
q = np.array([1.5,0.5])
K = np.array([[1,0],[2,1],[-1,1]])
V = np.array([[-1,1],[1,2],[2,-1]])
s = K @ q / np.sqrt(2)
a = np.exp(s - s.max()); a /= a.sum()
o = a @ V

a ≈ (18.9%, 77.8%, 3.2%)
o ≈ (0.654, 1.714)

Figure

Weights aValues v → output othe18.9%cat77.8%sat3.2%onmasked-1-1001122thecatsatono

Hover or focus a token row to trace it through the other views. The arrow is the weighted mixture of the value points.

One calculation, four views. Constructed vectors, fixed keys and values; no trained model. Displayed numbers are rounded; masked weights are exactly zero.Download this calculation in Python

Query sat, position 3. the: 18.9%; cat: 77.8%; sat: 3.2%; on: 0.0%, masked. Output o = (0.654, 1.714).

The notebooks follow the same plan at greater length: intuition, mathematics, code and a figure you can change. Begin from the intuition or from the equation. Read the attention notebook

The curriculum

Seven parts and 51 notebooks lead from mathematical foundations to the methods used to align and inspect today's models, and the course ends at open research questions. Read the parts in order, or begin wherever your question does.

  1. Foundations. The cubic f(x) = x³ − x and its tangent at x₀ = 0.82. The highlighted tangent has slope f′(x₀) = 1.0172; its triangle has run Δx = 0.28 and rise Δy = f′(x₀)Δx.IMathematical foundationsVectors, derivatives and probability: the language the rest is written in.13 notebooks
  2. Optimization. Ten gradient-descent steps on L(θ) = ½θᵀAθ, where A has eigenvalues 1 and 10 and axes rotated by −0.38 radians. A highlighted path with step size 0.18 zigzags toward the minimum through exact quadratic level sets.IILearning and optimizationHow a model improves: descending a loss surface, one step at a time.4 notebooks
  3. The transformer. A six-by-six causal attention matrix from fixed scores. Square area is proportional to the row-normalized softmax weight; every row sums to one. Future positions above the diagonal are masked to zero. The fifth row is highlighted.IIIThe transformerFrom text to tokens to attention: the architecture behind language models.5 notebooks
  4. Generative models. Forward diffusion of an equal mixture of N(−2, 0.42²) and N(2, 0.42²). The highlighted curve is the data density; signal retention ᾱ = 0.6, 0.18 and 0.0025 gives successively noisier densities under xₜ = √ᾱ x₀ + √(1−ᾱ) ε, with independent standard normal ε. The faint final curve is near N(0,1).IVGenerative modellingLearning a distribution well enough to draw new samples from it.5 notebooks
  5. Scale. Illustrative loss power law L(C) = 0.6 + 2.4 C^(−0.25), where C is dimensionless compute. Both axes are logarithmic. Highlighted example points lie on the law; the dashed line is the irreducible floor E = 0.6. These are illustrative values, not measured scaling results.VScale and generalizationWhy larger models trained on more data behave as they do.4 notebooks
  6. Systems. The signal tanh(x) and, highlighted, its three-bit uniform quantization Q(tanh(x)) to eight levels from −1 to 1. The marked spacing is Δ = 2/7; the absolute quantization error is at most Δ/2 for inputs in [−1,1].VIEfficient training and inferenceMaking models cheaper to run, and knowing exactly what that changes.13 notebooks
  7. Alignment. Reward over-optimization with d = √KL, the square root of policy KL divergence. Gold reward R(d) = d(1 − 0.6 log d), with R(0) = 0, uses the RL form of Gao, Schulman and Hilton (2023) with illustrative coefficients, not measured rewards. Its maximum, highlighted, is at d* = exp(1/0.6 − 1). The rising proxy is schematic; the empirical gold form is not reliable near the origin.VIIAlignment, reasoning and interpretabilityShaping behaviour with human preferences, and checking what was learned.7 notebooks
  8. Research. Ordinary least squares fitted to six synthetic observations. Past the last observed x, the fit becomes a dashed extrapolation in a tinted unknown region. A pointwise 95% prediction interval for one future observation widens away from the data mean. It assumes a linear mean and independent normal errors with common variance; it does not cover failure of the linear model.ResearchQuestions not yet settled, each written down with what would count against it before any study is run.Open questions

The full curriculum

From understanding to evidence

The lab poses questions that you work through in your browser, some on real benchmark data and some on small models you train yourself: predict, change one thing, compare, and decide what the result supports.

A prototype for learning alongside an AI assistant: predict whether one gradient step lowers a loss, check the calculation, then take the same example to an assistant you already use. Its hints are prepared in advance; no AI model is connected here. Open the prototype

Why this exists

People who build, evaluate or rely on AI systems need to be able to check what those systems do: follow the calculation behind a claim, test it on a small case, and see where the result stops holding.

Continuous Function is built to help with that. The curriculum teaches the mechanisms of machine learning precisely, and the lab gives practice in testing claims about them. The longer aim is research in which people work with AI tools and make the decisions themselves; the vision page describes that direction and how much of it exists today.