Learning and research in modern AI

Understand how modern AI works, precisely enough to question it.

Continuous Function teaches the mathematics and mechanisms of machine learning, from the dot product to open research questions. Every idea is shown four ways at once — as intuition, as mathematics, as code and as a working instrument — and the four always agree.

One idea, four views

Scaled dot-product attention is the operation at the centre of every transformer. Choose the word the model is reading: the sentence, the equation, the code and the figure change together, because they are one calculation.

Query token

A query can read only itself and earlier tokens.

Intuition

Reading sat, this query gives the most weight to cat (77.8%), followed by the (18.9%) and sat (3.2%). Future token on (0.0%) is masked.

Each weight tells how much of a token’s value enters the output. Select a query to change the mixture.

Mathematics

q is the query, k a key, v a value; each has two coordinates. The visible scores s become weights a; future scores are set to −∞. The output is o.

sj=q⋅kj2,a=softmax⁡(s)s_j=\frac{q\cdot k_j}{\sqrt{2}},\quad a=\operatorname{softmax}(s)
s≈(s\approx(1.0611.061,2.4752.475,−0.707-0.707,−∞-\infty))
a≈(a\approx(18.9%18.9\%,77.8%77.8\%,3.2%3.2\%,0.0%0.0\%))
o=∑j=13ajvj≈(0.654, 1.714)o=\sum_{j=1}^{3}a_jv_j\approx(0.654,\,1.714)

Code

NumPy · visible rows 1–3

import numpy as np
q = np.array([1.5,0.5])
K = np.array([[1,0],[2,1],[-1,1]])
V = np.array([[-1,1],[1,2],[2,-1]])
s = K @ q / np.sqrt(2)
a = np.exp(s - s.max()); a /= a.sum()
o = a @ V

a ≈ (18.9%, 77.8%, 3.2%)
o ≈ (0.654, 1.714)

Instrument

Weights aValues v → output othe18.9%cat77.8%sat3.2%onmasked-1-1001122thecatsatono

Hover or focus a token row to trace it through the other views. The arrow is the weighted mixture of the value points.

One calculation, four views. Constructed vectors, fixed keys and values; no trained model. Displayed numbers are rounded; masked weights are exactly zero.Python witness

Query sat, position 3. the: 18.9%; cat: 77.8%; sat: 3.2%; on: 0.0%, masked. Output o = (0.654, 1.714).

Every notebook here is built this way. Begin from the intuition or from the equation; change an input and see which numbers move. Read the attention notebook

The curriculum

Seven parts and 51 notebooks lead from mathematical foundations to the methods used to align and inspect today's models. Read them in order, or begin wherever your question does.

  1. Foundations. The cubic f(x) = x³ − x and its tangent at x₀ = 0.82. The warm tangent has slope f′(x₀) = 1.0172; its triangle has run Δx = 0.28 and rise Δy = f′(x₀)Δx.IMathematical foundationsVectors, derivatives and probability: the language the rest is written in.13 notebooks
  2. Optimization. Ten gradient-descent steps on L(θ) = ½θᵀAθ, where A has eigenvalues 1 and 10 and axes rotated by −0.38 radians. A warm path with step size 0.18 zigzags toward the minimum through exact quadratic level sets.IILearning and optimizationHow a model improves: descending a loss surface, one step at a time.4 notebooks
  3. The transformer. A six-by-six causal attention matrix from fixed scores. Square area is proportional to the row-normalized softmax weight; every row sums to one. Future positions above the diagonal are masked to zero. The fifth row is warm.IIIThe transformerFrom text to tokens to attention: the architecture behind language models.5 notebooks
  4. Generative models. Forward diffusion of an equal mixture of N(−2, 0.42²) and N(2, 0.42²). The warm curve is the data density; signal retention ᾱ = 0.6, 0.18 and 0.0025 gives successively noisier densities under xₜ = √ᾱ x₀ + √(1−ᾱ) ε, with independent standard normal ε. The faint final curve is near N(0,1).IVGenerative modellingLearning a distribution well enough to draw new samples from it.5 notebooks
  5. Scale. Illustrative loss power law L(C) = 0.6 + 2.4 C^(−0.25), where C is dimensionless compute. Both axes are logarithmic. Warm example points lie on the law; the dashed line is the irreducible floor E = 0.6. These are illustrative values, not measured scaling results.VScale and generalizationWhy larger models trained on more data behave as they do.4 notebooks
  6. Systems. The signal tanh(x) and its warm three-bit uniform quantization Q(tanh(x)) to eight levels from −1 to 1. The marked spacing is Δ = 2/7; the absolute quantization error is at most Δ/2 for inputs in [−1,1].VIEfficient training and inferenceMaking models cheaper to run, and knowing exactly what that changes.13 notebooks
  7. Alignment. Reward over-optimization with d = √KL, the square root of policy KL divergence. Gold reward R(d) = d(1 − 0.6 log d), with R(0) = 0, uses the RL form of Gao, Schulman and Hilton (2022) with illustrative coefficients, not measured rewards. Its warm maximum is at d* = exp(1/0.6 − 1). The rising proxy is schematic; the empirical gold form is not reliable near the origin.VIIAlignment, reasoning and interpretabilityShaping behaviour with human preferences, and checking what was learned.7 notebooks
  8. Research. Ordinary least squares fitted to six synthetic observations. Past the last observed x, the fit becomes a dashed extrapolation in a tinted unknown region. A pointwise 95% prediction interval for one future observation widens away from the data mean. It assumes a linear mean and independent normal errors with common variance; it does not cover failure of the linear model.VIIIResearchWhere the curriculum ends, the open questions begin.Open questions

The full curriculum

From understanding to evidence

Understanding shows in what you can do with it. The lab poses questions about real models and data that you work through in your browser: predict, change one thing, compare, and decide what the result supports.

Why this exists

AI systems are growing more capable and have begun to help build their successors. However far that goes, the people who work with them should still be able to see what they are doing, understand it, question it and steer it in time.

That is the human layer for AI, and it begins with understanding. Continuous Function starts as a place to learn the mechanisms precisely and to test claims about them, and is growing toward research carried out by people and AI together, with people deciding.