Curriculum
A course in modern AI, from first principles to open questions.
Seven parts and 51 notebooks, ordered so that each idea rests on the ones before it. Comfort with algebra and a little programming is enough to begin; later parts assume the earlier ones, and each notebook names its prerequisites.
How each notebook teaches
Every notebook presents one idea four ways. The four views describe the same object with the same numbers, so that a step you cannot follow in one can be checked in another.
- 01IntuitionWhat the idea does, in plain words and one picture, before any symbol is introduced.
- 02MathematicsThe definitions and the derivation, with every symbol defined and every assumption stated.
- 03CodeA short program that computes the same object, on the same numbers, so the equation can be checked line by line.
- 04InstrumentA figure you can change. Predict what will happen, change one input, and see which quantities respond.
Part I
Mathematical foundations
The dot product as a measure of agreement, the derivative carried backwards through a computation, and probability as the bookkeeping of uncertainty. The part ends with likelihood, cross-entropy and relative entropy, which return later as the losses that models are trained to minimise.
- I.1Vector SpacesA vector space is a set of objects that can be added and scaled by numbers, with both operations obeying a few consistent rules.Introductory
- I.2Dot ProductThe dot product sums coordinate-wise products and equals the two lengths multiplied by the cosine of the angle between them.Introductory
- I.3DerivativesThe derivative is the limit of a secant’s slope as its step shrinks to zero: the instantaneous rate of change at a point.Introductory
- I.4Computation GraphsA computation graph writes a calculation as small operations joined by dependencies, so values flow forward and derivatives back.Introductory
- I.5Reverse-Mode Automatic DifferentiationReverse mode records the operations that ran, then propagates adjoints backwards to give a scalar’s full gradient in one sweep.Intermediate
- I.6BackpropagationBackpropagation carries the chain rule backwards from the loss, reusing each node’s gradient so one pass reaches every weight.Intermediate
- I.7Probability BasicsProbability assigns mass to sets of outcomes; conditioning keeps only the outcomes the evidence allows and renormalises.Introductory
- I.8Random VariablesA random variable is a function that assigns a number to each outcome; its randomness comes only from which outcome occurs.Introductory
- I.9DistributionsA distribution says how probability is spread over the values a random variable can take, as point masses or as a density.Introductory
- I.10Maximum LikelihoodMaximum likelihood holds the data fixed and picks the parameters that make them most probable, minimising negative log-likelihood.Intermediate
- I.11Bayesian InferenceBayesian inference multiplies a prior by the likelihood and normalises, giving a distribution over the unknowns, not one estimate.Intermediate
- I.12Cross-EntropyCross-entropy is a model’s expected surprise under the true distribution: that distribution’s entropy plus the KL divergence.Intermediate
- I.13KL DivergenceKL divergence is the expected extra log-loss of using Q where P is true; it is zero only when they agree, and is not symmetric.Intermediate
Guided practice
Moving the camera or moving the object? — Four landmarks through a change of coordinate frame, in three dimensions and on the camera image, with the classic inverted-transform error beside the correct one.
Part II
Learning and optimization
Gradient descent and the adaptive methods used in practice, what the shape of the loss landscape suggests about stability and generalization, and how schedules for the step size are designed around it.
- II.1Gradient DescentGradient descent repeatedly steps against the gradient, scaled by a learning rate, using only the local slope to lower the loss.Introductory
- II.2Adam OptimizerAdam divides a running average of each parameter’s gradient by the root of a running average of its square, both bias-corrected.Intermediate
- II.3Loss Landscapes, Sharpness & Flat MinimaCurvature near a minimum measures its sharpness, a sensitivity to weight changes that also depends on how the weights are scaled.Intermediate
- II.4Learning Rate SchedulesA schedule varies the learning rate over training: usually a short warmup from small steps, then a decay, often along a cosine.Intermediate
Part III
The transformer
Scaled dot-product attention and the residual block first; then what the learned vectors represent, how text becomes tokens, and the normalization and position schemes that make deep stacks trainable. What attention costs in memory, and how that cost is engineered down, is taken up in Part VI.
- III.1Scaled Dot-Product Attention & Transformer LayersScaled dot products of a query with every key, passed through a softmax, give the weights for averaging the values.Intermediate
- III.2Representation Learning & Embedding GeometryA learned representation maps inputs to vectors whose distances and directions can encode meaning, context and similarity.Intermediate
- III.3Tokenization & Vocabulary DesignByte-pair encoding builds a vocabulary by repeatedly merging the most frequent adjacent pair; text is then read as token IDs.Intermediate
- III.4Layer Normalization & RMSNormLayerNorm centres each token’s features and divides by their standard deviation; RMSNorm divides by the root mean square alone.Intermediate
- III.5Rotary Position EmbeddingsRoPE rotates query and key coordinate pairs by angles proportional to position, so scores depend on position only via the offset.Intermediate
Guided practice
The transformer lab — Eight stations that rebuild a transformer block by prediction and calculation.
Part IV
Generative modelling
Latent-variable models and the evidence lower bound, invertible flows with exact densities, and the score-based, diffusion and flow-matching methods behind modern image and video generation, presented as one family of ideas.
- IV.1Variational AutoencodersA variational autoencoder maximises a lower bound on log-likelihood; the gap is the KL divergence between encoder and posterior.Advanced
- IV.2Normalizing FlowsA normalizing flow warps a simple density through an invertible map; the change-of-variables formula gives its exact likelihood.Advanced
- IV.3Diffusion, Score-Based Models & Flow MatchingDiffusion adds Gaussian noise to data step by step and trains a network to undo it, so that sampling can turn noise into data.Advanced
- IV.4Score Matching & Score-Based Generative ModelsThe score, the gradient of the log-density, does not involve the normalising constant and can be estimated by learning to denoise.Advanced
- IV.5Flow Matching & Rectified FlowsFlow matching regresses a velocity field onto known velocities along paths from noise to data, then integrates it to draw samples.Advanced
Part V
Scale and generalization
The double-descent curve, empirical scaling laws and what they do and do not predict, the infinite-width limit of the neural tangent kernel, and how additional computation at inference time can be traded for accuracy.
- V.1Overparameterization & GeneralizationTest error can peak where a model first fits its training data exactly, then fall again as capacity grows beyond that threshold.Intermediate
- V.2Scaling Laws & Emergent AbilitiesTest loss falls empirically as a power law in parameters, data and compute, so small runs can forecast larger ones.Intermediate
- V.3Neural Tangent Kernel (NTK) & Infinite-Width LimitsAt infinite width a network stays near its linearisation at initialisation and trains like kernel regression with a fixed kernel.Advanced
- V.4Test-Time ComputeSampling more answers helps only if some sample is correct and the verifier selects it; coverage and selection fail separately.Advanced
Part VI
Efficient training and inference
How generation runs, what attention costs in memory through the key–value cache, and how serving systems schedule prefill and decode; then the methods that make it cheaper, each with the error it introduces: quantization, pruning, distillation and sparse mixtures of experts. The part continues into long contexts and FlashAttention, and ends with speculative and structured decoding.
- VI.1EfficiencyFour moves make models cheaper at different costs: fewer bits per number, a smaller student, a low-rank update or sparse experts.Advanced
- VI.2Decoding & SamplingTemperature rescales logits before the softmax; top-p samples from the smallest set of tokens whose total probability reaches p.Intermediate
- VI.3Efficient Attention at ScaleDecoding reads past keys and values from a cache; grouped-query attention shrinks it by letting query heads share key–value heads.Advanced
- VI.4LLM Serving at ScaleServing splits into a compute-bound prefill of the prompt and a memory-bound decode, one token at a time, scheduled in batches.Advanced
- VI.5QuantizationQuantization stores weights and activations as low-bit integers with a scale, saving memory traffic at the cost of rounding error.Intermediate
- VI.6PruningPruning zeroes single weights or removes whole channels and heads; only the latter makes the dense tensors themselves smaller.Intermediate
- VI.7Knowledge DistillationDistillation trains a smaller student to match a teacher’s softened output distribution, which says more than the label alone.Intermediate
- VI.8Sparse Mixture of ExpertsA router sends each token to a few experts, so total parameters can grow while the computation spent on each token stays small.Advanced
- VI.9Long Context EngineeringLong context meets two limits: positions beyond those seen in training, and a key–value cache whose memory grows with each token.Advanced
- VI.10FlashAttentionComputes exact attention in tiles held in fast memory, trading extra arithmetic for far fewer reads and writes of slow memory.Advanced
- VI.11Speculative DecodingA small model drafts tokens, the large model checks them in parallel, and rejection sampling keeps the output distribution exact.Advanced
- VI.12Structured DecodingMasking every token that cannot lead to a valid completion guarantees well-formed output, though not that its content is true.Advanced
- VI.13MoE Serving & SchedulingServing sparse experts is a scheduling problem: uneven routing creates stragglers, and moving tokens between devices costs time.Advanced
Guided practice
Attention and serving — A guided route from attention through the cache to batched serving.
Part VII
Alignment, reasoning and interpretability
Reward models and reinforcement learning from human feedback, direct preference methods that skip the reward model, and the ways a proxy reward can be over-optimised; then step-level verifiers and the search over reasoning steps they guide. Sparse autoencoders close the part by asking what a network represents internally.
- VII.1RLHFRLHF fits a reward model to human comparisons, then moves the policy towards high reward with a KL penalty to a reference.Advanced
- VII.2Direct Preference OptimizationDPO fits preference pairs with a logistic loss on log-ratios to a reference model, without a separate reward model or an RL loop.Advanced
- VII.3Kahneman-Tversky OptimizationKTO learns from unpaired outputs labelled good or bad, pushing each implied reward above or below a reference point.Advanced
- VII.4Reward HackingOptimising hard against a flawed reward model moves probability onto its errors, so proxy reward can rise as true quality falls.Advanced
- VII.5Process Reward ModelsA process reward model scores each reasoning step, not only the final answer: denser feedback that is still a learned proxy.Advanced
- VII.6Tree Search ReasoningTree search spends inference budget on partial reasoning, expanding the prefixes a step-level verifier scores as most promising.Advanced
- VII.7Sparse AutoencodersA sparse autoencoder approximates an activation with a few directions from a learned dictionary, to separate superposed features.Advanced
Part VIII
Research
Each part ends at something that is not yet settled. Research here starts from such a question, written down with what would count against it before any study is run.
Reference
- Index of concepts — search every notebook, reference note and source by name.
- Mathematical reference — short entries on the definitions the notebooks rely on.
- Learning paths — shorter routes through the curriculum toward one goal.