FoundationsChecking saved investigationReading browser-local route memory before showing a continuation.

Foundations Atlas

Mathematical Foundations

A broad atlas of the mathematical objects behind modern AI: prerequisites, demos, papers, and routes into the newer domain notebooks.

100 concepts37 demos336 connections13 learning phases
1Core
2Optim
3Gen
4Rep
5Scale
6Systems
lab surfacepredict, drag, test

Atlas Instrument

Start with the map when the prerequisite is unclear.

Selected object100-concept prerequisite atlas37 runnable demos, 336 mapped edges, 13 ordered phases
Current question

What path gets me from a confusing modern AI mechanism to the smaller mathematical object that makes it understandable?

PredictionName the likely gap first.

Before opening a frontier page, predict whether the blocker is notation, optimization, probability, representation, scaling, or systems behavior.

EvidenceUse graph, paper, and demo together.

The map shows dependency edges, papers ground the object, and demos give a local witness when the idea needs to be manipulated.

InvariantCarry the mechanism forward.

A foundation earns its place when the same object explains several models, papers, or engineering tradeoffs.

Next moveChoose map, path, or lab.

Survey the graph for neighbors, follow the ordered route when sequence matters, or filter for demos when intuition needs a witness.

Recommended Study Order

Build understanding from fundamentals to frontier techniques. Each phase builds on the previous one.

1

Core probabilistic training + transformers

9

Advanced architectures & generation

11

Frontier research & scaling

Browse the Atlas

Search by concept, equation, source, or runnable witness.

1

ML/CE/KL

Maximum Likelihood, Cross-Entropy & KL Divergence

Core Training261 paper▶ Demo
Explore →
2

Attention

Scaled Dot-Product Attention & Transformer Layers

Core Training231 paper▶ Demo
Explore →
3

Adam

Adam & Adaptive Gradient Methods

Optimization🔗 42 papers▶ Demo
Explore →
4

Sharpness

Loss Landscapes, Sharpness & Flat Minima

Optimization112 papers▶ Demo
Explore →
5

Double Descent

Overparameterization & Generalization, Double Descent

Optimization🔗 42 papers▶ Demo
Explore →
6Θ

NTK

Neural Tangent Kernel & Infinite-Width Limits

Theory🔗 41 paper▶ Demo
Explore →
7

VAEs

Variational Autoencoders & Variational Inference

Generative Models🔗 51 paper▶ Demo
Explore →
8

GANs

GANs & Adversarial Divergence Minimization

Generative Models🔗 22 papers▶ Demo
Explore →
9

Diffusion

Diffusion, Score-Based Models & Flow Matching

Generative Models114 papers▶ Demo
Explore →
10

Embeddings

Representation Learning & Embedding Geometry

Representations141 paper▶ Demo
Explore →
11

Superposition

Superposition, Sparse Features & Monosemanticity

Representations🔗 32 papers▶ Demo
Explore →
12

Probing

Probing, Linear Classifier Probes & Activation Analysis

Representations🔗 52 papers▶ Demo
Explore →
13

Circuits

Transformer Circuits, Induction Heads & Mechanistic Interpretability

Representations🔗 72 papers▶ Demo
Explore →
14

Scaling

Scaling Laws & Emergent Abilities

Scaling & Alignment🔗 73 papers▶ Demo
Explore →
15

RLHF

Preference-Based Alignment: RLHF, Reward Modeling, Constitutional AI

Scaling & Alignment173 papers▶ Demo
Explore →
16

Efficiency

Efficiency: Quantization, Distillation, LoRA & Sparse MoE

Efficiency🔗 93 papers▶ Demo
Explore →
17

Theory

Theoretical Foundations: PAC Learning, MDL & Information Bottleneck

Theory🔗 93 papers▶ Demo
Explore →
19

Efficient Attention

Efficient Attention at Scale: KV Cache, GQA & FlashAttention

Efficiency113 papers▶ Demo
Explore →
18

RoPE

Rotary Position Embeddings (RoPE)

Representations🔗 33 papers▶ Demo
Explore →
20

Speculative Decoding

Speculative Decoding: Lossless Multi-Token Generation

Efficiency🔗 63 papers▶ Demo
Explore →
21

LLM Serving

LLM Serving at Scale: Prefill, Decode & Continuous Batching

Efficiency🔗 73 papers▶ Demo
Explore →
22🔀

MoE

Sparse Mixture of Experts: Routing, Load Balancing & Expert Parallelism

Efficiency🔗 83 papers▶ Demo
Explore →
23

MoE Serving

MoE Serving & Scheduling: Token Dispatch, All-to-All, Disaggregated Parallelism

Efficiency🔗 43 papers▶ Demo
Explore →
24🎯

DPO

Direct Preference Optimization: RL-Free Alignment from Human Preferences

Scaling & Alignment🔗 53 papers▶ Demo
Explore →
25👍

KTO

KTO: Alignment from Binary Feedback via Human-Aware Losses

Scaling & Alignment🔗 33 papers▶ Demo
Explore →
26⚠️

Reward Hacking

Reward Hacking & Overoptimization: Goodhart's Law in Preference Optimization

Scaling & Alignment🔗 33 papers▶ Demo
Explore →
27🔍

Sparse Autoencoders

Sparse Autoencoders at Scale: Feature Dictionaries for Mechanistic Interpretability

Representations🔗 63 papers▶ Demo
Explore →
28🔬

Circuit Discovery

Automated Circuit Discovery: Patching, Attribution & Decomposition at Scale

Representations🔗 43 papers▶ Demo
Explore →
29🎚️

Activation Steering

Activation Steering: Feature-Guided Interventions for Inference-Time Control

Representations🔗 43 papers▶ Demo
Explore →
30📏

Long Context

Long Context Engineering: RoPE Scaling, KV Compression & Memory Optimization

Efficiency103 papers▶ Demo
Explore →
31🔀

SSMs & Hybrids

State Space Models & Hybrid Architectures: Mamba-2, Jamba, Griffin

Core Training🔗 53 papers▶ Demo
Explore →
32🖼️

Multimodal VLP

Multimodal Foundations: Vision Encoders, Contrastive Learning & Cross-Attention Fusion

Representations🔗 53 papers▶ Demo
Explore →
33🔤

Tokens

Tokenization & Vocabulary Design

Representations🔗 43 papers▶ Demo
Explore →
34🎲

Decoding

Decoding & Sampling: Temperature, Top-p & Inference-Time Control

Core Training🔗 53 papers▶ Demo
Explore →
35

Backprop

Backpropagation & Automatic Differentiation

Optimization🔗 42 papers
Explore →
36∇p

Score Matching

Score Matching & Score-Based Generative Models

Generative Models🔗 52 papers
Explore →
37📝

ICL

In-Context Learning: Learning Without Weight Updates

Representations🔗 52 papers
Explore →
38⚖️

OT/Wasserstein

Optimal Transport & Wasserstein Distance

Theory🔗 32 papers
Explore →
39🌊

Flows

Normalizing Flows: Exact Likelihood via Invertible Transforms

Generative Models🔗 32 papers
Explore →
40🎯

PPO

PPO: Proximal Policy Optimization

Scaling & Alignment🔗 11 paper
Explore →
41

Residuals

Residual Connections & Skip Connections

Core Training🔗 21 paper
Explore →
42🎚️

CFG

Classifier-Free Guidance in Diffusion

Generative Models🔗 21 paper
Explore →
43🔍

RAG

Retrieval-Augmented Generation (RAG)

Representations🔗 31 paper
Explore →
44⚔️

Adversarial

Adversarial Examples & Robustness

Theory🔗 31 paper
Explore →
45💡

Grokking

Grokking: Delayed Generalization

Theory🔗 31 paper
Explore →
46🔬

Logit Lens

Logit Lens: Probing Intermediate Representations

Representations🔗 41 paper
Explore →
47📉

LR Schedules

Learning Rate Schedules: Warmup, Decay & Cycling

Optimization🔗 21 paper
Explore →
48🎲

Init

Weight Initialization: Xavier, He & µP

Optimization🔗 32 papers
Explore →
49🔗

Contrastive

Contrastive Learning & InfoNCE

Representations🔗 22 papers
Explore →
50🌐

Distributed

Distributed Training: Data, Tensor & Pipeline Parallelism

Efficiency🔗 32 papers
Explore →
51🌳

Beam Search

Beam Search & Structured Decoding

Core Training🔗 21 paper
Explore →
52💧

Dropout

Dropout: Stochastic Regularization

Optimization🔗 21 paper
Explore →
53

EBMs

Energy-Based Models & Score Functions

Generative Models🔗 21 paper
Explore →
54📏

LayerNorm

Layer Normalization & RMSNorm

Core Training🔗 22 papers
Explore →
55

Fisher Info

Fisher Information & Information Geometry

Theory🔗 42 papers
Explore →
56🧭

Natural Grad

Natural Gradient & Riemannian Optimization

Optimization🔗 32 papers
Explore →
57🏃

SGD+Momentum

SGD & Momentum: The Workhorses of Optimization

Optimization🔗 22 papers
Explore →
58⚖️

AdamW

Weight Decay & AdamW: Decoupled Regularization

Optimization🔗 22 papers
Explore →
59✂️

Grad Clip

Gradient Clipping & Explosion Prevention

Optimization🔗 22 papers
Explore →
60🎯

Label Smooth

Label Smoothing & Soft Targets

Optimization🔗 22 papers
Explore →
61📊

BatchNorm

Batch Normalization

Core Training🔗 12 papers
Explore →
62🧪

Distillation

Knowledge Distillation: Learning from Teachers

Efficiency🔗 62 papers
Explore →
63🔢

Quantization

Quantization: Compressing Models to Integers

Efficiency🔗 22 papers
Explore →
64✂️

Pruning

Pruning: Removing Unnecessary Weights

Efficiency🔗 22 papers
Explore →
65🔄

SSL

Self-Supervised Learning: Labels from Structure

Representations🔗 32 papers
Explore →
66🌡️

Calibration

Calibration & Temperature Scaling

Theory🔗 32 papers
Explore →
67🚪

SwiGLU

SwiGLU & Gated Activations

Core Training🔗 12 papers▶ Demo
Explore →
68

FlashAttn

FlashAttention: IO-Aware Attention

Efficiency🔗 22 papers
Explore →
69📜

Constitutional

Constitutional AI: Principles-Based Alignment

Scaling & Alignment🔗 62 papers
Explore →
70🪞

Bregman

Bregman Divergence & Mirror Descent

Theory🔗 22 papers
Explore →
71🎭

RKHS

Reproducing Kernel Hilbert Spaces

Theory🔗 22 papers
Explore →
72🕳️

TDA

Persistent Homology & Topological Data Analysis

Theory🔗 22 papers
Explore →
73🔀

Lie Groups

Lie Groups & Equivariant Networks

Theory🔗 22 papers
Explore →
74🧠

Test-Time

Test-Time Compute & Inference Scaling

Scaling & Alignment🔗 52 papers
Explore →
75💭

CoT

Chain-of-Thought Prompting

Scaling & Alignment🔗 42 papers
Explore →
76🌍

World Models

World Models & Model-Based RL

Theory🔗 22 papers
Explore →
77🏭

Synth Data

Synthetic Data & Self-Improvement

Scaling & Alignment🔗 22 papers
Explore →
78🎯

Consistency

Consistency Models: One-Step Diffusion

Generative Models🔗 22 papers
Explore →
79💾

Checkpointing

Activation Checkpointing & Memory Efficiency

Efficiency🔗 22 papers
Explore →
80👥

GQA

Grouped Query Attention (GQA)

Efficiency🔗 21 paper▶ Demo
Explore →
81📋

PRMs

Process Reward Models

Scaling & Alignment🔗 21 paper
Explore →
82🤖

RLAIF

RLAIF: AI Feedback

Scaling & Alignment🔗 21 paper
Explore →
83🌊

Flow Match

Flow Matching & Rectified Flows

Generative Models🔗 21 paper
Explore →
84📝

Instruct

Instruction Tuning

Scaling & Alignment🔗 21 paper
Explore →
85🎯

Deliberative

Deliberative Alignment

Scaling & Alignment🔗 21 paper
Explore →
86⚔️

Debate

AI Safety via Debate

Scaling & Alignment🔗 11 paper
Explore →
87🔄

IDA

Iterated Amplification

Scaling & Alignment🔗 31 paper
Explore →
88💪

Weak→Strong

Weak-to-Strong Generalization

Scaling & Alignment🔗 21 paper
Explore →
89🔴

Auto RedTeam

Automated Red Teaming

Scaling & Alignment🔗 21 paper
Explore →
90🔮

Mesa-Opt

Mesa-Optimization & Inner Alignment

Scaling & Alignment🔗 21 paper
Explore →
91😴

Sleepers

Sleeper Agents & Alignment Faking

Scaling & Alignment🔗 21 paper
Explore →
92⚖️

LLM-as-Judge

Model-Graded Evaluations

Scaling & Alignment🔗 21 paper
Explore →
93🔍

Elicitation

Capability Elicitation & ELK

Scaling & Alignment🔗 21 paper
Explore →
94🥪

Sandwich

Sandwiching Evaluations

Scaling & Alignment🔗 21 paper
Explore →
95📊

MoD

Mixture-of-Depths

Efficiency🔗 21 paper
Explore →
96🌲

MCTS-LLM

Tree Search over Thoughts

Scaling & Alignment🔗 21 paper
Explore →
97🎬

VideoWM

Video World Models

Generative Models🔗 21 paper
Explore →
98🔄

Self-Improve

Self-Improvement & Distillation Loops

Scaling & Alignment🔗 21 paper
Explore →
99📉

Collapse

Model Collapse & Synthetic Data

Theory🔗 21 paper
Explore →
100♾️

InfCtx

Infinite Context Architectures

Efficiency🔗 21 paper
Explore →

Editorial Contract

Intuition First

Each object should explain the felt problem before adding symbols.

Source Grounded

Canonical papers and equations keep broad navigation tied to evidence.

Runnable When Possible

Demo-bearing concepts are treated as local witnesses, not decoration.