Before opening a frontier page, predict whether the blocker is notation, optimization, probability, representation, scaling, or systems behavior.
Foundations Atlas
Mathematical Foundations
A broad atlas of the mathematical objects behind modern AI: prerequisites, demos, papers, and routes into the newer domain notebooks.
Atlas Instrument
Start with the map when the prerequisite is unclear.
What path gets me from a confusing modern AI mechanism to the smaller mathematical object that makes it understandable?
The map shows dependency edges, papers ground the object, and demos give a local witness when the idea needs to be manipulated.
A foundation earns its place when the same object explains several models, papers, or engineering tradeoffs.
Survey the graph for neighbors, follow the ordered route when sequence matters, or filter for demos when intuition needs a witness.
Recommended Study Order
Build understanding from fundamentals to frontier techniques. Each phase builds on the previous one.
Optimization & generalization
Generative modeling families
Representation & interpretability
Modern efficiency & inference
Alignment & RLHF
Scaling, theory & multimodal
Advanced architectures & generation
Mathematical foundations & information geometry
Frontier research & scaling
Advanced alignment & safety research
Browse the Atlas
Search by concept, equation, source, or runnable witness.
ML/CE/KL
Maximum Likelihood, Cross-Entropy & KL Divergence
Explore →Attention
Scaled Dot-Product Attention & Transformer Layers
Explore →Adam
Adam & Adaptive Gradient Methods
Explore →Sharpness
Loss Landscapes, Sharpness & Flat Minima
Explore →Double Descent
Overparameterization & Generalization, Double Descent
Explore →NTK
Neural Tangent Kernel & Infinite-Width Limits
Explore →VAEs
Variational Autoencoders & Variational Inference
Explore →GANs
GANs & Adversarial Divergence Minimization
Explore →Diffusion
Diffusion, Score-Based Models & Flow Matching
Explore →Embeddings
Representation Learning & Embedding Geometry
Explore →Superposition
Superposition, Sparse Features & Monosemanticity
Explore →Probing
Probing, Linear Classifier Probes & Activation Analysis
Explore →Circuits
Transformer Circuits, Induction Heads & Mechanistic Interpretability
Explore →Scaling
Scaling Laws & Emergent Abilities
Explore →RLHF
Preference-Based Alignment: RLHF, Reward Modeling, Constitutional AI
Explore →Efficiency
Efficiency: Quantization, Distillation, LoRA & Sparse MoE
Explore →Theory
Theoretical Foundations: PAC Learning, MDL & Information Bottleneck
Explore →Efficient Attention
Efficient Attention at Scale: KV Cache, GQA & FlashAttention
Explore →RoPE
Rotary Position Embeddings (RoPE)
Explore →Speculative Decoding
Speculative Decoding: Lossless Multi-Token Generation
Explore →LLM Serving
LLM Serving at Scale: Prefill, Decode & Continuous Batching
Explore →MoE
Sparse Mixture of Experts: Routing, Load Balancing & Expert Parallelism
Explore →MoE Serving
MoE Serving & Scheduling: Token Dispatch, All-to-All, Disaggregated Parallelism
Explore →DPO
Direct Preference Optimization: RL-Free Alignment from Human Preferences
Explore →KTO
KTO: Alignment from Binary Feedback via Human-Aware Losses
Explore →Reward Hacking
Reward Hacking & Overoptimization: Goodhart's Law in Preference Optimization
Explore →Sparse Autoencoders
Sparse Autoencoders at Scale: Feature Dictionaries for Mechanistic Interpretability
Explore →Circuit Discovery
Automated Circuit Discovery: Patching, Attribution & Decomposition at Scale
Explore →Activation Steering
Activation Steering: Feature-Guided Interventions for Inference-Time Control
Explore →Long Context
Long Context Engineering: RoPE Scaling, KV Compression & Memory Optimization
Explore →SSMs & Hybrids
State Space Models & Hybrid Architectures: Mamba-2, Jamba, Griffin
Explore →Multimodal VLP
Multimodal Foundations: Vision Encoders, Contrastive Learning & Cross-Attention Fusion
Explore →Tokens
Tokenization & Vocabulary Design
Explore →Decoding
Decoding & Sampling: Temperature, Top-p & Inference-Time Control
Explore →Backprop
Backpropagation & Automatic Differentiation
Explore →Score Matching
Score Matching & Score-Based Generative Models
Explore →ICL
In-Context Learning: Learning Without Weight Updates
Explore →OT/Wasserstein
Optimal Transport & Wasserstein Distance
Explore →Flows
Normalizing Flows: Exact Likelihood via Invertible Transforms
Explore →PPO
PPO: Proximal Policy Optimization
Explore →Residuals
Residual Connections & Skip Connections
Explore →CFG
Classifier-Free Guidance in Diffusion
Explore →RAG
Retrieval-Augmented Generation (RAG)
Explore →Adversarial
Adversarial Examples & Robustness
Explore →Grokking
Grokking: Delayed Generalization
Explore →Logit Lens
Logit Lens: Probing Intermediate Representations
Explore →LR Schedules
Learning Rate Schedules: Warmup, Decay & Cycling
Explore →Init
Weight Initialization: Xavier, He & µP
Explore →Contrastive
Contrastive Learning & InfoNCE
Explore →Distributed
Distributed Training: Data, Tensor & Pipeline Parallelism
Explore →Beam Search
Beam Search & Structured Decoding
Explore →Dropout
Dropout: Stochastic Regularization
Explore →EBMs
Energy-Based Models & Score Functions
Explore →LayerNorm
Layer Normalization & RMSNorm
Explore →Fisher Info
Fisher Information & Information Geometry
Explore →Natural Grad
Natural Gradient & Riemannian Optimization
Explore →SGD+Momentum
SGD & Momentum: The Workhorses of Optimization
Explore →AdamW
Weight Decay & AdamW: Decoupled Regularization
Explore →Grad Clip
Gradient Clipping & Explosion Prevention
Explore →Label Smooth
Label Smoothing & Soft Targets
Explore →BatchNorm
Batch Normalization
Explore →Distillation
Knowledge Distillation: Learning from Teachers
Explore →Quantization
Quantization: Compressing Models to Integers
Explore →Pruning
Pruning: Removing Unnecessary Weights
Explore →SSL
Self-Supervised Learning: Labels from Structure
Explore →Calibration
Calibration & Temperature Scaling
Explore →SwiGLU
SwiGLU & Gated Activations
Explore →FlashAttn
FlashAttention: IO-Aware Attention
Explore →Constitutional
Constitutional AI: Principles-Based Alignment
Explore →Bregman
Bregman Divergence & Mirror Descent
Explore →RKHS
Reproducing Kernel Hilbert Spaces
Explore →TDA
Persistent Homology & Topological Data Analysis
Explore →Lie Groups
Lie Groups & Equivariant Networks
Explore →Test-Time
Test-Time Compute & Inference Scaling
Explore →CoT
Chain-of-Thought Prompting
Explore →World Models
World Models & Model-Based RL
Explore →Synth Data
Synthetic Data & Self-Improvement
Explore →Consistency
Consistency Models: One-Step Diffusion
Explore →Checkpointing
Activation Checkpointing & Memory Efficiency
Explore →GQA
Grouped Query Attention (GQA)
Explore →PRMs
Process Reward Models
Explore →RLAIF
RLAIF: AI Feedback
Explore →Flow Match
Flow Matching & Rectified Flows
Explore →Instruct
Instruction Tuning
Explore →Deliberative
Deliberative Alignment
Explore →Debate
AI Safety via Debate
Explore →IDA
Iterated Amplification
Explore →Weak→Strong
Weak-to-Strong Generalization
Explore →Auto RedTeam
Automated Red Teaming
Explore →Mesa-Opt
Mesa-Optimization & Inner Alignment
Explore →Sleepers
Sleeper Agents & Alignment Faking
Explore →LLM-as-Judge
Model-Graded Evaluations
Explore →Elicitation
Capability Elicitation & ELK
Explore →Sandwich
Sandwiching Evaluations
Explore →MoD
Mixture-of-Depths
Explore →MCTS-LLM
Tree Search over Thoughts
Explore →VideoWM
Video World Models
Explore →Self-Improve
Self-Improvement & Distillation Loops
Explore →Collapse
Model Collapse & Synthetic Data
Explore →InfCtx
Infinite Context Architectures
Explore →Editorial Contract
Intuition First
Each object should explain the felt problem before adding symbols.
Source Grounded
Canonical papers and equations keep broad navigation tied to evidence.
Runnable When Possible
Demo-bearing concepts are treated as local witnesses, not decoration.