The Atlas

Everything readable here, in one register.

Concept notes grouped by mathematical territory, the foundations underneath them, and five long-running areas that connect them. Search first if you know what you want; otherwise the whole corpus is below.

Notes
51 published
Topics
12
Foundations
100

Linear Algebra

Vectors, matrices, and linear maps: the language of representations, optimization, and modern deep learning. Open the topic.

Calculus

Rates of change and accumulation. Calculus is the language behind gradients, optimization, continuous-time dynamics, and why backprop works as efficiently as it does. Open the topic.

Optimization

How we train models: gradients, learning rates, curvature, and the practical tricks that make deep nets converge. Open the topic.

Probability

Uncertainty made precise: events, random variables, expectations, and the distributions that models learn. Open the topic.

Information Theory

How we measure information and mismatch between distributions: entropy, cross-entropy, KL divergence, mutual information, and why they appear everywhere in ML. Open the topic.

Attention & Transformers

The sequence model backbone: tokenization, self-attention, positional encodings, and the transformer block that powers modern LLMs. Open the topic.

Representation Learning

Embeddings and the geometry of meaning: similarity, normalization, contrastive objectives, and why vector spaces become usable interfaces for models. Open the topic.

Generative Models

How models generate: likelihood, latent variables, diffusion/score models, flows, and the training tricks that make sampling work. Open the topic.

Scaling

How loss and capability change with parameters, data, and compute; how to allocate a training budget; and why some abilities appear suddenly at scale. Open the topic.

Alignment

How we shape model behavior: preference learning, reward modeling, KL-regularized fine-tuning, and the failure modes that appear when you optimize the wrong thing. Open the topic.

Efficiency

How we make models cheaper to train and serve: quantization, distillation, low-rank adapters, sparsity, and the memory/latency tradeoffs that dominate real deployments. Open the topic.

LLM Systems

How models run in production: prefill vs decode, KV cache memory, batching and scheduling, and the techniques that make latency and throughput practical. Open the topic.

Foundations

The mathematical objects underneath the models. Also readable as a connected map.

Maximum Likelihood, Cross-Entropy & KL DivergenceCore TrainingScaled Dot-Product Attention & Transformer LayersCore TrainingAdam & Adaptive Gradient MethodsOptimizationLoss Landscapes, Sharpness & Flat MinimaOptimizationOverparameterization & Generalization, Double DescentOptimizationNeural Tangent Kernel & Infinite-Width LimitsTheoryVariational Autoencoders & Variational InferenceGenerative ModelsGANs & Adversarial Divergence MinimizationGenerative ModelsDiffusion, Score-Based Models & Flow MatchingGenerative ModelsRepresentation Learning & Embedding GeometryRepresentationsSuperposition, Sparse Features & MonosemanticityRepresentationsProbing, Linear Classifier Probes & Activation AnalysisRepresentationsTransformer Circuits, Induction Heads & Mechanistic InterpretabilityRepresentationsScaling Laws & Emergent AbilitiesScaling & AlignmentPreference-Based Alignment: RLHF, Reward Modeling, Constitutional AIScaling & AlignmentEfficiency: Quantization, Distillation, LoRA & Sparse MoEEfficiencyTheoretical Foundations: PAC Learning, MDL & Information BottleneckTheoryEfficient Attention at Scale: KV Cache, GQA & FlashAttentionEfficiencyRotary Position Embeddings (RoPE)RepresentationsSpeculative Decoding: Lossless Multi-Token GenerationEfficiencyLLM Serving at Scale: Prefill, Decode & Continuous BatchingEfficiencySparse Mixture of Experts: Routing, Load Balancing & Expert ParallelismEfficiencyMoE Serving & Scheduling: Token Dispatch, All-to-All, Disaggregated ParallelismEfficiencyDirect Preference Optimization: RL-Free Alignment from Human PreferencesScaling & AlignmentKTO: Alignment from Binary Feedback via Human-Aware LossesScaling & AlignmentReward Hacking & Overoptimization: Goodhart's Law in Preference OptimizationScaling & AlignmentSparse Autoencoders at Scale: Feature Dictionaries for Mechanistic InterpretabilityRepresentationsAutomated Circuit Discovery: Patching, Attribution & Decomposition at ScaleRepresentationsActivation Steering: Feature-Guided Interventions for Inference-Time ControlRepresentationsLong Context Engineering: RoPE Scaling, KV Compression & Memory OptimizationEfficiencyState Space Models & Hybrid Architectures: Mamba-2, Jamba, GriffinCore TrainingMultimodal Foundations: Vision Encoders, Contrastive Learning & Cross-Attention FusionRepresentationsTokenization & Vocabulary DesignRepresentationsDecoding & Sampling: Temperature, Top-p & Inference-Time ControlCore TrainingBackpropagation & Automatic DifferentiationOptimizationScore Matching & Score-Based Generative ModelsGenerative ModelsIn-Context Learning: Learning Without Weight UpdatesRepresentationsOptimal Transport & Wasserstein DistanceTheoryNormalizing Flows: Exact Likelihood via Invertible TransformsGenerative ModelsPPO: Proximal Policy OptimizationScaling & AlignmentResidual Connections & Skip ConnectionsCore TrainingClassifier-Free Guidance in DiffusionGenerative ModelsRetrieval-Augmented Generation (RAG)RepresentationsAdversarial Examples & RobustnessTheoryGrokking: Delayed GeneralizationTheoryLogit Lens: Probing Intermediate RepresentationsRepresentationsLearning Rate Schedules: Warmup, Decay & CyclingOptimizationWeight Initialization: Xavier, He & µPOptimizationContrastive Learning & InfoNCERepresentationsDistributed Training: Data, Tensor & Pipeline ParallelismEfficiencyBeam Search & Structured DecodingCore TrainingDropout: Stochastic RegularizationOptimizationEnergy-Based Models & Score FunctionsGenerative ModelsLayer Normalization & RMSNormCore TrainingFisher Information & Information GeometryTheoryNatural Gradient & Riemannian OptimizationOptimizationSGD & Momentum: The Workhorses of OptimizationOptimizationWeight Decay & AdamW: Decoupled RegularizationOptimizationGradient Clipping & Explosion PreventionOptimizationLabel Smoothing & Soft TargetsOptimizationBatch NormalizationCore TrainingKnowledge Distillation: Learning from TeachersEfficiencyQuantization: Compressing Models to IntegersEfficiencyPruning: Removing Unnecessary WeightsEfficiencySelf-Supervised Learning: Labels from StructureRepresentationsCalibration & Temperature ScalingTheorySwiGLU & Gated ActivationsCore TrainingFlashAttention: IO-Aware AttentionEfficiencyConstitutional AI: Principles-Based AlignmentScaling & AlignmentBregman Divergence & Mirror DescentTheoryReproducing Kernel Hilbert SpacesTheoryPersistent Homology & Topological Data AnalysisTheoryLie Groups & Equivariant NetworksTheoryTest-Time Compute & Inference ScalingScaling & AlignmentChain-of-Thought PromptingScaling & AlignmentWorld Models & Model-Based RLTheorySynthetic Data & Self-ImprovementScaling & AlignmentConsistency Models: One-Step DiffusionGenerative ModelsActivation Checkpointing & Memory EfficiencyEfficiencyGrouped Query Attention (GQA)EfficiencyProcess Reward ModelsScaling & AlignmentRLAIF: AI FeedbackScaling & AlignmentFlow Matching & Rectified FlowsGenerative ModelsInstruction TuningScaling & AlignmentDeliberative AlignmentScaling & AlignmentAI Safety via DebateScaling & AlignmentIterated AmplificationScaling & AlignmentWeak-to-Strong GeneralizationScaling & AlignmentAutomated Red TeamingScaling & AlignmentMesa-Optimization & Inner AlignmentScaling & AlignmentSleeper Agents & Alignment FakingScaling & AlignmentModel-Graded EvaluationsScaling & AlignmentCapability Elicitation & ELKScaling & AlignmentSandwiching EvaluationsScaling & AlignmentMixture-of-DepthsEfficiencyTree Search over ThoughtsScaling & AlignmentVideo World ModelsGenerative ModelsSelf-Improvement & Distillation LoopsScaling & AlignmentModel Collapse & Synthetic DataTheoryInfinite Context ArchitecturesEfficiency

Areas

Five long-running storylines that thread the topics together.