Lab · Training

Train a language model in your browser

A transformer begins with random weights and no knowledge of words. Give it a passage, ask it to predict the next character, and correct its errors. Here you can watch that process, read the text it produces, and keep a record of the run. Training takes place on your device, without an account or a compute bill.

A small transformer, trained here

in your browser · $0

Choose a recipe, then let the model read an excerpt of Jane Austen’s Pride and Prejudice. The model predicts one character at a time. Your device does the computation.

Training recipe
2 blocks · 64-character context
Logarithmic scale · 0.0001–0.01
25–1,000 · 10-minute limit
Same recipe, same random draws
Ready

Start with random weights. Each update asks the model to predict the next character in a passage it has just read.

Step
0 / 400
Training loss
—
Validation loss
—
Characters / second
—
Elapsed
0 s
Training remaining
—
Learning to predict. Cross-entropy in nats per character; lower is better.
Training and validation loss by update stepSolid line: mean training loss since the last checkpoint. Dashed line: four fixed held-out windows. Uniform guessing has loss 4.277 nats. The latest measurements and a data table follow.3.03.74.35.0uniform guessing0update step400
TrainingValidation

Validation uses four fixed windows from the held-out final 10%, so it is a small estimate. Training loss is the mean since the previous checkpoint. Speed and time remaining exclude evaluation and text sampling.

A sample of its predictions

Temperature divides the model’s scores for each character (its logits) before one is drawn. Lower values favor likely characters; higher values allow more variation.

A sample will appear here before the first update, then every 100 steps and at the final checkpoint.

This is generated text, not a quotation from the book.

Read the measurements as a table
Recorded checkpoints; loss in nats per character
StepTrainValidationChar / s

Run record

The record keeps the recipe, the hash of the training text, the seed, the cost and the measurements together. Download it before you leave: this tab does not save the model or the record. You can read a downloaded record back in the run-record viewer.

Parameters
113,480
Where it runs
Your browser, in a background worker
Cost
$0
Record status
No run yet
Recipe, data and reproducibility

A decoder-only transformer: token and learned position embeddings, 2 pre-norm blocks, final layer normalization, and an untied output head. Each block has causal multi-head attention and an MLP with the tanh approximation to GELU. All weights train; there is no dropout.

AdamW uses β₁ = 0.9, β₂ = 0.95, ε = 10⁻⁸, matrix-only weight decay of 0.1, and global gradient clipping at norm 1. The learning rate warms up linearly, then follows a cosine to 10% of its peak. Float32 arithmetic may vary slightly across browser engines; the seed fixes the random draws.

Pride and Prejudice, by Jane Austen · Public domain in the USA. The 399,800-byte ASCII excerpt has 72 symbols. The first 90% is training text; the final 10% is held out. No training or validation window spans the split. Corpus provenance · Download text.

Pressing Start records your approval of this free run in the record. The approval is not linked to an account and is not checked by anyone else. Samples never feed back into training; the latest sample, its time and its temperature are saved in the record’s summary. After training ends, sampling uses the final weights and does not train further.

Training text SHA-256
c7754e51600e0c77c312a7b5fb160957452eeda555cdf8ff8110a10b1f7c0f7e
Specification SHA-256 (recipe, data, seed and budget)
Computed when you start

What you just did

For a sequence of characters, the model assigns a probability to the next one. The loss is the mean negative natural logarithm of the probability assigned to the observed character. Each update follows the gradient of that loss. A lower validation loss means better predictions on the sampled held-out passages; it does not by itself establish reasoning, factual knowledge, or useful language generation.

The same mathematical objects appear in the curriculum:

Try the same recipe with another seed. Does the held-out loss reach a similar value, and does the sample tell the same story?