Lab · Training
Train a language model in your browser
A transformer begins with random weights and no knowledge of words. Give it a passage, ask it to predict the next character, and correct its errors. Here you can watch that process, read the text it produces, and keep a record of the run. Training takes place on your device, without an account or a compute bill.
A small transformer, trained here
in your browser · $0Choose a recipe, then let the model read an excerpt of Jane Austen’s Pride and Prejudice. The model predicts one character at a time. Your device does the computation.
Start with random weights. Each update asks the model to predict the next character in a passage it has just read.
- Step
- 0 / 400
- Training loss
- —
- Validation loss
- —
- Characters / second
- —
- Elapsed
- 0 s
- Training remaining
- —
Validation uses four fixed windows from the held-out final 10%, so it is a small estimate. Training loss is the mean since the previous checkpoint. Speed and time remaining exclude evaluation and text sampling.
A sample of its predictions
Temperature divides the model’s scores for each character (its logits) before one is drawn. Lower values favor likely characters; higher values allow more variation.
A sample will appear here before the first update, then every 100 steps and at the final checkpoint.
This is generated text, not a quotation from the book.
Read the measurements as a table
| Step | Train | Validation | Char / s |
|---|
Run record
The record keeps the recipe, the hash of the training text, the seed, the cost and the measurements together. Download it before you leave: this tab does not save the model or the record. You can read a downloaded record back in the run-record viewer.
- Parameters
- 113,480
- Where it runs
- Your browser, in a background worker
- Cost
- $0
- Record status
- No run yet
Recipe, data and reproducibility
A decoder-only transformer: token and learned position embeddings, 2 pre-norm blocks, final layer normalization, and an untied output head. Each block has causal multi-head attention and an MLP with the tanh approximation to GELU. All weights train; there is no dropout.
AdamW uses β₁ = 0.9, β₂ = 0.95, ε = 10⁻⁸, matrix-only weight decay of 0.1, and global gradient clipping at norm 1. The learning rate warms up linearly, then follows a cosine to 10% of its peak. Float32 arithmetic may vary slightly across browser engines; the seed fixes the random draws.
Pride and Prejudice, by Jane Austen · Public domain in the USA. The 399,800-byte ASCII excerpt has 72 symbols. The first 90% is training text; the final 10% is held out. No training or validation window spans the split. Corpus provenance · Download text.
Pressing Start records your approval of this free run in the record. The approval is not linked to an account and is not checked by anyone else. Samples never feed back into training; the latest sample, its time and its temperature are saved in the record’s summary. After training ends, sampling uses the final weights and does not train further.
- Training text SHA-256
- c7754e51600e0c77c312a7b5fb160957452eeda555cdf8ff8110a10b1f7c0f7e
- Specification SHA-256 (recipe, data, seed and budget)
- Computed when you start
What you just did
For a sequence of characters, the model assigns a probability to the next one. The loss is the mean negative natural logarithm of the probability assigned to the observed character. Each update follows the gradient of that loss. A lower validation loss means better predictions on the sampled held-out passages; it does not by itself establish reasoning, factual knowledge, or useful language generation.
The same mathematical objects appear in the curriculum:
- Tokenization gives each character an index; causal attention lets it read earlier characters.
- Layer normalization rescales each token’s features before attention and the MLP.
- Gradient descent supplies the direction; Adam adapts its scale, with decoupled weight decay in this run.
- Learning-rate schedules control the size of updates as training proceeds.
Try the same recipe with another seed. Does the held-out loss reach a similar value, and does the sample tell the same story?