Open the black box

From scattered tokens to surprisingly fluent sentences, follow the hidden steps behind an AI answer. Play with the pieces, watch the patterns emerge, and discover how language models really work.

Your route through the model

Start with Tokens and follow the lessons in order. You do not need to code or know calculus: a vector is a list of numbers, and a probability describes how much of a distribution belongs to an outcome. Each lesson includes an experiment, a learning goal, and a question with feedback.

How the demos fit together

These are separate experiments illustrating connected ideas. They use different vocabularies, vectors, and architectures; one demo's output is not passed into the next. Within a real model, the tokenizer, embedding table, transformer, and output vocabulary must match.

Models and representations used by each demo
LessonDemoWhat it teaches
Tokenso200k_base byte-pair tokenizerText → vocabulary IDs
Embeddings139 GloVe word vectors, 50 dimensionsStatic word geometry
AttentionDistilBERT with WordPiece tokensBidirectional transformer attention
BackpropagationSmall network learning XORGradients → weight updates
Training & scalingCharacter MLP with 3 characters of contextLearning and held-out prediction
GenerationGPT-2 small with r50k_base tokensCausal prediction and sampling
A few words you will meet
Parameter / weight
A learned number that controls a model's computation.
Representation / activation
A value computed for the current input; it changes with context.
Logit
A raw score for a candidate, before conversion to probabilities.
Loss
A numerical penalty for a prediction; training tries to reduce its average.
Gradient
How loss changes with a small change in a parameter.
Context
The information supplied for this prediction, within the model's input limit.

Lesson 01 / 08

Tokens

Your goal: Explain why token IDs depend on a tokenizer and why token count differs from character count.

Before a model can do anything with text, it breaks that text into tokens: chunks that are often whole words, but just as often fragments, punctuation, or whitespace. This demo uses the o200k_base byte-pair encoding vocabulary; the GPT-2 demo later uses a different vocabulary, r50k_base.

Why tokenize at all?

A model doesn't read letters, and it can't keep a slot for every possible word, since real languages have millions, and people invent new ones constantly. So instead of words, the tokenizer works from a fixed vocabulary of subword chunks. Vocabulary sizes vary by model. Common words get their own token; anything rarer is stitched together from smaller known pieces. That way even a word the model has never seen, or one you just made up, can still be built out of chunks it already knows.

Try this: type strawberry, then step through 1 to 4 to watch the text become the plain list of numbers the model actually reads. Press Play sequence to run the four steps on their own.

Try:

Cut: the tokenizer slices the text into chunks that each exist in its fixed vocabulary.

3 tokens10 characters200,006-token vocabulary

Character counts include spaces and punctuation. Combined emoji and accented letters each count as one character.

strawberry

Each colored chunk is one token. Its id is its position in the fixed vocabulary, and a leading space is part of the token, drawn as · inside the chunk. The four steps take the same text you typed from an unbroken string, to chunks, to ids, to the bare numbers the model receives.

Why counting letters is hard for LLMs

Try it above: “strawberry” splits into st + raw + berry. The model never sees the individual letters; it sees those three whole chunks, so the count of r's isn't something its input directly exposes. Modern models often do get this right, by reasoning through it step-by-step or from spelling patterns they've memorized. It's not impossible, just not automatic, the way it would be if the model read raw letters. That difficulty is a direct consequence of tokenization.

Why token count matters in practice

A context window limits the tokens available to a model at once. Instructions, conversation history, retrieved text, and generated output all use token space. Many language-model APIs charge for input and output tokens, so token counts also help estimate cost. The tokenizer count here covers the text you enter; a chat application can add role markers and other framing tokens.

Same idea, different price

Here is one short sentence written five ways, each counted live with the same tokenizer. The bars show tokens, and the character count sits beside each one so you can see the contrast.

English12 tokens · 49 characters
The weather is lovely today, let's go for a walk.
Spanish12 tokens · 49 characters
El clima está precioso hoy, vamos a dar un paseo.
Japanese15 tokens · 21 characters
今日は天気がいいので、散歩に行きましょう。
Emoji10 tokens · 5 characters
☀️😊👟🌳🚶
Code14 tokens · 42 characters
if (weather === "lovely") { goForWalk(); }

Character counts include spaces and punctuation. Combined emoji and accented letters each count as one character.

Compare the actual counts for these examples. Language, script, spacing, and formatting affect which chunks exist in this vocabulary. Some Unicode characters span several byte tokens, while some tokens contain multiple characters. These examples do not establish a universal ranking of languages: the text and the tokenizer both matter.

Good to know: tokenizers differ by model

The splitter above uses one specific GPT-family encoding. Different models can use different tokenizers. The exact same text can come out to a different number of tokens depending on which model processes it, so a token count is always relative to a particular model family.

Check your understanding

Two tokenizers give the word cat different IDs. What can you infer?
Sources & further reading
Within a model, its vocabulary IDs index its matching embedding table. Next: explore the idea of learned vectors using separate word embeddings.→

Lesson 02 / 08

Embeddings

Your goal: Distinguish a token's learned input vector from its later representation in context.

Tokenization left us with a list of ids: plain integers, one per chunk of text. But a model can't do math on “token #4826” as if the number meant anything. The first thing it really does is swap each id for a vector: a list of numbers that places the token at a specific point in space. Those input embeddings start the model's computation. Later layers produce new representations for each token using its position and context.

Why embeddings?

A token id is arbitrary. “cat” might be 9707 and “dog” 3899, but those numbers are just slots in a vocabulary, being close in id says nothing about being close in meaning. An embedding replaces that arbitrary label with a learned position in space, fitted during training so that words used in similar ways land near each other. Now distance and direction carry meaning: “cat” sits beside “dog”, both far from “Tuesday”, and the model has something it can actually compute with.

Explore a word's neighborhood

Pick a word, follow its neighbors, and compare the numbers behind their meaning.

Search this teaching collection of 139 words.

Word in focus

cat

Animals

Six closest neighbors

Cosine similarity

Within this collection. Tap a word to explore from there.

Bars run from −1 to +1. The center mark is 0.

How close are two words?

Your word

cat

0.922cosine similarity

−10+1

Closer to +1 means the vectors point in similar directions. This is not a probability. Words with opposite meanings can still score highly when used in similar contexts.

The numbers behind the words

Each fingerprint shows all 50 learned values, 25 per row. Both words use the same color scale.

cat0.4528
dog0.1101
● Negative● PositiveBrighter = larger magnitude

A single cell is not an “animal” or “emotion” switch. Meaning is encoded in the full pattern.

These are genuine pretrained GloVe embeddings from Stanford, using 50 numbers per word. The explorer searches a curated collection of 139 words. Neighbor rankings and comparisons use all 50 values, and the fingerprints show those same values directly. GloVe is an older word embedding model used here to explain the idea. A language model learns its own input embeddings and updates token representations through transformer blocks. These whole-word GloVe vectors are a separate teaching dataset; they are not the embeddings of the IDs from the Tokens lesson.

What do the numbers actually mean?

It's tempting to imagine each dimension stands for one human idea, an “animal-ness” axis, a “royalty” axis. It almost never works that way. Individual dimensions rarely line up with a single concept we'd name. Meaning lives in the combination of many dimensions at once: it's the overall pattern of the vector, and its position relative to every other vector, that encodes what a word is like, not any one number on its own. That's also why the interesting structure shows up as directions and clusters rather than as readable coordinates.

Similar usage does not mean identical meaning

Try comparing happy with sad. They describe opposite feelings, but both appear in similar sentences about emotions. Their vectors can therefore be close. An embedding captures patterns of use, not a dictionary definition or a guarantee that two words are interchangeable.

Cosine similarity measures the angle between the full vectors. A score of +1 means the same direction, 0 means perpendicular directions, and −1 means opposite directions. It is a geometric score, not a percentage of shared meaning.

Embeddings are everywhere, not just in LLMs

The same idea, turning something into a vector whose position encodes meaning, powers a huge amount of modern software. Semantic search finds documents whose embeddings sit near your query's. Recommendation systems place songs, products, or videos in a space and suggest nearby neighbors. Retrieval-augmented generation (RAG) can use a dedicated embedding model to compare a query with passages, then supply selected text to a language model. Document embeddings are trained for a different use from token input embeddings. Similarity helps find candidates; it does not guarantee that a passage answers the question. The Assistants lesson follows that process further.

Check your understanding

Why might happy and sad be close in this explorer?
Sources & further reading
Now every token is a vector in space. Next: how those vectors look at each other, so each word's meaning can shift based on its context.→

Lesson 03 / 08

Attention

Your goal: Trace queries, keys, values, masking, and the other operations inside a transformer block.

Embeddings gave every token a starting vector, but that vector is the same no matter what surrounds it. Position information and transformer blocks turn that starting vector into a representation of this occurrence in context. Attention computes weights over allowed positions and mixes information from them.

Why attention?

Without it, bank would carry the exact same vector in “river bank” and “savings bank”, since the embedding table has only one entry for the word. Attention is what lets the surrounding context influence the representation: “river” and “savings” supply different information. More precisely, attention mixes learned value vectors, then projects and adds the result to the representation. The rest of the block transforms that result further.

Try it: click a word and watch what it looks at.

Load the model, then click it in the default sentence, and compare the weights on animal and the other words. The brighter a word, the larger its displayed attention weight. Change the sentence or layer and observe how that pattern changes.

Try:

Kept to 14 words, since the model runs live in your browser, so short sentences stay snappy.

These weights are not a similarity heuristic: the demo downloads a real DistilBERT transformer with its own WordPiece tokenizer. The displayed weights average heads, combine subword pieces, remove framing tokens, and renormalize over the displayed words. They summarize real attention rather than showing an individual head's untouched distribution. A bright word shows a large mixing weight; it does not prove a grammatical link or explain the model's final decision.

One catch: this model looks both ways

Notice that it can attend forward to tired, a word that comes later in the sentence. That's because the demo runs a bidirectional model (DistilBERT), which sees the whole sentence at once, handy for showing the mechanism clearly. GPT-2 and other autoregressive decoder models use causal attention: each position is masked from future positions, so it can only look back at itself and the words before it, never ahead. Attention there flows in one direction only.

Many heads, many kinds of looking

A model doesn't run just one of these attention passes; it runs several in parallel, called heads. Each head has its own learned projections and can emphasize different patterns, such as nearby positions or particular relationships. Heads do not have fixed human-assigned jobs, and their behavior depends on the model, layer, and input. The demo averages them together, which can hide differences between individual heads.

And it happens layer after layer

One round of attention isn't the whole story either. The model stacks the same mechanism in layers (DistilBERT has six), and each layer refines the token representations before passing them up. Each block also contains a feed-forward network, residual connections, and normalization. Patterns across layers can differ; a later layer is not guaranteed to reveal a clearer grammatical relationship. Flip open step through the layers to watch the same sentence's attention shift as it moves up the stack.

Inside a transformer block

Step through a decoder like GPT-2. This uses its normalization-before-sublayer order; DistilBERT places normalization differently and allows attention in both directions.

1. Token + position

Look up each token's learned embedding. GPT-2 adds a learned vector for its position. This gives the same token different starting representations at different positions. Some newer architectures instead apply rotary position information to queries and keys.

Work through one attention head

These three words stand in for three tokens. The two-dimensional Q, K, and V vectors below are hand-chosen examples of what learned projections could produce. They come from a teaching calculation, independent of DistilBERT's weights.

Choose the position to update:

Q for bank = [1, 1]

Scaled dot-product attention for bank
TokenKey KValue VQ · KScore ÷ √2Weight
river[1, 0][2, 0]10.70750.0%
bank[0, 1][0, 2]10.70750.0%
rises[1, 1][3, 1]2masked (−∞)0.0%

weighted sum of V = [1.000, 1.000]

With bank selected and the mask on, river and bank receive 50% each, giving 0.5 × [2, 0] + 0.5 × [0, 2] = [1, 1]. Turn the mask off: rises can now contribute. These weights sum to 1 over the allowed positions; the result is a value-vector mixture, which later operations project and add to the running representation.

Optional math: the attention equation

Q = XWq, K = XWk, V = XWv; Attention = softmax(QKᵀ / √d + mask)V

X contains the input representations and the W matrices are learned parameters. A masked score of −∞ gets zero softmax weight. Attention changes activations for this input; the W matrices change during training.

Check your understanding

In causal attention, which tokens can bank use in the sequence river bank rises?
Sources & further reading
Attention, embeddings, all of it is steered by millions of weights. Next: how those weights actually get set: error flowing backward through the network.→

Lesson 04 / 08

Backpropagation

Your goal: Separate computing gradients with backpropagation from applying an optimizer update.

Attention, embeddings, all of it runs on weights, and at the start those weights are just random numbers. A forward pass produces an output, and training compares it to the correct answer to get a single number: the error, or loss. But one error number doesn't obviously tell you how to fix any particular weight. Backpropagation computes how that loss changes with each weight: its gradient. An optimizer uses those gradients to choose the weight updates.

Why backpropagation?

A real model can have billions of weights. You could imagine testing each one: nudge it, run the whole network again, see if the error went down, but that means billions of full forward passes per step, which is hopeless. Backprop gets the answer for every weight in a single backward pass. It uses the chain rule from calculus to share work across the network: it figures out how wrong the output was, then hands each layer its share of the blame, reusing the layer above instead of recomputing from scratch. That efficiency is the whole reason training large networks is even possible.

Try it: run one prediction, then send the error back.

Click Step forward to push an input through the network to a real prediction, then Step backward to watch the error compute a gradient for every weight and nudge each one. Then hit Train and watch the loss fall.

0.300.50-0.30-0.400.200.600.40-0.500.301x10x2h1h2h3ŷinputshidden layeroutput
Input [1, 0]Prediction –Target 1Loss (all 4) 0.1252
━ positive weight━ negative weight■ gradient (how to nudge that weight)

Ready. Step forward runs the current weights on the highlighted example below.

x1x2targetprediction
0000.52click to trace
0110.52click to trace
1010.51shown above
1100.51click to trace
Loss over training steps0 steps
0.1400

Every point is the network's real average error, recorded after an actual weight update. Compare the curve over many steps; individual updates are not guaranteed to reduce loss for every initialization.

This is a real network: two inputs, a hidden layer of three neurons, one output, learning XOR (output 1 only when the inputs differ), a task a single linear classifier cannot solve. The predictions, loss, gradients on each edge, and the weight updates are all computed live in your browser with actual chain-rule math. Nothing here is a pre-recorded animation; hit Reset and it learns again from new random weights.

Gradient descent: many small nudges

A gradient is a local rate of change of loss with respect to a weight. Plain gradient descent subtracts learning rate × gradient from each weight. With a sufficiently small step this moves downhill; a large step can overshoot. A single step doesn't fix the network; it just moves every weight a little bit downhill. Do that over and over and the errors shrink, which is exactly what the falling loss curve is showing you: not one clever fix, but thousands of tiny, honest corrections adding up.

The same thing, at unimaginable scale

Every weight, in every layer, gets its own gradient and its own nudge, all at once, every single training step. The toy above has around a dozen weights, so you can watch each one move. A transformer has parameters in embedding tables, attention projections, feed-forward networks, and normalization operations. The chain rule applies across those operations too. The XOR network illustrates differentiation and updates; it is a different architecture from an LLM.

One honest caveat: the backward pass computes gradients by the chain rule. The operations being differentiated depend on the architecture. The update rule is where real training differs. Instead of the plain fixed-size step the toy above takes, LLMs use smarter optimizers like Adam (and AdamW), which add momentum and give each weight its own adaptive step size. The gradients come from backprop; Adam just decides how to spend them.

Check your understanding

A weight has a positive loss gradient. What does a plain gradient-descent step do?
Sources & further reading
One backward pass nudges the weights once. Next: what happens when you repeat it across a whole dataset, millions of times: training.→

Lesson 05 / 08

Training

Your goal: Compute a next-token loss and distinguish learning the training data from generalizing to unseen data.

In the last lesson, one backward pass nudged every weight a tiny bit. That single update barely changes anything. Training is that exact loop: forward pass, measure the error, backprop, nudge, repeated millions of times over an enormous amount of text. Do it enough and a network of random numbers gradually turns into something that can write.

What does an LLM train on?

A common pretraining objective for causal language models is predict the next token. Show the model a stretch of text with the next piece hidden, let it predict a distribution, and score the probability assigned to the observed next token. Training text can include books, code, and web pages. Many useful patterns develop through prediction, while post-training can further teach instruction following and task-specific behavior. Other model families can use different objectives, such as masking tokens.

Watch a real model learn, live.

Below is a tiny language model training in your browser: no pretrained weights, no server, no canned animation. Press Start training and read the sample text as the loss curve falls: it goes from random characters, to strings with the right letter frequencies, to word-shaped fragments, to real words. It is a character-level feed-forward network (MLP) with three characters of context, rather than a transformer. The prediction, loss, and gradient-update loop illustrates the same general learning process.

Try it: press Start training and watch the model learn to write.

It starts knowing nothing: the sample text is pure random characters. As the loss curve falls, watch the samples turn into letter-frequency soup, then word-shaped fragments, then real words and phrases. Same predict-the-next-character task, repeated thousands of times.

Step 0Loss 3.580
Model's writing right nowstep 0
,yl!g'!aatg.klyfqkqmwphbc.ymi!fnbbjbxpq'eey?g
gk-cv'ul!:; y,?.vqyih'lqh?:imkoxn.ys,sy?a'h't fpt xp t
gdxlaig,qqtv?wohq ba,-phdsylipohailb-r:o!
k xcr;.anpky 
;si:d
v-lluip,;
maybwdk

Generated live: the model predicts a next-character probability for the last 3 characters and one is sampled from it, repeatedly. Nothing here is pre-written; reset and it learns from a brand-new random network.

Cross-entropy loss over training steps3.58 → 3.58
3.580

How the writing evolved

A saved sample from each stage of training, so the progression is visible at a glance.

step 0 (random)

kb:!u.yy:-ddxerogu!t tg-a.-iwn'qighwkcvvl?jiey:cgqdhphv'j jgkhpxdwhgmtoj?u;!stu?xjgms m;ftdh' ttnwx,elkavb -sm

now

,yl!g'!aatg.klyfqkqmwphbc.ymi!fnbbjbxpq'eey?g gk-cv'ul!:; y,?.vqyih'lqh?:imkoxn.ys,sy?a'h't fpt xp t gdxlaig,q

What you're actually training

A tiny character-level feed-forward language model (MLP), 4,155 parameters, a 35-character vocabulary, learning from 4.7 KB of text (Excerpt from Alice's Adventures in Wonderland by Lewis Carroll (1865, public domain).). It reads the last 3 characters and predicts the next one; gradient descent adjusts its parameters each step. It has no attention layers; the short context limits it to local character patterns. Its loss curve measures sampled training examples, rather than a held-out evaluation.

Advanced: the learning rate

The learning rate sets how big each downhill step is. Too low and the loss crawls; too high and the updates overshoot and the loss explodes. Change it, hit Reset, and train again to feel the difference.

lr = 0.20

Applied on the next step; use Reset to compare cleanly from a fresh network.

One prediction, one loss

For the made-up training example The pet is a → cat, pretend the vocabulary contains just cat, dog, and fish. Raise cat's score and watch its probability rise and its loss fall. The scores here are chosen for teaching.

cat (target)logit 1.0 → 42.2%
doglogit 1.0 → 42.2%
fishlogit 0.0 → 15.5%

loss = −ln(0.4223) = 0.862 nats

Softmax exponentiates the logits and divides each by their sum, giving positive probabilities that sum to 1. Cross-entropy penalizes low probability on the observed token: 90% gives about 0.105 nats; 10% gives about 2.303. A batch's loss averages these penalties before the backward pass.

Optional math: softmax and perplexity

pᵢ = exp(zᵢ) / Σⱼ exp(zⱼ); perplexity = exp(mean loss)

Perplexity summarizes average predictive uncertainty; lower is better on the same evaluation data. Here exp(loss) = 2.37 for this one target. Scores from different tokenizers, vocabularies, or datasets are not directly comparable.

Where do the targets come from?

The text supplies its own labels. For a sequence such as the cat sat, each position is trained to predict the token at the next position. During training, the input contains the actual previous tokens, a setup called teacher forcing. A transformer's causal mask lets it compute predictions for many positions in one forward pass without using the future token as input to its prediction.

During generation, the next input includes tokens the model itself selected. Errors can therefore change the context for every later prediction. Training and generation share a prediction objective but use different sources for those previous tokens.

Epochs, batches, learning rate, in plain words

Each step above doesn't look at the whole text at once; it grabs a small random batch of examples, because averaging the error over a handful is far cheaper than over everything and works nearly as well. One full pass over all the training text is an epoch. Recipes can use one pass or reuse data; the random sampling here does not visit examples in a fixed epoch order. The learning rate is how big each update is. As the advanced control in the demo shows, too small can learn slowly, while too large can make the updates diverge.

Learning the examples versus generalizing

This demo plots loss on sampled training examples. A falling curve shows a better fit to those examples. To test generalization, keep other text out of training. The Scaling lesson measures loss on an unseen portion of its corpus.

Training set
Examples that produce gradients and change the weights.
Validation set
Held-out examples used to compare settings or decide when to stop.
Test set
A separate set reserved for the final assessment after those choices.

If training loss falls while validation loss rises, investigate overfitting: fitting the training examples too closely to transfer well. If test examples or near-duplicates enter training, data leakage can make the score look better than it should. Prediction loss is useful, but it still does not measure every task an assistant must perform.

The dataset is part of the model's education

Preparation includes selecting sources, cleaning text, removing duplicates, filtering unsuitable content, and balancing useful examples. Data quality and coverage influence what gets learned, including errors and social biases. This demo uses a small public-domain text sample cleaned to a limited character set; it cannot demonstrate broad language competence.

This is only pretraining: chatbots need more

What you trained here is a raw next-token predictor: it continues text, but it doesn't know it's supposed to be helpful, follow instructions, or answer a question rather than ramble on. That first stage is called pretraining. Turning a pretrained model into an assistant takes further stages: fine-tuning on curated examples of good responses, and learning from human and AI feedback about which answers people prefer. The model you meet in a chat window often uses some combination of these stages; the demo illustrates pretraining. The final Assistants lesson explains the stages and the chat system that uses the trained model.

The gap to a frontier model

The demo model has a few thousand parameters and learns from a few kilobytes of text. Larger language models use much larger datasets and computation budgets, often with transformer architectures, subword vocabularies, adaptive optimizers, and distributed training. They share the predict, measure, backprop, update loop, but the architecture and training recipe differ substantially from this three-character MLP.

Check your understanding

Training loss keeps falling while validation loss rises. What should you investigate?
Sources & further reading
What happens when we vary size while keeping an architecture and objective fixed? Next: scaling, and how to interpret the experiment.→

Lesson 06 / 08

Scaling

Your goal: Interpret a controlled size experiment without confusing parameters, compute, and useful capability.

In the last lesson you trained a real model with a few thousand parameters. Now keep that toy architecture and objective fixed while increasing its width. Does its prediction on unseen text improve? This small experiment introduces the broader study of how model size, data, and computation affect prediction loss.

Why scale works: scaling laws

Researchers have found approximate power-law relationships between prediction loss and model size, data, or compute within studied regimes. After accounting for a loss floor, a power law appears straight on logarithmic axes for both variables. Those empirical fits can help forecast loss within their assumptions; they do not guarantee every downstream skill. The Chinchilla study examined how to allocate a fixed training-compute budget between model parameters and data. It found that training a smaller model on more tokens could outperform a larger, less-trained model.

Run a small size experiment.

The demo below trains the same character-level network from the Training lesson at six sizes, from a few hundred parameters to tens of thousands, each on the same text for the same number of steps. It then plots the held-out loss of each. This is a genuine experiment running on your machine. Observe the trend and any exceptions; six points from one run do not establish a general scaling law.

Try it: press Run the experiment and compare held-out loss as the models get bigger.

Six models are trained one after another, smallest to largest, each on the same text for the same number of steps. The only thing that changes is embedding and hidden width; initializations and sampled batches also vary. A point drops onto the chart as each one finishes: do they land on a smooth downward curve?

2.12.32.52.7232XS696S1.7KM4.2KL9.8KXL22.4KXXLparameters (log scale) →held-out loss (lower is better)Press “Run the experiment” to plot real loss vs. size.
Smallest: 232 params1×4

Press Run to generate a sample.

Largest: 22.4K params16×256

Press Run to generate a sample.

Both samples are generated live by the two trained models after the run: same task, text, and training steps, with different widths and random initialization.

This is a real experiment, run in your browser

Every model here is the exact same character-level network from the Training lesson: reused code, not a copy, trained by real gradient descent on 137 KB of text (The complete text of Alice's Adventures in Wonderland by Lewis Carroll (1865, public domain, via Project Gutenberg).). For each size only the embedding and hidden widths grow. Every model sees the same text for the same 1,500 steps, and the loss you see is measured on a held-out 10% slice the model never trained on. Lower loss means better average prediction of those held-out characters. Equal steps do not mean equal compute; wider models do more work per step. Nothing is precomputed; press Run again and six brand-new networks train from scratch.

One variable at a time: read this carefully

This experiment deliberately changes only the embedding and hidden widths, holding the data, context length, batch size, learning rate, and number of steps fixed. Wider models generally require more arithmetic per step, so this is not an equal-compute comparison. Each size also has a different random initialization and sampled batches. Repeat runs to judge variability. The fixed three-character context and finite dataset can limit gains; larger models can also overfit. Undertraining and overfitting are different problems: a model can need more training without having memorized its data.

Selected milestones: size and efficiency

The timeline follows selected published model sizes from 2018 through February 2026. It is a history of different architectures and training recipes, rather than a capability leaderboard. Every entry links to its original report or model card. Some later models use a mixture of experts (MoE): they store many parameters but activate only a subset for each token.

Before the timeline: how we got to transformers

Language models existed long before the modern giants, but each step introduced new ways to represent and process language. The transformer became a foundation for many later models, including newer hybrids.

  1. ~1990sn-gram models: Predict the next word from raw counts of the last few, with no learning of meaning.
  2. 2013word2vec: Learned word vectors where meaning became geometry, so that king − man + woman ≈ queen.
  3. 2014RNNs / LSTMs & seq2seq: Networks that read a sequence one step at a time; powered the first neural translation.
  4. 2017“Attention Is All You Need”: The transformer combines attention with position information and feed-forward layers, using no recurrence. Later architectures also explore hybrids.

Selected published sizes, 2018–2026. Hover or tap a point, or choose a model below.

100M1B10B100B1T201820192020202120222023202420252026GPT-3year releasedparameters (log scale)
GPT-3OpenAI· May 2020175B total parameters

Few-shot learning from the prompt alone, with no fine-tuning needed.

Read the original report or model card (opens in a new tab) ↗

Every plotted count is from its creators' report or model card. This is a selected history, updated September 30, 2026, with entries through February 2026; it is not exhaustive. Models without a sourced public parameter count are omitted, and each point refers to the specified variant. The connecting line follows releases, not a fitted scaling law or ranking. Chinchilla was smaller than Gopher yet performed better after training on more data.

“Emergent” abilities, and the argument about them

Something odd shows up as models grow: certain skills, like multi-step arithmetic, following an unusual instruction, translating a rare language, seem largely absent in smaller models and then appear fairly suddenly past some size, rather than improving gradually. These are often called emergent abilities. Their interpretation is debated: some of the “sudden” jumps are partly an artifact of how the ability is measured : a strict all-or-nothing score can flip from zero to one abruptly even when the underlying skill was improving smoothly all along. Evaluate the actual task and metric, rather than assuming a particular parameter threshold guarantees an ability. Both the emergence study and the critique are linked in the sources below.

Scale isn't free

Larger training runs require more arithmetic, memory, and energy. Data quality, architecture, and optimization can change what that budget achieves. Deployment adds another tradeoff: a smaller model trained for longer may cost less to run for many users. MoE can reduce active computation relative to total size, but storing weights and routing between experts still have costs. Parameter count alone describes neither quality nor operating cost.

Check your understanding

These models take the same number of training steps. Do they use the same amount of compute?
Sources & further reading
A big trained model outputs a probability for every next token. So how does it turn that into actual words? Next: generation.→

Lesson 07 / 08

Generation

Your goal: Separate next-token probability, token selection, and factual correctness.

Every lesson so far has been building to this. Text is split into tokens, each token becomes an embedding, attention lets those vectors mix in context, and training tuned every weight along the way. And it all ends at one thing: a probability for every possible next token. Generation is what happens when you act on that number, over and over.

From a vector to a vocabulary distribution

After the transformer blocks and final normalization, the last token's representation is mapped to one raw score, or logit, per vocabulary entry. This output projection is sometimes called unembedding. Softmax turns the logits into probabilities that sum to 1; the worked loss example in Training uses the same conversion.

This demo uses GPT-2's own embeddings, transformer blocks, and 50,257-entry r50k_base vocabulary together. Its IDs differ from those in the Tokens lesson, and its vectors differ from the GloVe explorer.

Inspect predictions across layers

Token representations pass through repeated transformer blocks. The logit lens applies the final output readout to intermediate representations, letting us inspect how those representations project into vocabulary scores. Compare the readings across layers; they need not improve smoothly or monotonically.

What am I looking at?

A layer is one processing step inside the model. GPT-2 small has 12 of them stacked in a row. Each layer reads the running representation of your text, mixes in a bit more context, and passes it up to the next. The model only commits to a next-token probability after the final layer, but there is a clever trick, called the logit lens, that lets us peek early: take the half-finished representation at any layer and run it through the same output step the model uses at the end. This produces a diagnostic prediction from that representation.

A guess at a middle layer is not the model's real answer. It is a probe reading, not a literal record of thoughts. The output readout was trained for the final representation, so applying it earlier can give poorly calibrated or unrelated predictions. A change in a bar does not establish exactly when the model learned or retrieved a fact.

What the bars measure: probability assigned by the readout to a token. The final layer gives GPT-2's actual next-token prediction at temperature 1; the earlier readings are diagnostic. Token probability describes a continuation under the model, not confidence that a statement is true.

Try this: compare readouts across the layers.

Load GPT-2, then read the stack from the bottom up. Each row is one of the model's 12 layers, with the final vocabulary readout applied at that depth. Look for changes in the leading token and its probability. Intermediate readouts are probes; they need not approach the final prediction smoothly.

Try:

Compare a bare question with a prompt containing a worked example. The example can change the model's continuation without changing any weights: this is in-context learning. GPT-2 can still produce an incorrect or incomplete answer in either case. Observe what it predicts rather than assuming the prompt succeeds.

One token at a time

How a sentence gets built

In standard autoregressive decoding, the model predicts a single next token, one is chosen, that token is appended to the input, and the model predicts again, now conditioned on that added token. This demo recomputes the full sequence each time. Other systems can reuse intermediate computations. A model can generate plans or intermediate steps in its context, but this demo exposes next-token predictions rather than a complete description of planning.

Below is the real thing: GPT-2, the original 124-million-parameter model, downloaded and run entirely in your browser. The bars are the actual softmax probabilities it computes for the next token, not a mock-up. Click a bar to commit that token and see the distribution for what follows.

Try it: build a sentence one token at a time.

Load the model, then click any bar to append that token to the text, and the chart instantly recomputes what GPT-2 thinks comes next. Or hit Generate 20 tokens and watch the loop run itself.

Try:

Temperature & sampling, in plain terms

Once you have a probability for every token, you still have to pick one. Greedy decoding always takes the single highest bar: fully deterministic, and quick to fall into repetitive loops. Sampling instead rolls a weighted die, so a token with 40% probability is picked about 40% of the time, and even long-shot tokens occasionally come up.

Temperature is the knob that reshapes those odds before the die is rolled. Turn it down and the probabilities sharpen toward the top token, concentrating mass on the leading candidates. Turn it up and they flatten out, handing rare tokens a real chance: more creative, but more likely to select unlikely continuations. Sampling is one reason repeated prompts can produce different output. Chat applications can also retrieve information or call tools before generating. Temperature alone does not guarantee creativity, correctness, or safety. With fixed logits, positive temperature does not change which token greedy decoding selects.

Some decoding configurations trim the distribution before rolling. Top-k keeps only the k most likely tokens; top-p (nucleus) keeps the smallest set of tokens whose probabilities add up to some threshold, say 0.9, and throws the rest away. A repetition penalty can also adjust scores of previously used tokens. The surviving probabilities are renormalized before sampling. These are optional strategies; their order and implementation vary. This demo uses greedy selection or temperature sampling from the full vocabulary, without top-k, top-p, or a repetition penalty.

Why “predict the next token” can sound so capable

It's fair to wonder how something this simple produces essays, code, and arguments. Part of the answer is scale: the model you just ran is tiny and often clumsy, but the same next-token objective, trained on far more text with vastly more parameters, absorbs an enormous amount of the structure of language and the world. To predict the next token well across billions of examples, a model has to pick up grammar, facts, styles, and patterns of reasoning, not because anyone programmed them in, but because they help lower the prediction error.

The other part is post-training. A raw predictor will happily continue your text in any direction; the helpful, instruction-following assistant you talk to is that predictor after extra rounds of fine-tuning and human feedback that steer it toward being useful and on-topic. It's worth being clear-eyed, though: because the core act is predicting plausible text rather than checking truth, a model can state false things with complete fluency. The widely-used term is hallucination. Sounding confident and being correct are simply not the same operation here.

Context, caching, and stopping

The prompt and selected tokens must fit within a context limit. This browser demo uses a short limit for responsiveness. Generation stops at an end-of-text token, a length limit, or your Stop action; chat systems may use other end-of-turn markers.

Many decoder implementations cache previously computed keys and values (the KV cache) so each new token can attend to them without recomputing earlier positions. Longer contexts still use memory and attention computation. This cache stores activations for the current sequence, rather than changing the learned weights.

Check your understanding

GPT-2 assigns a token 80% probability. What does that number describe?
Sources & further reading
Next: how post-training, chat context, retrieval, and tools turn a text predictor into an assistant.

Lesson 08 / 08

Assistants

Your goal: Explain what post-training changes and how retrieval, tools, and chat context support an assistant.

GPT-2 completes a piece of text. An assistant is a model trained for conversation, wrapped in a system that supplies instructions, history, and sometimes outside evidence or tools. Understanding that surrounding system explains much of the difference between this course's generation demo and a chat application.

How training changes the behavior

  1. 1. Pretraining learns broad patterns. A causal model learns from text by predicting next tokens. This creates a base model that can continue many kinds of text, with no guarantee that it will follow a user's request.
  2. 2. Supervised fine-tuning demonstrates responses. Curated prompt-and-response examples show how an assistant should answer. The model trains to assign higher probability to these responses. Training changes its weights; writing an example into a prompt alone does not.
  3. 3. Preference learning compares responses. People or other systems can judge which of two answers is better. RLHF can train a reward model from those comparisons and use reinforcement learning to update the assistant. DPO uses preference pairs directly to optimize the model. Recipes vary, and some also reward verifiable outcomes, such as passing a coding test.

These stages steer behavior according to the examples and feedback supplied. A preference for a fluent or pleasing answer does not automatically make that answer correct. Post-training can improve usefulness while leaving errors and bias to measure.

What enters a chat model

A chat template turns structured messages into the token sequence the model expects. Role markers distinguish instructions, user messages, assistant replies, and tool results. For illustration:

System / application instruction
Answer from the provided library document and cite the relevant line.
User message
When does the library close?
Retrieved context
Library hours: 9am–6pm, Monday to Friday.
Assistant response begins
The library closes at …

The exact formatting and instruction roles vary by model. The application chooses which messages and passages fit in the context window. Long conversations may be shortened or summarized. A product can save memories and retrieve them later, but the model's context is limited and those memories are supplied by the application.

Retrieval adds evidence; tools perform operations

In retrieval-augmented generation (RAG), a system searches a collection, selects relevant passages, and inserts them into the model's context. A dedicated embedding model can rank passages by similarity, often combined with keyword search or reranking. These document embeddings serve a different task from the token input vectors in the Embeddings lesson.

A tool flow lets the model request an operation, such as a calculation or database lookup. The application executes the request and returns its result; the model then continues with that result in context. For example: question → retrieve the hours document → supply the passage → generate an answer with a citation.

Retrieval and tool calls usually happen with fixed model weights. Fine-tuning changes weights. In-context learning changes how an existing model responds to the current examples. Retrieval can miss the right document, and tools can fail, so outside information still needs checking.

Reasoning uses a budget too

Some systems spend extra inference compute generating intermediate steps, testing candidates, or using tools before responding. That can help on difficult tasks, with a cost in time and tokens. A visible explanation is another generated output; it does not provide a complete record of the model's internal computation.

Evaluate the answer and the system

A hallucination is an unsupported or incorrect generated claim. Neither fluent wording nor high next-token probability establishes truth. Evaluate the behavior you need on examples that were held out from training and development:

  • Correctness: use known answers, executable tests, or qualified review. Check whether citations actually support the claims.
  • Retrieval quality: did the system find the relevant passage, and did the answer use it accurately?
  • Coverage and bias: compare performance across languages, groups, and unfamiliar cases; an average can hide substantial failures.
  • Reliability and cost: repeat important cases when sampling varies, record failure rates, and measure time and token use.

Human and model judges can disagree or favor particular styles. Document the rubric and inspect failures. Also distinguish the model from the application's data handling: this site's demos run locally after downloading model files, while a hosted assistant may send prompts and documents to its service. Data quality, permissions, and privacy are choices throughout the system.

Check your understanding

An assistant retrieves a new document and answers from it. Have its model weights necessarily changed?
Sources & further reading

Final exercise: trace one answer

Use the library-hours example above. Explain how the question and retrieved passage become an answer. Include the tokenizer, positions and attention, output probabilities, token selection, whether the weights change, and how you would check the result.

Reveal a worked answer
  1. The application retrieves the hours document and formats instructions, the question, and the passage using the model's chat template. Its tokenizer encodes that context into vocabulary IDs.
  2. The model looks up matching input embeddings and incorporates position information. Its decoder blocks use causal attention to mix earlier information, plus feed-forward transformations, normalization, and residual connections.
  3. The final representation at the current last position is projected into vocabulary logits. Softmax and the decoding settings produce a next-token distribution; greedy decoding or sampling selects a token.
  4. The selected token is appended and generation continues until a stop condition. A system may reuse cached keys and values. The passage affects the activations, while the trained weights stay fixed during this inference.
  5. Check that the answer says 6pm on Monday to Friday, cites the supplied hours, and does not invent weekend hours. An answer of 6pm without the weekday qualification may be incomplete, even if every individual token had high probability.