Probability Primer

The probability a Transformer actually runs on, built out of one picture: a row of bars whose heights add to one. Distributions, expectation and variance, Bayes, likelihood, sampling, entropy and cross-entropy — every number on the page is computed by the figure beside it, so you can drive it anywhere and it stays true.

01

A probability is a share

One number between 0 and 1. Everything else on this page is arithmetic on that one number.

Two different things get called a probability. One is a frequency: run the experiment a hundred thousand times and count. The other is a degree of belief about something that happens once. They share a name because they obey the same three rules, and the frequency reading is the one you can draw. So we draw it — a hundred squares, one per run, and the filled ones are the runs where it happened:

p = 0.35 — 35 of 100 squares

Notice there is nowhere past the last square for the slider to go. A probability lives in [0, 1] because it is a share of something, and a share cannot exceed the whole: at the right-hand readout stops printing a fraction and says certain, and at it says never.

The frequency reading comes with a warning attached. Flip a fair coin and the running share of heads wanders before it settles onto the probability — and it settles far more slowly than people expect. Push the flip count out and watch the wander shrink:

after 1 flip

After one flip the share reads 1.000: the coin came up heads, and the evidence so far says it always will. It takes to pull the share to 0.570 and 10,000 to reach 0.501. The band is one standard error, 0.5/√n — four times the data for half the error.

The third rule is the one every later section leans on. Give each outcome a share of one bar — rain, cloud and sun tile it exactly — so moving a boundary takes width from a neighbour rather than lengthening the bar:

rain 0.35 · cloud 0.36

Watch the right end of the readout: sun is never chosen, it is computed, 1 − 0.35 − 0.36. That is what sums to 1 means in practice — one of the numbers is not free. A forecast of 0.35, 0.36 and 0.40 is not optimistic; it is arithmetic that does not close.

Almost nothing arrives summing to 1, so we divide by the total — which works until it doesn't. Drag the third score below zero:

score = 3.0

At c = 3 the scores 3, 4, 3 normalise to 0.30, 0.40, 0.30 over a total of 10. At the total is 6 and the third share is −0.17. Three numbers come back, nothing throws, and the shape is nonsense. Normalising does not check anything — and on an all-zero score vector it returns NaN with no message at all.

02

A distribution is a shape

One number per outcome, none of them negative, adding to exactly one.

Cut the bar from §01 into as many pieces as there are outcomes, stand the pieces up as heights, and you have the picture every plot in machine learning is a variation on. The dashed line here is where a fair die puts every face — raise the loading and watch the six take its extra mass out of the other five:

p(loaded face) = 0.17

Watch the sum in the readout as you drag: it never moves. At the six holds 1.00 and the other five are exactly 0 — still a legal distribution, with nothing random left in it. That degenerate corner matters later: a zero temperature, an argmax and a one-hot label are all this same shape.

Models do not emit probabilities. They emit logits — unbounded real scores, one per token — and softmax exponentiates them and divides by the total. Drag the second token's logit and watch the row underneath follow it:

z = 1.0

Because exp is monotonic the two rows are in the same order, so softmax never changes the argmax. What it changes is spacing. The token under your hand moves in both rows at once, a logit gap of 1 becomes a probability ratio of e ≈ 2.72, and the gap of 5 between the first token and the last is therefore a factor of 148 — 0.577 against 0.0039.

Every real implementation subtracts the largest logit before exponentiating. It cancels in the numerator and the denominator, so the answer is identical; what it buys is that exp(800) is Infinity in float64 and exp(0) is 1. Over a 50,257-token vocabulary the whole pass is one exponential and one divide per token, which is nothing beside the matrix multiply that produced the logits. Dividing the logits by a temperature first is a real change, though — pull it toward zero and watch the top token eat the row:

T = 1.00

As soon as T reaches the top token holds 1.000 to three places — that is greedy decoding, and what temperature=0 means in an API. At it is down to 0.278, heading for 1/5. Temperature changes not what the model believes but how much of it survives into the sample.

Continuous outcomes need one more idea, and it is the one that trips people. A density is not a probability; it is probability per unit of x, so it is free to be larger than 1. Narrow the bell and drag the window along it:

centre = 0.00
area 0.383

At σ = 1 the peak reads 0.399 and the window holds 0.383. At σ = 0.2 the peak is 1.995 — a value no probability may take — while the same window holds 0.988, which is perfectly legal. The area is the probability; the height is only its rate. A density above 1 is not a bug report.

03

Where it sits, how far it spreads

Two numbers stand in for a whole distribution: a balance point and a squared distance.

Put the outcomes on a line, hang each one's probability above it as a weight, and the whole distribution acquires a balance point. Slide mass onto one face and watch the fulcrum follow it along the beam:

E[X] = 3.50

Because every face is equally likely, a fair die balances at 3.50 — a value it cannot roll. That is the first thing to accept about an expectation: it is a weighted address, not an outcome. Load and the fulcrum reaches 6.00; the same loading on the one drags it to 1.00 instead.

Spread needs a second number, and squaring is what makes it work. Every outcome pays its distance from the mean, squared, weighted by its own probability. Push the mass out to the ends and watch the bill grow:

σ² = 2.92

Notice that the middle faces contribute almost nothing even while they are the likeliest ones: half a face away the squared term is 0.25, against 6.25 out at the ends. The squaring is not cosmetic. It is what makes variance add across independent variables, which absolute distance does not do.

The price of squaring is the units. A variance of 4.15 is in faces squared, which is not a thing anyone can picture. Take the square root and the number comes back onto the axis it belongs on:

σ = 1.71

At the uniform setting σ is 1.71 and the interval μ ± σ covers 0.67 of the mass — four faces out of six. Push the spread and σ grows to 2.38 while the covered mass falls to 0.13. σ is a ruler, not a container.

One more, because it is the mistake people actually make. Estimate a variance by dividing the squared deviations by n and the answer comes back too small. Drag the sample size:

1.458

At n = 2 the ÷n estimator averages 1.458 against a true 2.917 — exactly half, and it never says so. At it reaches 90.0% of the truth, at n = 20, 95.0%: the bias is a factor of (n−1)/n, shrinking but never closing. Dividing by n − 1 lands on the line at every n.

The variance of a sample mean is σ²/n, so its standard error is §01's 1/√n again — which is why batch size behaves as it does. GPT-3 trained at a 3.2M-token batch: a thousandfold larger batch buys a gradient only thirty-twofold quieter.

04

Given that something else happened

Conditioning is renormalisation — delete everyone the evidence rules out, and rescale what is left.

Two events fit in one square. Cut it left to right by who is ill and top to bottom by what the test says, and every tile's area is a joint probability. The test here catches 99% of the ill; drag the vertical cut to change how common the illness is:

prevalence = 10%
prevalence = 10%

At a 10% prevalence the ill-and-positive tile holds 0.099 and the healthy-but-positive tile holds 0.090 — almost the same area, out of a test that is right 99% of the time on the people it is looking for. The healthy column is nine times wider, so a small error rate inside it makes a comparable slab.

A positive result deletes the whole negative row, leaving one tile and one slab. What is left is not a distribution yet — it sums to 0.189, not 1 — so we stretch it back out until it does. Drag the lower bar's right-hand end across the stage:

stretched 0.00% of the way
stretched 0.00% of the way

Notice the ratio inside the bar never changes while you drag. P(ill | +) = P(ill ∧ +) / P(+) divides both parts by the same number, so the proportions are fixed and only the meaning of everything moves. The answer reads 0.524 at both ends of the drag.

Now make the illness rare, which is what a screening programme does. Hold the test at 99% in both directions and pull the prevalence down the log axis, out of the one in ten it opens on:

prevalence = 10%

At a prevalence the answer is 0.090. A 99%-accurate test, a positive result, and you are still 91% likely to be healthy — because 10.1 false positives arrive for every true one. Break-even sits at a prevalence of exactly , where the two error rates cancel; below it the base rate wins and no amount of test accuracy changes that, only moves the crossing.

This is the one failure on the page that is dangerous because it is quiet. Nothing throws, no number looks wrong, and P(+ | ill) and P(ill | +) are 0.99 and 0.09 for the same test. Independence is the special case where none of this arithmetic is needed — and it has a picture too. Pull the two cuts out of line:

offset = 0.00

At an offset of zero, P(A | B) = P(A) = 0.50 and the square is a plain grid. Pushed to , P(A | B) is 0.80 while P(A) has not moved — the marginal is held fixed by construction, so the step between the cuts is the whole of the dependence. Independence is the claim that that step is zero, and nothing looser.

05

Which model made this likely

Turn the question round: not what a parameter predicts, but which parameter makes what we saw most likely.

Suppose we flipped a coin twenty times and got thirteen heads. Every possible bias assigns that outcome some probability, and the curve of those probabilities — read as a function of the bias, with the data held fixed — is the likelihood. Drag the bias along the axis:

θ = 0.50
θ = 0.50

Watch the peak: it sits at 0.65, which is exactly 13/20 — for a coin, the maximum-likelihood estimate is the sample frequency. At θ = 0.50 the curve has fallen to 0.40 of the peak — fair is not ruled out, only outbid. Push the data to and the peak jumps to θ = 1.00: a model saying tails can never happen.

Two things about that curve are inconvenient: it is not a distribution over the bias, and its values are already minute at twenty flips. Take the logarithm:

θ = 0.50
log L = −13.86

Because log is monotone, the peak has not moved: whatever maximises L maximises log L. And the values are readable now — −12.95 at the peak against −13.86 at fair, where the raw likelihoods print as 2.4e−6 and 9.5e−7.

That is not a stylistic preference. Chain independent probabilities together and the product runs out of float64 long before a sequence gets interesting. Take a per-token probability a language model would be pleased with and push the number of terms out from one:

1 term

At p = 0.02 a token, the product is 1.3e−170 after and reaches exactly zero at — float64's smallest denormal is 4.94e−324. The sum of the logs reads −747.2 and carries on. Nothing throws: the product becomes 0, log 0 becomes −∞, and the first NaN turns up downstream in a gradient.

Flip the sign on that log and you have the loss every language model trains on: one curve, read at the probability the model gave the token that came next:

p = 0.500
p = 0.500

At p = 1 the loss is 0, the only free answer. At 0.5 it is 0.693 nats, at 0.2 it is 1.609, and at 0.002 it is 6.215. The loss is unbounded, so confident and wrong is punished without limit while uncertain is cheap — and that asymmetry is the entire training signal. It is also why a model that assigns exactly 0 to a token that then occurs produces an infinite loss and a dead run.

06

Drawing from it

A distribution describes. A sampler acts — and it needs exactly one random number to do it.

Lay the five probabilities end to end along a single unit of length. Ask the generator for a uniform u in [0, 1), drop it on the strip, and whichever block it lands in is the token. Drag the dart along:

u = 0.320
u = 0.320

That is inverse-transform sampling, and it is the whole algorithm. The block boundaries are the cumulative sums — the one the dart is in is written beside the strip as 0.000 → 0.577 — so a binary search over them turns one uniform into one token in O(log V). Each block is hit exactly as often as it is wide, which is the only property a sampler needs.

Sampling is also how we measure things too big to sum over: draw n, count, divide. The answer is only as good as n. Push the sample count out:

10 samples

Since both axes are logarithmic, the straight line is itself the claim: the interval falls one decade for every two the sample count climbs. Ten samples give ±0.31, give ±0.031, and ±0.01 takes 9,604. Monte Carlo does not care about the dimension — it cares about 1/√n.

Left alone a sampler will eventually reach into the tail, and the tail of a language model is mostly nonsense, so decoders cut it off. Slide the keep count down and watch the discarded mass appear:

k = 10

At the kept mass is 0.913 with half the vocabulary gone; at it is 0.448 and there is nothing left to sample — greedy decoding again, from the other side. Nucleus sampling inverts the control: name the mass you want, 0.9, and let k be whatever that costs, which here is 5.

Whether the tail matters at all turns on a number people rarely bother to compute. A one-in-a-thousand token is rare; a completion is long. Push the length out along the log axis:

1 token

At one token the chance is 0.001, as advertised. At it is 0.394, and at it is 0.632: 1 − (1−q)ⁿ passes 1 − e⁻¹ the moment the product qn in the readout reaches 1. Rare per draw and rare per generation are different claims, and the second is the one your users actually meet.

07

How surprised, on average

Entropy is the average surprise. Cross-entropy is the bill for being surprised by the wrong things.

Learning that a certainty happened tells you nothing; learning that a long shot happened tells you a great deal. −log p is that idea, with the arithmetic that makes independent surprises add. Drag the probability along the axis:

p = 0.500
p = 0.500

At p = 1 the surprise is 0, and at 0.002 it is 6.215, climbing without limit as p goes to zero. The dashed line at 0.693 is one bit — the surprise of a fair coin — and it is also the unit conversion, since nats times 1.4427 are bits.

Entropy is that surprise averaged over the distribution's own weights. Each outcome contributes −p log p; lay the six contributions end to end and the total length is the entropy. Concentrate the distribution and watch the bar shorten:

H = 1.79

Uniform is the maximum. Six equally likely outcomes give H = 1.79 nats, exactly log 6, and the bar reaches the dashed mark. Push the concentration and p(1) climbs to 0.831 while H falls to 0.73: a distribution that has nearly made up its mind costs almost nothing to transmit.

Now the model. It does not know p — it proposes q, and it pays −log q on outcomes that keep arriving from p. Drag the model away from the truth and watch a second block appear on the end of the bar:

H(p, q) = 1.67

At an offset of zero the two rows coincide and the bar is exactly H(p) = 1.67 — the model pays only what the uncertainty costs. Push it to and the bill is 2.05, of which 0.38 is the mismatch. That excess is the KL divergence: never negative, zero only when q equals p — which is why minimising cross-entropy is the right objective.

Nats are hard to feel, so the field reports exp(H): the number of equally likely options this hard to choose between. The marker opens on a real model's loss:

H = 2.86 nats

A loss of 0 is a perplexity of 1: no choice at all. Guessing blind over GPT-2's 50,257-token vocabulary is — the dashed line. The 1.5B GPT-2 scores 17.5 on WikiText-103 zero-shot — 2.86 nats, or 4.13 bits a token — so out of fifty thousand options it has narrowed the field to about eighteen. Cross-entropy and perplexity are one measurement in two units.

08

All of it, in one frame

Every section above turns up in the last four lines a language model runs.

A forward pass ends in exactly this: a row of logits, a softmax over them, one draw, and a loss. The true next token here is cat and the sampler has been handed a fixed u = 0.62. Drag that token's logit and watch all four stages move together:

loss 2.05

At the opening logit of 0.5 the model gives the true token 0.129 and the draw returns a different one — the loss reads 2.05 nats, a perplexity of 7.8. Raise the logit to and the probability goes to 0.973, the draw lands on cat, and the loss collapses to 0.03. Drop it to and the loss is 6.41.

p = softmax(z / T)      # z -= z.max() first
i = searchsorted(cumsum(p), uniform())
loss = -log(p[true])    # nats, 0 at p = 1
ppl = exp(mean(loss))   # effective choices

Training moves that logit; sampling reads that row. Run the four lines twelve times and the losses simply add, which is §05's sum of logs with a name on it. Score the positions one at a time:

1 of 12 token

The total after twelve tokens is 19.70 nats, a mean of 1.641, and — while position 7 alone, at p = 0.02, cost 3.91 of it. That is the whole subject in one bar: probability describes, the log makes it add, and the average is the number the paper reports.