Supervised Learning Primer

Nine houses, one number each, and everything a training run does. This page proves three things on those nine sales: that a model is a function picked out of a family you chose, that training is one arithmetic rule repeated until a slope reaches zero, and that a fit which is perfect on the data you have is the single most reliable sign that it will be useless on the data you do not.

01

What a label buys you

Nine houses sold, and we know the area and the price of each. Supervised learning turns those nine facts into a price for a house that has not sold yet.

A supervised problem always arrives as pairs. Each example carries a feature — something you can measure before the answer is known — and a label, which is the answer itself. Here the feature is floor area, the label is the sale price, and there are exactly nine of them.

The slider asks about one area at a time. On the nine areas somebody actually sold at, the label is there to read off; drag the query anywhere else and it is not:

1,500 ft² → 1,070k

Notice how little we are given: nine prices, and between them nothing at all. At the figure shows a question mark, because no such sale exists. Every method in this primer is a different answer to what goes in the gaps, and none is given more than this.

The cheapest answer is to draw one straight line and read prices off it. The line is ŷ = w·x + b; here b is pinned and the slider moves w, the price of one more square foot:

w = 200 $/ft² · rmse 592k

Watch the number on the right. That is the fit error — root-mean-square, so it is in dollars, like the prices themselves. At w = 200 the line is $592k out per house; drag to and it falls to $74k, which is where this family runs out of room.

Pinning b was a convenience, not a law. The second parameter lifts the whole line without tilting it, so grab the line itself: sideways tilts it, up and down lifts it:

w 200 · b 700k. drag the line: sideways tilts it, up and down lifts it; arrow keys nudge, Home restores the opening line
w 200 · b 700k · rmse 338k

Because both knobs move, the search is over a plane of lines rather than a row of them, and the best point on that plane is w = 484, b = $377k, at $74k. Two numbers. A frontier language model has on the order of a trillion, and the search is the same search.

A family is a bet about the shape of the world, and the bet can be lost. The switch trades the fitted function between four families over the same nine sales:

degree 0 · 1 number to learn

Notice that flat cannot get below $448k however you set its one number: that family has no way to say price rises with area at all. This is underfitting, and it is a property of the family rather than of the training. Wiggle, with nine numbers, reaches zero — and §04 is about why that is the worse failure.

So a supervised model is three decisions, in this order: what you measure, which family you search, and what counts as a good fit. The first is fixed here by the data we happen to have. The second is the bet above. The third is the whole of the next section.

02

What “wrong” is worth

A fit needs a score before it can be improved. Which score you pick decides which houses the model is allowed to get wrong.

For one house the mistake is easy to name: the gap between what the model predicted and what the house actually sold for. The whole design question is what to do with nine of them at once.

The gaps hang between each sale and the line. The slider moves the line, and all nine change together:

w = 300 $/ft² · gap 341k

The number on the right is their plain average, $341k at w = 300. Adding the gaps with equal weight gives mean absolute error — a perfectly good loss and almost nobody's default, because every house gets one vote no matter how badly it is missed.

The default squares each gap first, and that is not a formality: a squared gap is an area, and the loss is the total area the nine squares cover. The slider is the same slope:

w = 300 $/ft² · mse 149,965

Watch the worst house as you drag. Squaring makes the penalty grow faster than the mistake does, so even at one house of nine owns 21% of the loss. That is the personality of mean squared error: it would rather be a little wrong nine times than badly wrong once.

The model has one free number here, so the loss is a function of one number too — and a function of one number can simply be drawn. The slider is the same w; the curve is every fit error you could have had:

w = 300 $/ft² · mse 149,965

Because the model is linear and the loss is squared, this curve is exactly a parabola — not roughly, exactly — and its lowest point sits at w = 484. Training is the search for that point. the loss reads 1,006,452; at the bottom, 5,516.

Squaring has a failure mode, and it is one you can drive. One sale price is under your hand: drag it up and watch the squared-error line chase it while the absolute-error line ignores it:

1,333k. drag up and down to change what this house sold for; arrow keys nudge, Home restores the real price
1,333k · squared w 484

A square's derivative grows with the gap, so one bad label drags the whole fit toward itself; the absolute-error line stops caring the moment it is on the wrong side. Real datasets have bad labels. Squared error is the right default and the wrong one to leave unexamined.

Not every label is a number. When the answer is a class, the model outputs a probability for it and the cost is minus the logarithm of that probability — the slider sets it:

p = 0.50 · 0.69 nats

Notice the asymmetry. Being right costs almost nothing; being confidently wrong costs without bound. At the cost is 4.61 nats, nearly seven times the 0.69 of a coin flip — and that unboundedness is what makes a language model care about the token it nearly missed.

The number you actually care about is usually not the one you can descend. The slider moves one bias through eight scored examples, with the 0/1 loss and cross-entropy both watching:

threshold 0.00 · 0/1 loss 0.25 · cross-entropy 0.49

The 0/1 loss only counts, so it is a staircase: flat everywhere, and undefined exactly where it moves. There is no direction to follow. Cross-entropy is a smooth stand-in that agrees about which way is better, and it is what training actually descends.

So the loss is not a scoreboard bolted on afterwards; it is the surface the optimiser walks. Squared error makes that surface a parabola. Cross-entropy makes it convex in the scores. Accuracy makes it a flight of stairs, which is why nobody trains on it.

03

Getting to the bottom

Nothing so far has trained anything. Training is one rule that turns the shape of the loss into the next value of the parameter.

You cannot solve for the bottom in general — a linear fit has a closed form, a neural network does not — so the working method is local: measure which way is downhill, and take a step.

The bowl is the loss from the last section. Drag along it and the tangent follows — its slope is the gradient, the one number the whole of training runs on:

w = 200 $/ft². drag sideways to move along the curve; arrow keys nudge, Home returns to the opening point
w 200 · slope of the curve: -2,427

The sign alone tells you which way to move: slope negative, go right. At w = 200 it reads −2,427; at the bottom of the bowl it reads 0, which is not a coincidence but the definition of the bottom.

The magnitude decides how far. One step is w ← w − η·slope, and η — the learning rate, under the slider — is the only part of that rule you get to choose:

η 0.020 · w 100 → 166

Watch the arrow at η = 0.020: from w = 100 the step lands on 166 and the loss falls from 635,423 to 438,344. It undershoots because the slope was measured where the step started. Push to and one step clears most of the distance.

Undershooting is fine when you get to repeat. The transport runs twenty steps and the second slider sets η; the run starts at w = 0, which is the flat line:

step 0 of 20 · w 0

Each step multiplies the remaining distance by the same factor, so convergence is geometric rather than linear: at η = 0.020 the run is still at 472 after twenty steps, while at η = 0.050 it reaches 483 in twelve. Doubling η roughly halves the steps — up to a point.

That point is exactly locatable, and it is not a matter of taste. Push η past it and the same rule that was converging starts climbing out of the bowl:

η = 0.100

Because the curve is a parabola of curvature 8.553, each step multiplies the distance to the bottom by |1 − η·8.553|, so the run converges only below η = 0.234 and diverges above it. Nothing warns you at : the loss just becomes NaN a few hundred steps later.

And 0.234 is not a property of the problem — it is a property of the units. The switch remeasures the same nine areas in four different units and redraws the step sizes that still converge:

÷ 1 · η < 2.34e-7

The axis is logarithmic — every tick is ten times the last — so read the spans as ratios. Raw square feet cap η at 2.34e−7; dividing area by a thousand caps it at 0.234, a million times larger, because curvature carries the square of the feature. This is why pipelines normalise before they take a step.

Two things here survive into every optimiser you will meet. The step is proportional to the gradient, so a flat region moves slowly however long you wait; and the safe step is bounded by curvature, which is why Adam estimates a per-parameter scale rather than trusting one η across a billion of them.

04

Fitting is not learning

A fit error of zero proves nothing. The only question that matters is what the model does on a house it was never shown.

Which we can actually answer, because nine more houses sold in the same market over the same period and we kept every one of them out of the fit.

The held-out sales are drawn as rings, and the fit — an ordinary straight line — has never seen one of them. Drag them onto the plot one at a time:

0 held out

The figure now carries two numbers instead of one: fitted, measured on the nine sales the line was built from, and unseen, measured on the nine it was not. For a straight line they land at $74k and $76k. Close together, which is the whole virtue of a straight line.

Now spend capacity. It opens on the same straight line, and the slider sets the degree of the fitted polynomial — flat at 0, through every training sale at 8:

degree 1 · fitted 74k · unseen 76k

Watch fitted fall the whole way — to $0 at — while unseen turns around at degree 4 and ends at $220k. The model did not learn the market. It memorised nine houses and invented whatever it liked between them.

Both numbers, plotted against every degree at once: fitted falling, unseen turning, and the slider's own degree ruled through both:

degree 1 · fitted 74k · unseen 76k

The vertical axis is logarithmic — each label is twice the one below it — so what you read between fitted and unseen is a ratio. The two run together to degree 4 and separate after it. That separation is the generalisation gap, and it is the only diagnostic this section has.

The gap has a shape, and you can put your hand on it. The query area drags along the axis; past the last sale, the degree-8 fit is answering with nothing at all to go on:

2,000 ft². drag sideways to move the area being asked about; arrow keys nudge, Home returns inside the data
2,000 ft² → 1,391k

At — 500 ft² past the largest house it has ever seen — the fit predicts $81,854k, more than forty times the price of the most expensive house in the data. It raises no error and reports no low confidence. It returns a number, and a downstream system will use it.

There are two ways out, and the first is the one you cannot always buy. The slider grows the training set from eight sales to four hundred, with capacity held fixed at degree 4:

8 sales · unseen 108k

Both curves move, in opposite directions: fitted rises, because eight points are easy to please and are not, and unseen falls to $69k and flattens out. That floor is the noise — the part of every price our one feature does not contain, about $70k a house — and more data does not move it.

One more failure, and it is the one that actually ships. The slider makes some of the held-out sales duplicates of houses already in the fit: the same listing, entered twice:

0 duplicated · reported 220k · honest 220k

Watch the reported error fall while the honest one does not move at all. At nine duplicates of nine the report claims $12k on data the model has “never seen”, and every one of those nine is a house it memorised. This is leakage, and it fails silently.

The invariant is one sentence: nothing that touched the fit may appear in the score. Not the row, not a near-duplicate of it, and not a feature computed using it — a column normalised over the whole dataset before splitting has already leaked the test set's mean into training.

05

Pulling it back in

More data is the honest fix and often the unavailable one. The alternative is to make nine sales buy a smaller model.

Degree 8 is not wrong for having nine coefficients. It is wrong for being allowed to make them enormous: the wiggle between the training points is what a coefficient of 48,365 buys.

So charge for size. The fit now minimises the squared error plus λ times the sum of its squared coefficients, and the slider raises λ from almost nothing:

λ = 1.0e-7 · fitted 11k · unseen 174k

Watch the curve relax. Fitted gets worse the whole way — it must, because a penalised fit is by construction no longer the best fit — while unseen improves to before it, too, gets worse. Nothing about the family changed. Only what it costs to use it.

The penalty acts on every coefficient at once, so the honest picture is all eight of them, against the same slider. The largest is drawn in full and the other seven in grey:

λ = 1.0e-7 · max |c| 48,365

Both axes are logarithmic here — each label a hundred times the last — so a straight fall on this plot is a power law. The largest coefficient goes from 48,365 to 25 across the sweep. The fit does not lose the ability to bend; it loses the ability to bend sharply.

The trade has a best point, found the way degree 4 was found in §04 — by watching the held-out error and never the fitted one. The slider is λ again:

λ = 1.0e-7 · unseen 174k

The minimum in unseen — $69k at λ = 0.010 — is within a thousand dollars of what degree 4 already gave us. So the penalty did not beat a well-chosen family; it removed the need to have chosen one, which is worth a great deal when the families cannot be enumerated by hand.

The cheapest regulariser of the three is not a penalty at all. This is the same degree-8 fit, fitted by gradient descent from all-zero coefficients, and the slider is how many steps have run:

step 0 · fitted 1,351k · unseen 1,394k

Because descent closes the easy directions of a problem first, the fit passes through a good model on its way to the memorised one: unseen bottoms at $65k after and is back to $78k by 2,048. Stopping early is a regulariser you get for free, and it beats both of the others here.

All three moves shrink the set of functions the fit is allowed to reach. Weight decay is this penalty under another name; dropout, augmentation and a smaller model are the same move spelled differently. The knob is always chosen on held-out data, never on the training loss.

06

One step, end to end

Everything above is one loop. Here it is on the nine sales, with every number the loop produces on screen.

Real training does not use all the data on every step. It takes a minibatch — three of the nine here — because for the same arithmetic a noisy gradient computed thirty times beats an exact one computed once.

The transport walks one step through its five stages, over the batch, with w starting at 300 and η at 0.05:

stage 1 of 5 · take a batch

Notice the gradient at stage four: −1,698 on this batch against −1,572 on all nine. The batch is wrong by 8%, and that error is the entire cost of not looking at the whole dataset — repaid three times over in steps taken.

for batch in loader:            # 1
    yhat = model(batch.x)       # 2
    loss = mse(yhat, batch.y)   # 3
    g    = grad(loss, w)        # 4
    w   -= eta * g              # 5

That 8% is one draw. The switch takes the same nine sales in batches of one, three or nine and draws every batch's estimate of the exact slope:

batch of 9 · spread ±0

At a batch of nine there is one estimate and it is exact. At three the spread is ±311. At one it is ±1,716 — larger than the slope itself, so a single sale can point twice as steeply as the truth or barely at all, and averaging is what buys the accuracy back.

Those five lines are supervised learning. A frontier run changes only their size: line 2 becomes a trillion parameters, line 3 cross-entropy over a vocabulary, line 4 backpropagation, and eta a schedule — GPT-3's 175-billion-parameter run peaked at η = 6 × 10⁻⁵ and warmed up over its first 375 million tokens (Brown et al., 2020). The questions do not change: is the held-out loss falling, and is anything leaking into the batch.