Gradient Descent Primer

One line of arithmetic, repeated a few hundred thousand times, trains every model you have heard of. This primer takes it apart: the step itself, the one number that decides whether it converges or explodes, why momentum and Adam exist, how much data each step gets, and the schedule that carries a run to the end.

01

One step downhill

Every model you have heard of was trained by repeating one line of arithmetic: measure the slope, step against it, measure again.

Picture a marble on a hillside. The marble is the parameter vector θ; the hillside is the loss L(θ), one number for each choice of θ; downhill is the direction opposite the gradient ∇L. In one dimension the hillside is a curve and the marble is a dot on it.

The slope under the marble is the algorithm's only input: the tangent is steep on the flanks and flat at the minimum. Drag the ball and read ∇L off the tangent under it:

θ = 3.20. drag the ball along the curve; the arrow keys move it 0.2 at a time and Home returns it to 3.2
θ = 3.20

Notice that ∇L is signed, not a direction word: it reads +3.20 at θ = 3.20 and turns negative on the left flank. The gradient always points uphill. It also shortens as the ball nears the minimum, and that is the first thing the update rule exploits.

The rule is θ ← θ − η∇L: scale the gradient by the learning rate η, subtract. At η = 0 nothing moves. Raise it and watch the step grow:

η = 0 — no step yet

The minus sign is what turns a positive gradient into a leftward arrow. Watch the landing point: it reaches the minimum exactly at η = 1.00 and shoots past it at 1.05. The step is a product — a knob we choose times a slope we measure — so the same η behaves differently on different problems.

Everything rests on that sign. The switch runs the same four steps with θ ← θ − η∇L and with θ ← θ + η∇L:

θ ← θ − η∇L — downhill

The plus sign does not fail loudly. It runs, it returns numbers, and the loss climbs from 5.12 to 15.66 in four steps — a run that reports a rising loss and no error at all. Frameworks bury the sign inside the optimizer, so the version you meet is a hand-written gradient with its own sign wrong.

Back on the descending side, one property falls out for free. Take twelve steps one at a time at a fixed η = 0.35 and read this step's length as you go:

0 steps taken

Because the gradient shrinks with the distance to the minimum, so does the step — 1.12 on the first, 0.010 on the twelfth. Nobody scheduled that; it is the rule's own arithmetic, and it is why a converged run looks converged long before it is.

So the invariant, stated the way you would assert it: for a loss whose curvature is at most a, every step with η < 2/a moves the parameter closer to the minimum — L(θ − η∇L) − L(θ) ≤ −η(1 − ηa/2)‖∇L‖². That bound has an a in it, and nobody hands you a.

02

How big a step

One number decides whether training converges, crawls, or explodes — and the same number means different things on different losses.

The bowl from §01 has curvature a = 1, so the update is exactly θ ← (1 − η)θ. That single factor is the whole story: the distance to the minimum is multiplied by |1 − η| on every step.

Twelve steps, same bowl, with the rate under your hand. Small rates crawl; lands in one; past the ball leaves the bowl:

η = 0.20

Watch what happens between 1.00 and 2.00: the steps overshoot the minimum and land on the far side, still closer than before. Overshooting is not divergence. The run that alternates sides and keeps shrinking is converging; the one that alternates and grows is not, and the boundary between them is exact.

That boundary is a number you can plot. |1 − ηa| is what the distance gets multiplied by per step — drag the marker along it:

η = 0.20. drag left and right along the curve; the arrow keys move η by 0.05 and Home returns it to 0.20
η = 0.20

The V has a floor at η = 1/a, where the factor is zero and one step solves the problem exactly. Either side of it progress is geometric: at η = 0.20 the factor is 0.80 and 1% takes 21 steps; at η = 0.05 it is 0.95 and the same 1% takes 90. A quarter of the rate costs four times the steps.

None of that is what you see while training. What you see is the loss, one number per step, on a logarithmic axis. Set η and read the shape:

η = 0.20

Three shapes, one knob. Below η = 1 the trace falls straight, because on a log axis geometric decay is a line. Between 1 and 2 it saws downward. Past 2 it climbs, and it climbs geometrically — which is why a diverging run reaches inf in tens of steps rather than thousands.

The bound is 2/a, so the safe rate is a property of the surface, not of the optimizer. Drag the point across curvature and rate, and cross the boundary:

a = 1.0 · η = 0.20. drag left and right along the curve; the arrow keys move η by 0.05 and Home returns it to 0.20
a = 1.0 on the right

Notice that both curves are the same shape: η = 2/a is where it diverges, η = 1/a is where it solves in one step, and both move with the curvature. Double the curvature and you must halve the rate — which is why a rate found on one model does not transfer to another at a different scale.

In practice a is not one number: it is the largest curvature the run meets, and it moves as the run does. So the usual failure is not a clean divergence at step 0 but a run that trains for an hour and then blows up on a sharper part of the landscape.

03

Two directions, one rate

A real model has millions of parameters and they do not share a curvature. One learning rate has to serve all of them at once.

Two dimensions show the whole problem. The loss is ½(axx² + ayy²) — a valley, long one way and steep across the other. The ratio κ = ay/ax is the condition number.

Start with the geometry. Drag the point anywhere in the valley and watch which way the gradient would send it:

θ = (-0.88, 0.62). drag the point anywhere in the valley; the arrow keys move it and Home puts it back
θ = (-0.88, …)

The arrow points almost straight across the valley, not along it. At the opening pose it is 8.5 times longer across than along, so the step is essentially vertical — spent on the direction that is already nearly right. Gradient descent does not point at the minimum; it points down the steepest wall.

Forty steps from that point, with the rate under your hand. The steep direction caps how large it may be — push past and it diverges:

η = 0.40

Because the steep direction sets the ceiling, the best rate this valley allows is , and it still needs 28 steps to close the last 1%: the shallow direction is being crawled down at a rate chosen for the steep one. Turn η down to 0.40 and the zigzag goes away, but the count goes to 130.

How bad it gets is a function of one number. Take the best rate for each κ and stretch the valley out from a circle to :

κ = 1

Watch the count climb with κ. A circle is solved in a single step — one curvature, one perfect rate — and from there it is 5 steps at κ = 2, 28 at 12, 93 at 40, because the per-step contraction is (κ − 1)/(κ + 1), which approaches 1 from below. Conditioning, not dimension, is what makes optimization slow.

One cheap change fixes most of it. Keep a fraction β of the previous step and add it to this one — velocity — then step along that instead of along the raw gradient:

β = 0 — no momentum

Watch the crossing components cancel. Successive steps point up and down the same wall, so they subtract; along the valley they all point the same way, so velocity accumulates. At the same run takes 9 steps instead of 28 — the zigzag is spent on progress rather than on itself.

Momentum has its own failure, and it is not divergence. Push β past what this valley wants and read the distance to the minimum on a logarithmic axis:

β = 0.45

At β = 0.90 the run still converges — in 74 steps instead of 9, most of them spent travelling away from the minimum and back. Too much momentum is a heavy ball in a bowl: it does not fall out, it rolls around. The tell is a loss that oscillates over tens of steps.

Theory says the right pair turns κ into √κ: with η = 4/(√amax + √amin)² and β = ((√κ−1)/(√κ+1))² the contraction is (√κ−1)/(√κ+1) — 12 steps against 28 at κ = 12. Both settings need κ, and nobody hands you κ.

04

A rate for every parameter

The condition number of a real network is enormous, and worse: different parameters need steps of wildly different sizes.

Take six parameters out of one model — a LayerNorm gain, a bias, an attention projection, a rare token's embedding row. Their gradients span two orders of magnitude in one backward pass, and a single η multiplies all six.

Suppose each of them wants a step near 0.05, within a factor of two. Move and count how many of the six land in that band:

η = 0.02

Watch the count as η moves: at best two of six. The gradients span 120×, so the steps span 120×, and the band is only a factor of four wide — no choice of η beats that arithmetic. Either the largest parameter takes a step that overshoots, or the smallest takes one it will never finish.

Adam's answer is to divide by the gradient's own size. It keeps a running mean m and a running mean square v per parameter; scrub the counter and watch √v walk onto |g|:

after 1 step

Because √v converges to the gradient's own magnitude, m̂/√v̂ is ±1 whatever the scale, and every parameter moves one learning rate per step. That is what Adam buys: not speed, but a step size that no longer depends on a scale you did not choose. η stops meaning “times the gradient” and starts meaning “per step”.

Both accumulators start at zero, so early on they under-report — and their two decays are not the same speed. Leave the correction off, then walk the counter through the first two hundred steps:

step 1

β₂ = 0.999 warms up a hundred times slower than β₁ = 0.9, so the denominator is the one that is too small. Uncorrected, Adam's first step is 3.16 learning rates and it peaks at . It fails silently: the run trains, the loss falls, and the first few hundred steps were taken at six times the rate you set.

That state is not free. Both accumulators are a full copy of the parameters, in fp32. Pick a model and an optimizer and read the bill against one 80 GB card:

model 6.7B

Adam costs 16 bytes per parameter against SGD's 8 — two fp16 copies for weights and gradients, an fp32 master copy, and two fp32 moments. A model needs 100 GB for Adam and 50 GB for SGD, so the model that does not fit does fit the moment you give up the thing this section exists to explain.

This is also where the trip-up lives. Adam's weight decay is not L2 regularization: adding λθ to the gradient sends it through the same 1/√v̂ division, so parameters with large gradients end up decayed less. AdamW subtracts λθ from the parameters directly, after the division — a one-line difference worth real accuracy.

05

How much data per step

The gradient in the update rule is an average over the whole dataset. Nobody computes it. Every real step uses a sample.

One parameter, sixty-four training examples, one gradient per example. The full-batch gradient is their mean — the number the rule actually wants. A mini-batch takes B of them and uses that mean instead.

Grow the batch and watch the estimate walk onto the truth:

B = 1

At B = 1 the estimate reads 0.452 against a true 0.528 — and the sign of that error changes with every draw. What makes SGD work is that the error is zero on average: the estimate is unbiased, so over many steps the wrong parts cancel and the right part accumulates. Noise is not bias.

How fast the error shrinks is the ordinary square-root law. Drag the marker along the line and read the standard error off it:

B = 1. drag left and right along the line; the arrow keys change the batch size and Home returns it to 1
B = 1

Both axes are logarithmic, so the power law is a straight line of slope −½: buys half the noise, not four times less. Doubling the batch costs twice the compute and returns 1.41× the precision, and that ratio does not improve at scale.

The vocabulary falls out of the same picture. Cut one pass over the data into batches: each batch is one iteration and one optimizer step, and the whole row is one epoch:

B = 1

One epoch is one pass over the data; one step is one update. They are different units, and the ratio between them is ⌈N/B⌉ — which is why “we trained for ten epochs” says nothing about how many updates happened until you also know B. Schedules are indexed by steps, never by epochs.

Put real numbers on it. ImageNet's train split is 1,281,167 images and the classic recipe runs ninety epochs — drag the batch size and read both counts:

B = 256. drag left and right along the line; the arrow keys double or halve the batch and Home returns it to 256
B = 256

At B = 256 that is 5,005 steps per epoch and 450,450 updates in the run. Take the batch to and the same ninety epochs is 28,170 updates — sixteen times fewer chances for the optimizer to move, on exactly the same data.

It is bought for a reason. A step has a fixed cost that does not depend on B — kernel launches, the optimizer's own element-wise pass, one gradient all-reduce — plus arithmetic that does, which is what pushes throughput toward the ceiling:

B = 32. drag left and right along the curve; the arrow keys change the batch size and Home returns it to 32
B = 32

Notice where the curve bends. Below B = 80 the fixed 12 ms dominates and throughput is nearly proportional to the batch: 32 → 256 is 2.7× faster. Above it arithmetic dominates and the curve flattens onto the ceiling: is four times the batch for 1.22× the throughput and a quarter of the updates.

So batch size is a three-way trade: noise falls as 1/√B, wall clock per epoch falls until the hardware saturates, and updates fall as 1/B. The usual resolution is to raise the rate with the batch — linearly, to a limit — so the same distance is covered in fewer, larger steps.

06

Getting to the end

A run that converges is not a run that finishes. Three things decide when it stops and what it stops at.

Nothing so far has changed η during the run. Every large model does: it starts near zero, climbs for a few thousand steps, then decays to almost nothing. Each half of that shape fixes a different problem.

The warmup half is first. Drag the warmup length and watch what the first step is allowed to be:

no warmup

With no warmup the first step is taken at the peak rate — into a random initialization, on a gradient estimated from one batch, with Adam's second moment still empty. §04 showed that last one is worth 3–6× on its own. Warmup is not superstition; it is the cheapest way to survive the first few hundred steps.

The decay half is fixing something else. Run two hundred noisy steps at a constant rate and watch where the loss stops falling:

η = 0.10

The loss does not converge to the minimum; it converges to a cloud around it whose size the learning rate sets. The floor is ησ²/2(2 − ηa) — and the floor halves. That is what a decay schedule buys, and why the last fifth of a run, at almost no learning rate, still moves the number.

One thing left, and it is the one that wakes people up. A bad batch arrives at step 24 with a gradient twenty-six times the usual. Set the clipping threshold:

clipping off

Unclipped, one batch takes a step of 2.00 and throws the loss from 0.004 back to 1.84 — hours of progress, from one example. Clipping rescales any gradient longer than the threshold back to it, so is η·c whatever the data does.

So how do you know it converged? Not from the loss, which the noise floor dominates. The gradient norm is the honest signal: at a real minimum it goes to zero while the loss does not.

07

The loop, and the three optimizers

Two lines in the wrong place, three optimizers on one problem, and the eight lines that are every training loop.

Both mistakes below are one line out of position, and both run without an error. The switch puts one of them back:

the loop as written

Missing zero_grad() is the loud one — gradients accumulate, so the effective rate grows every step and the loss parks at 2.0. Stepping the schedule per epoch is the quiet one: it converges, to 8.7e−3 instead of 6.1e−4, because it never leaves the noise floor §06 measured.

Every line below has been one of the sections above:

for epoch in range(epochs):
  for x, y in loader:      # 1 iteration
    opt.zero_grad()        # grads add up
    loss = criterion(model(x), y)
    loss.backward()        # fills p.grad
    clip_grad_norm_(params, 1.0)
    opt.step()             # theta -= lr*g
    sched.step()           # per step

Two more that survive review: Adam's weight decay is not AdamW, and the rate for batch 256 is not the rate for 2,048. Last, the three optimizers on §03's valley, each at its own best constant rate:

SGD

Momentum wins outright: 12 steps against SGD's 28 and Adam's 98. Adam is a normalized method, moving each coordinate about η per step whatever the gradient, so its count scales with the distance rather than the conditioning. Its win is §04's: nothing was tuned per tensor.