Calculus Primer
The calculus a training loop actually runs on, built out of three pictures: a curve with a line touching it, a bowl seen from above with an arrow on it, and a graph with the derivative walking back down it. Sensitivity, the chain rule, the gradient, the step, and the four ways it fails — every number on the page is computed by the figure beside it, so you can drive it anywhere and it stays true. No integrals: you can train a billion-parameter model without ever computing one.
Slope is sensitivity
One number that answers the only question an optimiser ever asks: if I nudge this input, how far does the output move?
Training is a search, and at every step something asks the same question of every one of a billion knobs — turn this one a hair; does the loss get better or worse, and by how much? The answer is one number per knob. Everything below is the machinery that produces those numbers cheaply, built on one curve. Drag the point along the curve and watch its two coordinates follow:
A function is only a rule that turns one number into another, and the point is that rule seen once. It is not a rate yet: a rate needs two readings and the distance between them. Take a second point h further along, join the two, and the line you get has a slope you can compute. Shrink the step and watch the gap at its far end close:
Notice how fast the readings settle. At h = 1.40 the chord reads 2.05; at it reads 0.21; at , 0.05. The chord is walking towards 0, and that limit is the derivative — written f′(1), or df/dx when you want the two quantities named. Same object, three spellings, all three in every paper.
The limit turns the chord into a line that touches the curve at exactly one place and leans the way the curve leans there. That is the tangent, and its slope is the derivative at that point. Drag the point along and watch the line tilt with it:
Notice that the slope is a number at every x, so it is a function in its own right: f′(x) = x² − 1. It is positive where the curve climbs, negative where it falls, and exactly zero at the two grey ticks — and — where the curve has levelled off. Those are the flat spots an optimiser is hunting for.
Because it is a function, it can be drawn. Here is the curve above and its own slope below, on a shared input axis; drag anywhere and the two move together:
The lower curve crosses zero exactly where the upper one turns over. It is worth being clear that the two frames are not drawn to one scale — read the slope off the readout, not off the angle. What the picture does say is that one formula carries the climbing rate of every point at once, which is what makes a derivative worth computing symbolically rather than by measuring chords one at a time.
There is a stronger way to say what the tangent is, and it is the property everything later depends on: near enough to the point, the curve and the line are the same object. Halve the window and see what happens to the worst gap between them:
Because the gap falls by roughly four each time the window halves — gives 0.112, gives 0.027 — the linear model wins as you zoom in. That is the whole content of differentiable: f(x + h) = f(x) + f′(x)·h + (error) where the error dies faster than h does. Every gradient step in every framework is that line being trusted for one small h.
So the failure mode is a function that never becomes straight, no matter how far you zoom. ReLU is exactly that at the origin, and it is the most-used activation in deep learning. Slide the point onto and read the two one-sided slopes:
At the corner the answer from the left is 0 and the answer from the right is 1, so there is no single tangent and f′(0) does not exist. Nothing raises an error: PyTorch, TensorFlow and JAX all return 0 for relu'(0), silently taking one of the two answers and dropping the other. JAX only since 2022, when it pinned the derivative with a custom JVP; before that it returned 1. Any value in [0, 1] is a valid subgradient, so it costs nothing — but a library that changed its mind at a version boundary is worth knowing about before you go hunting elsewhere.
Many knobs at once
A model has billions of inputs, not one. The fix is to ask the one-variable question once per variable, and to be careful about what “holding the rest still” costs.
Give the loss two parameters instead of one and it stops being a curve and becomes a landscape: every pair (w₁, w₂) has a height L. We will look at it from directly above, so each ring joins the pairs that cost the same. Drag the point around and read the height it stands on:
This is the shape a squared-error loss really has near its minimum, and the two things to see are both about the rings. They are ellipses, not circles — the surface is stiffer across the valley than along it — and they crowd together where it is steep. At the centre they close on L = 0, the minimum training is looking for.
Now pin w₂ and let only w₁ move. That cuts one curve out of the landscape, and on that curve we are back in section 01 with a single variable. Slide along the cut, then change which cut you are on:
The slope of that tangent is the partial derivative, written ∂L/∂w₁ with a curly ∂ instead of the straight d. The curl is not new mathematics; it is a note to the reader saying which variables were held still. Computing one is the same: treat every other variable as a constant and use the section-01 rules.
Standing at one point there are two such questions, one per axis, and two answers. Move the point and read both — the rate along w₁ and the rate along w₂, as two arrows leaving it:
At the opening point the first reads −2.34 and the second 6.52: pushing w₁ up lowers the loss, pushing w₂ up raises it steeply. A function of D variables has D of these, one per variable, and the recipe never changes — only the bookkeeping grows.
One thing about “held still” catches people, and it is worth feeling rather than reading. ∂L/∂w₁ is itself a function of every variable. Leave w₁ exactly where it is and move the other one:
Watch the tangent tilt even though nothing touched w₁. At w₂ = 1.60 the slope reads −0.76; at it reads 2.70; at , 6.16. It even changes sign, at w₂ = 1.25. So a gradient is a fact about a point, not about a parameter: move any other weight in the network and every partial you already computed is stale.
Rates multiply
A network is functions feeding functions. The rule for getting a derivative through the whole stack is one multiplication per link — and nothing harder than section 01 ever appears.
Two links: x goes into g to make u, and u goes into h to make y. Both frames below share their middle axis, so a nudge entering on the left leaves on the right having been resized twice. Shrink the nudge from its opening width and watch all three intervals close together:
Watch which numbers agree. At the opening nudge Δu/Δx = 1.80 and Δy/Δu = 0.086, and their product, 0.155, is exactly what Δy/Δx reads. That identity is not an approximation and it does not depend on the nudge being small — Δu cancels, the way any two fractions multiplied end to end cancel their shared term.
Shrink the nudge to nothing and each of those three ratios becomes a derivative, so the identity survives the limit intact. Here are the two local slopes as tangents; drag the input and watch both of them change together:
That is the chain rule: dy/dx = h′(g(x)) · g′(x), or in the form that shows the cancellation, dy/dx = dy/du · du/dx. At the opening point the two factors are 0.142 and 1.40, and the answer is 0.198. N links, N factors; the recipe never gets harder.
There is one trap in it, and it is the mistake everyone makes once: h′ must be evaluated at the value the forward pass left there, not at x. Drive the input to and the figure reads 0.083; writing h′(x)·g′(x) instead gives 0.190, which is 2.3× too big and silently wrong. The forward pass supplies the right evaluation points, and that is why a framework runs forwards before it differentiates backwards.
Because the factors multiply, their sizes compound. A squashing link contributes a factor under one. The logistic curve is the classic offender: drag along it and watch its own slope below:
The lower curve peaks at exactly 0.25, at , and it is under 0.007 by . That number is the whole story: σ′ = σ(1−σ) is a parabola in σ with its maximum at σ = ½, so no sigmoid layer anywhere can contribute a factor larger than a quarter.
Stack those factors and watch what a long chain does to a rate. Set the gain each layer contributes and how many layers there are, and watch what each link bites out of what reached it:
At a gain of 0.75 — a mild squash — twenty-four layers already leave one part in a thousand — a beam flat on the floor of the frame long before the chain ends — and leave 3.19×10⁻⁸. Drop the gain to a sigmoid's best case, , and ten layers alone leave 9.54×10⁻⁷. That is the vanishing gradient, and it fails silently: no error, no NaN, just early layers whose updates are rounded to nothing while the loss sits still. Push the gain above one instead and the same multiplication explodes — which does announce itself, as a NaN.
All the partials, one arrow
Pack the per-variable answers into a vector and it acquires a property none of them had alone: it points the steepest way up.
Write the two partials from section 02 as the two components of one vector and you have ∇L = (∂L/∂w₁, ∂L/∂w₂), the gradient. It lives in the same space as the parameters, so it can be drawn on the map as an arrow. Move the point and watch the arrow against the ring it stands on:
The arrow meets the ring at a right angle everywhere, and that is forced rather than lucky: moving along a ring does not change L at all, so the rate in that direction is zero, so the gradient has no component along it. At the opening point ∇L reads [−2.34, 6.52] with length 6.92.
Calling it the steepest direction is a claim, and it is checkable. The climb rate along any unit direction û is ∇L · û — the gradient's shadow on that direction. Sweep the shadow right round and find its maximum:
Notice where the shadow is longest — , where it reads 6.92 — the gradient's own length, in the gradient's own direction. It reads 0.03 at — along the ring, where L does not change at all — and it is most negative at the opposite pole, which is the steepest way down. A projection cannot beat the length of the thing being projected; that is the whole proof.
One thing the gradient is not is a route. It is the best direction for an infinitesimal step, not a bearing to the minimum, and the two part company as soon as the bowl stops being round. Stretch the bowl and watch the angle between downhill and the minimum:
At the rings are circles and the two arrows agree exactly, at 0°. By κ = 6 they are 37° apart, and at , 44°. Steepest descent walks off across the valley instead of down it, and that one number — the ratio of the surface's stiffest curvature to its softest, its condition number — is what the next section is about.
The step
One line — w ← w − η ∇L(w) — and two ways to get it wrong, both of which you can cause here by hand.
The gradient points uphill, and we want down, so we subtract it. How far we move is not something calculus decides: the derivative is only valid in the limit, and any real step is a bet on how far the tangent stays honest. The learning rate η is that bet. Set it and watch where one step lands:
Because the surface curves away from the tangent, bigger is not simply better. From L = 4.63, a step at η = 0.10 lands at 1.23; at it lands at 0.53, which is the best a single step can do from here — the figure marks it, because that is where the line of possible landings just grazes the smallest level curve it can reach instead of cutting through it, at η★ = ∇ᵀ∇ / ∇ᵀH∇ = 0.171. Past it the step climbs back out: at the step lands at 0.64 again, and at at 5.52 — higher than it started.
Training does this over and over: compute the gradient at wherever you now are, step, repeat. Run it and drive both ends of the trade-off — the step size and how many steps you take:
Notice the zig-zag. The path crosses the valley instead of running down it, for exactly the reason section 04 gave — and it is the shape of the bowl that puts it there, not the step size. Shrinking η damps the swing and slows the descent in the same move: twelve steps reach L = 0.223 at , 0.0606 at η = 0.10, and 0.0036 at . All three are below the ceiling this bowl imposes, so the small rate is not the safe one, only the slow one — the quickest is , twelve steps to 6.6×10⁻⁴.
There is a hard ceiling above that, and it is not a matter of taste. Along each curvature direction a step multiplies the error by 1 − ηλ, which shrinks it only while |1 − ηλ| < 1. Plot that multiplier for the stiffest direction and the softest, and watch where it crosses one — drag η and find the crossing:
At the opening rate the stiff direction shrinks its error by ×0.40 a step and the soft one by ×0.90, so the soft direction is what costs the time. The stiff curve reaches one at exactly 2/6 = 0.333, and past it every step makes that direction worse: at twelve steps land at 9.92, at at 1.31×10⁶. In a real run that is the loss reaching NaN a few hundred steps in — the loud failure, and the easy one to diagnose.
The quiet one is the ceiling itself moving. λ_max is a property of the surface, so a bowl that is narrow in one direction forces a small η on every direction, including the ones that needed a big one. Stretch the bowl and count the steps:
Since each run uses the best η for that bowl, this is conditioning alone, with the learning rate already tuned: needs 7 steps, κ = 6 needs 21, needs 70 — growing in proportion to κ. Measured Hessian spectra for trained image classifiers put the largest eigenvalue in the hundreds against a bulk sitting near zero (Ghorbani, Krishnan and Xiao, 2019), so the ratio a real run faces is far worse than six — which is why plain gradient descent is not what anyone ships: momentum, Adam and layer norm are all, in the end, ways of making the bowl rounder.
Backpropagation
Nothing in this section is new mathematics. It is section 03 applied to a graph, in the one direction that makes it affordable.
Here is the smallest network with something to say: z₁ = w₁x, then a = tanh(z₁), then z₂ = w₂a, then L = ½(z₂ − y)². Running it forwards fills in the values; then walk the derivative back from the loss, one node per step:
Each backward step is one multiplication by a local derivative, taken at the value the forward pass left there. ∂L/∂z₂ = z₂ − y gives 0.834; times w₂ gives ∂L/∂a = 1.334; times tanh′ = 1 − a² gives ∂L/∂z₁ = 0.407; times x gives ∂L/∂w₁ = 0.407. Four multiplications, no calculus past section 01.
The direction is the whole trick. The chain rule works either way round, but a network has one loss and millions of parameters, and that asymmetry decides everything. Compare one pass per parameter against one pass for all of them on one graph of six weights — walk the slider up and count what each direction costs you:
Notice what six drags bought and what one did. Going forwards you propagate one input's influence through the graph, so every notch lays down another route across the spine and lights exactly one answer: with n parameters, n traversals. Going backwards you propagate one output's sensitivity, and one traversal hands you every ∂L/∂wᵢ at once. For a 7-billion-parameter model that is the difference between one backward pass and seven billion forward ones.
What reverse mode costs instead is memory: every intermediate from the forward pass has to stay alive until the backward pass reaches it. Set the depth and compare keeping all of them against keeping a few and recomputing:
At the plain backward pass holds 48 tensors and the checkpointed one holds 14 — √n checkpoints plus the √n activations of the one segment being replayed, so the peak is 2√n, for about 30% more compute (Chen et al., 2016). That trade is why activation memory sets the batch size.
The whole of it
Six rules, one loop, four ways to break it.
d/dx [c] = 0 d/dx [eˣ] = eˣ d/dx [xⁿ] = n·xⁿ⁻¹ d/dx [ln x] = 1/x d/dx [f + g] = f′ + g′ d/dx [f(g)] = f′(g)·g′ loss = forward(w) # every intermediate kept alive g = backward(loss) # one pass, every ∂L/∂wᵢ w -= eta * g # eta < 2 / largest curvature
All of it rests on one property: near a point, a function and its tangent agree to better than first order. Four things break it — a kink, a product of small gains, a step past 2/λ_max, an ill-conditioned surface — and only the third one announces itself.