Neural Net Primer
A network built from one unit up, with every claim on the page drawn as something you can drive. What a single unit computes, why a stack of them is worth nothing without a fold, which fold to pick, why a layer is one matrix multiply, what width and depth each buy — and the forward pass that ties them together.
What one unit computes
A unit takes a few numbers in and sends one number out. Three choices make it: a direction, an offset, and a fold.
The line everyone quotes is output = act(W·x + b). Written that way it is arithmetic, and arithmetic is easy to memorise and hard to picture. Two inputs are enough to draw every part of it, so this whole section lives on one plane with the input as a point on it.
Start with the multiplication. The input is a point on that plane, the weight vector is a direction, and w·x is how far along that direction the input reaches — drag the input anywhere and watch the thick amber segment:
Notice what w·x does not measure. It is not how big the input is: move it perpendicular to the weight arrow and the number does not change at all. A dot product is a similarity score, and the weights are the pattern this unit is looking for.
On its own that score has no zero point. The bias supplies one: z = w·x + b, and every input with z = 0 lies on a straight boundary. Drag the bias and watch the boundary leave the origin:
The boundary never turns — it slides, by exactly −b/|w|, which is what the violet arrow measures. At the teal sample lands on the silent side without having moved. Without a bias every boundary in the network would be forced through the origin.
Turning is the weights' job. Leave the bias alone and rotate the weight vector: the boundary stays perpendicular to it and sweeps the plane, and the sample's score changes sign twice per turn:
So the n weights choose a direction in n-dimensional space and the one bias chooses how far along it the boundary sits. That is n + 1 numbers per unit, and it is exactly what training has to learn. Everything a single unit can express is in those two choices.
One piece is still missing. Walk the input along the weight direction and plot what comes out — the pre-activation is a straight line, and the rectifier bends it once, at the boundary:
Watch the kink. A rectified unit is a hinge: flat on one side, a straight ramp on the other, jointed exactly where the boundary crosses, at t = −0.23. Slide to and the unit outputs nothing at all. That one bend is the whole non-linearity — and the next section is about why a network without it is not a network.
Why the fold is not optional
Two linear layers in a row are one linear layer. The rectifier is the only part of a network that changes what can be represented at all.
Stack a hundred matrices with nothing between them and you have bought nothing. Matrix multiplication is associative, so W₃(W₂(W₁x)) is (W₃W₂W₁)x — one matrix of the same shape, with the same expressive power a single layer had.
Here are two of them. A grid of input points goes through layer B, then layer A; the amber outline is where the single product matrix sends the same square. Drag the slider through both layers:
Notice that the grid lands exactly on the outline — , every time. Two 2×2 matrices spent eight numbers to produce a map that four numbers express. A deeper stack changes nothing: a composition of linear maps is a linear map.
The limit is easy to feel. XOR wants a 1 when exactly one input is on. Turn and slide the boundary and try to get all four labelled points onto the side that matches their label:
Three. Always three — sweeping every angle and every offset never reaches four, and the fourth point is always wrong. One unit can only cut the plane in two, and no single cut separates two diagonally opposite corners from the other two. That objection stopped neural networks for most of the 1970s.
The rectifier's contribution is that it is not a cut, it is a fold. Everything on the silent side of a unit's boundary is flattened onto it. Drag the slider to close the fold:
Watch where the clipped points go. They do not slide along the plane, they collapse onto the boundary, because their output is zero and zero is one number however you reached it. The layer throws information away on purpose — that is what makes it non-invertible, and useful.
Two folds are enough for XOR. Both hidden units look in the same direction, x₁ + x₂; the second one subtracts the far corner back off. Install them one at a time:
Watch the score rather than the lines. Nothing installed gets two of four for free, h₁ gets three, and h₂ changes no answer at all until the output subtracts it — and the score is four. The region the network answers 1 in is the strip between the two folds, which no single unit could express.
Choosing the fold
Which non-linearity you pick decides how far a gradient can travel back, and how many units die on the way.
Four functions cover almost everything that ships: relu, sigmoid, tanh and gelu. What matters about each is not the curve but its slope, because the slope is the only thing backpropagation multiplies by — once per layer, on the way back.
The amber curve is the activation and the teal one is its slope. Move the marker along z and switch between the four; the readout gives both values at the point you are standing on:
Look at the slope, not the output. relu's is exactly 1 wherever the unit is on, so a gradient crosses it unchanged. Sigmoid's peaks at 0.25 at and falls away on both sides. That factor of four per layer decides whether a deep stack trains at all.
Follow the sigmoid out from zero on an axis of its own. The vertical one is logarithmic — each tick down is a factor of ten — and the rose rule marks a slope of one part in a hundred:
At its very best, z = 0, the slope is a quarter, so ten stacked layers leave a millionth of the gradient. Push to and one layer alone costs a factor of four hundred. That is saturation, and it is why the 1990s stopped at two hidden layers.
relu has the opposite failure and it is permanent. Its slope is exactly 0 on the silent side, so a unit whose b has drifted below every input in the batch has no gradient left to climb back with:
Drag to and all sixteen bars go rose: the unit outputs zero for every example, so its gradient is zero for every example, so nothing updates it ever again. It is not slow, it is dead. Leaky relu and gelu exist because a slope of exactly zero is one bad step away from a permanently lost unit.
A layer is a matrix
Not a bag of neurons. One matrix multiply, whose shape is the whole of a layer's cost.
Put m units side by side, each reading the same n inputs, and their n-long weight vectors stack into an n × m matrix. One product computes all m pre-activations at once, which is why a GPU can run a layer as a single instruction stream.
Before the counting, the picture. A matrix sends every input direction to some output direction — drag the teal input around the unit circle and watch the amber output trace what it becomes:
Notice that the circle becomes an ellipse and never anything else: a linear map stretches and turns, and nothing more. The output runs from 0.94 to 1.30 times as long depending on direction, and the ratio of those two is the layer's condition number — the factor by which it distorts a gradient.
Now the counting. Each unit owns one column of amber weights and one violet bias. Drag units into the layer and read the parameter count off the grid you are building:
over eight inputs is 8×12 + 12 = 108 numbers, and both terms matter: weights grow as n·m and biases only as m, so a wide layer over a wide input is almost entirely weights. Double the width and you double the parameters and the arithmetic together.
Batching changes the arithmetic and not the weights. Stack B inputs as rows and the layer becomes one matrix–matrix product — but only if the inner dimensions agree. Drag W's rows away from 128:
The two inner edges are drawn to one scale, so is a picture rather than an assertion: the sides do not meet, and there is no product. Batching also buys arithmetic intensity — at batch 1 the layer does 0.5 flops per byte of weights it reads, at batch 32 it does 16. An H100 SXM does about 67 TFLOP/s of plain fp32 against 3.35 TB/s of HBM3, so under roughly 20 flops per byte it is waiting on memory, not computing.
Width against depth
Both buy straight pieces, and not the same number of them. Width is linear in the parameters; depth is exponential.
A relu network with one input and one output is a piecewise-linear function — always, whatever its weights. So “how expressive is this architecture” has an exact answer for once: how many straight pieces can it produce?
One hidden layer of k units gives at most k + 1 of them. Add units and watch the network close on the target curve; the readout gives the worst error anywhere on the interval:
Each unit contributes exactly one kink. Going from to takes the worst error from 0.182 to 0.059 — about threefold for a doubling of the layer, which is the usual return on width. Pieces arrive one at a time.
Depth buys them differently. Two relu units fold the interval in half; the next layer folds the folded thing again, so the pieces multiply instead of adding. Add layers:
of two units — ten units, 31 parameters — draw 32 straight pieces. A single layer needs 31 units and 94 parameters for the same count. Depth composes where width concatenates; this is Telgarsky's separation result, and the figure is its construction.
Put both on one axis. The vertical one is logarithmic, ten times per tick, and the slider spends a single parameter budget two ways — all on width or all on depth:
At the wide layer reaches 21 pieces and the deep stack reaches 1,024. Read the shape, not the height: on a log axis a straight line is exponential growth. What this does not say is that training finds those pieces — it is a bound on what the architecture can express, not on what gradient descent will build.
Depth has a bill, and it arrives on the way back. Every layer multiplies the gradient by one more factor, so the product is exponential in exactly the same way. Push the layer count past the safe band:
At a gain of 0.8 per layer, reach 1.4×10⁻⁵ and 100 layers reach 2×10⁻¹⁰ — the update is a rounding error. At 1.2 the same 50 layers reach 9,100 and the loss goes to NaN on the first big step. Residual connections, normalisation layers and careful initialisation all exist to hold that per-layer gain near 1.
The forward pass
Nothing left to invent: §4's matrix three times, §2's fold between them, and a vector changing shape at every step.
Take a four-number input and a 4 → 6 → 4 → 3 network. Every intermediate value is a vector, and drawing those vectors as bars rather than as numbers makes the thing that matters visible: how much of each layer's output survives the rectifier.
Here is the first layer alone. The top row is the six pre-activations for one input, the bottom is what the rectifier leaves of them. Sweep the input along one direction:
Notice how many bars are gone. At the resting input three of the six units are above zero, so half the layer's output is exactly zero — not small, zero. Relu networks are sparse by construction, and that sparsity is why the bottom row carries strictly less information than the top one.
Now the whole pass. Seven stages, each one either a matrix multiply plus a bias or a rectifier; press play, or drag the scrubber, and watch the vector change length and shape:
Watch the widths rather than the values: four numbers in, then six, then four, then three out. Every stage is one of two operations, and sixty multiply-adds and ten comparisons produce the whole answer at this size. Every one of those vectors also has to be kept, and that bill scales with the batch where the weights do not:
Real models are these shapes with bigger numbers. Here is the classifier every course starts with, 784 → h → h → 10, one bar per weight matrix. Drag the hidden width:
Notice where the parameters are. At h = 128 the model holds 118,282 of them and 85% sit in the first layer: 784 inputs is a great deal to read. That is 236,032 flops and 462 KB of weights at fp32. A batch of 128 carries 525 KB of activations on top — more than the weights — which is why costs training far more than it costs inference.
The output layer, and the seven lines
The last layer produces logits — numbers with no scale. Softmax gives them one.
Logits are unbounded and unnormalised: nothing in the network has any reason to produce a value in [0, 1]. Softmax exponentiates and divides, which turns any vector into a distribution while leaving the ranking untouched.
Drag the first logit up and watch the other three probabilities pay for it, then drag the temperature — dividing every logit by T before exponentiating is the one knob between a confident answer and a flat one, and at the top class takes almost everything:
Notice that the probabilities always sum to 1 and the order never changes. The whole forward pass is seven lines:
def forward(x, layers):
for W, b in layers[:-1]:
x = np.maximum(0, x @ W + b)
W, b = layers[-1]
z = x @ W + b
z -= z.max()
return np.exp(z) / np.exp(z).sum()z -= z.max() is not decoration. Softmax is shift-invariant, so subtracting the maximum changes no answer — and without it exp(1000) overflows and the vector comes back NaN, with no error raised.