Transformer Block Primer
A block is where the pieces finally meet: attention and a feed-forward network as two sublayers of one function, a residual connection so depth doesn't erase the gradient, a norm in exactly the position that keeps that connection honest, and N independent copies of the whole thing stacked into a model. Every number on this page is read off the figure beside it, derived from the same six lines of pseudocode the page ends on.
Two jobs, one shape
Attention moves information between tokens; the feed-forward network moves none at all. A Transformer block runs both in sequence, at one unchanged width throughout.
Everything on this page happens inside one repeating unit: the block. A modern language model is dozens of identical copies of it, stacked front to back, each one taking a sequence of vectors and handing back a sequence of vectors in exactly the same shape.
Inside that block, two very different operations run one after another: attention, which lets every position read every other, and the FFN, which never lets a position see its neighbors at all — toggle between them and watch what feeds token 2's output:
Notice that switching to FFN collapses the fan-in to a single vertical line: only position 2 feeds position 2. That is the entire difference between the two sublayers in one picture — attention is the only place information crosses between positions, and it runs first, so the FFN afterward always has whatever context attention already gathered.
Follow one token's vector through that sequence: it arrives, attention updates it, the FFN updates it again — drag through the three stages and watch d_model itself, not just the values inside it:
Watch the width, not the numbers: six bars enter, six leave attention, six leave the FFN. Every operation preserves d_model exactly — the only contract it has with whatever runs next.
That contract holds at any width. Slide through six real models, from up to , and watch the same three checkpoints stay locked together:
That is why attention and the FFN are “sublayers” of one thing: they share an input and output space, so a block is a single function from R^d_model to R^d_model — the property that lets N of them stack without reconciling shapes at the seams.
Keeping the scale honest
Left alone, activations drift by orders of magnitude with depth. Normalization is the fix — and LayerNorm and RMSNorm differ by exactly one step.
A block has no built-in floor or ceiling on its numbers. Each sublayer multiplies, sums, and reweights the vector it receives, and nothing forces the result back to a sane range — left unchecked across enough layers, that compounds. Model each layer as multiplying the vector's size by a fixed factor, then stack 24 with nothing to correct it: drag the factor slightly above or below 1 and watch where 24 layers lands:
Watch how little room there is on either side of 1.0: a factor of 1.4 per layer reaches over three thousand times the starting size by layer 24, and a factor of 0.6 falls under five millionths of it — and a real block runs dozens of operations per layer, any one of which can nudge that factor away from 1 with nobody watching.
LayerNorm is the fix, in exactly two steps on one eight-dimensional vector: subtract its own mean, then divide by its own standard deviation — step through and read the numbers it actually computes:
Notice both computed numbers belong to this one vector alone: μ and σ are recomputed independently for every token, at every position, every time. Nothing here is shared across the sequence — which is exactly what makes LayerNorm safe on a batch of different-length inputs, unlike BatchNorm, which needs the batch.
RMSNorm keeps only the second step. Toggle between the two on the same vector and read the mean each one leaves behind:
Because RMSNorm never subtracts a mean, its output keeps whatever mean the input had — the dashed line sits off zero. That recentering matters less than it looks: both reach comparable loss, and RMSNorm gets there in fewer operations.
Fewer operations, concretely: LayerNorm computes a mean, a variance, and a rescale; RMSNorm skips the mean step entirely. Drag through six real widths and compare the two op counts directly:
At GPT-3's width the gap is tens of thousands of elementwise operations per vector, on every one of its , every token, every forward and backward pass — the arithmetic behind the oft-cited “RMSNorm is roughly 7–10% faster end to end”: not less work in principle, measurably less of it at every call site.
The identity is free
Add the input back to the output, and a block's easiest possible behavior becomes doing nothing at all. That one addition is why deep stacks train.
Take any sublayer that maps a vector x to some f(x), and change what it hands onward to f(x) + x. The block no longer has to learn the whole output from scratch — it only has to learn the correction on top of what it already received.
Toggle the residual connection off and back on and watch what the block hands onward:
Notice that with the residual on, an f that outputs all zeros makes the whole block the identity — y = x, exactly. A stack of such blocks does nothing at all end to end, which is a far friendlier place to start optimizing from than a stack that has to learn a faithful copy of its input as step one of learning anything else.
Identity-as-default is only half of it. The chain rule on y = f(x) + x gives dy/dx = df/dx + 1, and that +1 is the residual path's whole contribution — model each layer's own gradient as a shrinking factor and drag the depth the gradient has to cross:
Watch the two curves split apart within a handful of layers, not gradually: past the plain chain has already lost most of what a residual chain still carries whole. A gradient that reaches layer 1 at effectively zero strength cannot update layer 1's weights — the layer is present in the model and absent from training.
That +1 is a real second path, not a brightness dial on the same one: drive a stack of blocks and switch the skip connection between present and structurally removed:
Because the rail is either drawn or not, there is nothing to misread as “a little weaker” — remove it and the gradient reaching the first block is exactly the plain chain's number, at any depth.
This is not unique to Transformers. Compare how deep a plain stack could be trained, before and after residual connections existed:
Before 2015, adding layers to a plain convolutional stack past roughly twenty made training worse — the deeper network could not even match a shallower one, let alone beat it. ResNet-152 trained without special tricks the year after. Transformers inherit this exactly: GPT-3's 96 layers and Llama 70B's 80 both train because every one of them is a residual block.
Where the norm goes
Every modern block normalizes before the sublayer, not after. The difference sounds cosmetic. It decides whether depth trains at all.
The original 2017 paper normalized after the residual sum — post-norm. Almost every model since 2020 normalizes before the sublayer instead — pre-norm. Same two ingredients, reordered — and the order changes what the residual connection protects. Walk the same token through both layouts side by side:
Watch where the last step lands: pre-norm's residual sum is the block's raw output, never touched by a norm; post-norm's sum gets normalized right there, one step later. The skip connection is identical in both — what differs is whether anything sits between it and the exit.
That one step matters because backprop runs the block in reverse: model each norm's backward pass as keeping a fraction c of the gradient that passes through it, and compare pre-norm against post-norm as depth grows:
Because post-norm's +1 sits inside the norm, it gets multiplied by c at every single layer — turn c down and the post-norm curve, residual connection and all, folds back into the same vanishing shape a plain chain has all along. Pre-norm's +1 never sits inside anything; it stays exactly one, no matter how deep the stack or how small c gets.
Pre-norm has one bill still to pay: nothing inside the block ever rescales the residual stream itself. Watch its own magnitude grow, block by block, with no correction in sight:
Because the norm only reads a copy of x on its way into a sublayer, it never touches x on the through-path — why every pre-norm model needs one last norm after the final block, before anything downstream reads the stream directly.
That is why a block carries both: normalization controls the scale of what flows through it, and the residual connection controls whether that flow reaches layer 1 at all. Put the norm on the wrong side of the sum and it quietly compromises the second job while doing the first — and nobody notices until the model is deep enough for it to matter.
The full block, once
Put the pieces in order — norm, attention, add, norm, FFN, add — and the result is the pre-norm block every modern decoder actually runs.
Two sublayers, two residual connections, two norms, one pass each, always in this order — the composition the last three sections built toward, now applied twice inside one block.
Step through the six stages on one token's vector and watch x get read, updated, and read again:
Notice x is reused, not replaced: the same variable gets normalized, transformed, and added back twice in a row. From the residual connection's point of view, attention and the FFN each contribute one small correction to a running total — neither produces the block's output from scratch.
Those two corrections cost very different amounts. Drag through six real widths and compare what attention spends against the FFN, one block at a time:
Notice the FFN bar is always exactly twice the attention bar, at every width: fourd_model×d_model matrices against two d_model×4d_model ones is 4d² against 8d², and the ratio cancels d entirely. The FFN is not incidentally bigger — it is bigger by an exact, width-independent factor.
Zoom out to a whole model and the same pattern repeats at a different scale. Compare the embedding table against every block combined:
In a Transformer, attention feels like “the model,” but it is the FFN that holds most of a block's parameters and does most of the computation: attention decides which information to route, the FFN decides what to do with it — including, researchers now think, a model's factual recall.
Depth is where composition happens
Stack N of these blocks and no single one changes — only how many times the transformation composes with itself.
Every block has its own weights — its own attention patterns, its own FFN, its own norm parameters — but every block reads and writes R^d_model, which is why §01's contract mattered. Build a stack one block at a time, up to a real model's depth:
Notice nothing changes shape as N grows — only silhouette count: , . Depth and width both grow a model — slide one factor and compare N against d_model:
Watch the gap widen as the line and the curve pull apart: quadruple the depth and parameters go up exactly 4×; quadruple the width and they go up exactly 16×, because every matrix in a block is d × d or d × 4d — width costs quadratically, depth costs linearly.
Which means the same parameter budget buys very different things depending on where you spend it. Drag between a deep, narrow stack and a shallow, wide one, holding the total roughly fixed:
Watch how many more blocks a narrow stack affords for the same budget — width's quadratic cost is exactly what makes depth the cheap way to buy more sequential composition, and width the expensive way to buy more capacity at one depth.
“More composition” is not just a metaphor. Unravelled, a stack of N residual blocks is a sum over every subset the signal could have taken — drag N and watch how fast the path count grows:
By — that is already 16.8 million distinct paths, all inside one stack of blocks with completely ordinary weights (Veit, Wilber & Belongie, 2016). Width never produces this: doubling a layer's neurons doubles its capacity at one depth. Doubling the number of layers multiplies how those layers can compose.
Six lines, and the model they become
Every claim on this page reduces to two lines, repeated N times, wrapped in four more.
Build the wrapper one stage at a time, from the block itself to a complete model:
That is six lines total: the two-line block from §05, repeated N times inside the stack just built —
x = x + attn(norm(x)) x = x + ffn(norm(x))
wrapped in four more: an embedding on one end, norm-then-head on the other — a complete decoder-only language model:
x = embed(ids) for blk in blocks: x = blk(x) x = final_norm(x) logits = head(x)
The head is a single d_model × vocab_size linear layer, often tied to the embedding table to save parameters twice over.
One piece was left for the self-attention primer: watch which positions the causal mask inside attention lets a given position read:
Notice can read all six once it arrives, but could only ever read itself. At inference the K and V vectors behind every already-visible position are cached rather than recomputed — the reason a longer generated reply gets slower to extend, and the target of MQA, GQA, and sliding-window attention alike.