Transformer Forward Pass

Twenty-one primers built self-attention, multi-head attention, positional encoding, the block, and the three architecture shapes as separate mechanisms. This one runs them together — one sentence, six stages, end to end — then keeps going past where every one of them stops: sampling a token, feeding it back in, the cache that makes that affordable, and what the whole machine actually costs to run.

01

The shape of the stack

Twenty-one primers built the pieces. Here they run as one machine — six stages, one sentence, start to finish.

The stages are always the same: tokenize, embed and add position, run N copies of one block, normalize, unembed, read off logits. Attention lives inside the block; everything before and after it is bookkeeping — and the bookkeeping is where the shapes actually live, so that's where we start.

A token isn't meaning yet — it's an id, and the id is a row number in a table every distinct word owns one line of. Drag through the sentence and watch the lookup land:

position 0 · "the" · row 0 of the table

Notice position 0 and both read "the," and both land on the identical row — same id, same vector, no notion of where in the sentence either one sits. That's deliberate: order isn't in the picture yet. It arrives next, and it's the whole subject of positional encoding.

Zoom out and the six stages are one drawing: the active stage lit, the stages behind it already run. Step through the whole pass:

stage 1 of 6 — tokenize

Watch the block band — it's drawn taller on purpose. It isn't one stage, it's N stacked copies of an identical block, and N is a knob we turn next. Everything self-attention, multi-head attention and positional encoding built lives entirely inside that one wider band.

Turn that knob and watch the parameter count — width held at GPT-2 small's own 768, only depth moving:

N = 12 · 124M params

GPT-2 small's own 12 layers land at 124M parameters — read it straight off the curve, and it's the number the paper reports. Depth alone got us there: no vocabulary changed, no width changed, just twelve copies of one 7-million-parameter block on top of a shared embedding table. Push to and the count roughly triples — width held fixed, the curve is still a straight line.

Every one of those N blocks is handed the residual stream and has exactly two moves available: add to it, or replace it. Toggle between them:

block 1 · add · stream width 4

Watch what "replace" does: the stream's own width never changes — it's still four slots wide — but everything the earlier blocks wrote is simply gone. That's the invariant worth pasting as an assertion: at every layer boundary the stream keeps its own shape, and every sub-layer only ever adds to it. §02 opens the block up and shows exactly what gets added.

02

One block, twice

Every one of those N copies runs the same four moves. Once through them is the whole section — the mechanisms themselves already have their own primers.

A block reads the stream, normalizes a copy of it, runs attention on that copy, and adds the result back. Then it does it again with the FFN. Two sub-layers, each wrapped in the same add-not-replace shape §01 just proved.

The active move lights up as you step; the straight line down the middle is the stream itself, and it never bends — every branch peels off it and rejoins it:

1 of 4 — norm

Notice both branches end the same way: attention adds, then the FFN adds. That's the invariant from §01, twice per block — nothing here replaces the stream, ever.

The two sub-layers don't do the same job. Pick a token and watch what each one is allowed to read: attention reads every other position; the FFN reads only the token it's sitting on:

token 0 · attention reads all 4 · FFN reads only itself

As soon as you move off token 0, attention's lines fan out to all four cells while the FFN's never leaves its own column. Say it plainly: attention moves information between tokens; the FFN moves it between channels, one token at a time.

That second move is expensive on purpose. Of one block's 12d² parameters, drag the width and watch how they split:

attention 2.4M (33%) · FFN 4.7M (67%)

The FFN holds two-thirds of every block's weights — the fact each block's hidden layer is 4× wider than d_model, paid for twice, once up and once down. Try : the counts shrink, but the split itself never moves. It's a ratio, not a count.

One more thing a block chooses: where the norm sits. Toggle it, and watch the gradient's own path back to the embedding:

N = 12 · pre-norm · gradient crosses 0 norms

Under post-norm — the original 2017 design — the residual path itself runs through a normalization at every single block, N of them end to end. Under pre-norm, the norm only touches the branch, never the stream, so the gradient crosses zero. That single choice is most of why GPT-2, Llama and every modern stack train reliably past a hundred layers. §03 leaves the stack behind and asks what comes out the far end.

03

From logits to a token

The stack ends in logits, one score per vocabulary entry. Turning that into a single chosen word is its own small machine — the first thing on this page that nothing earlier in the stack builds.

Take one point in the story: the prompt is "the cat sat on the ___", and the model has scored five candidates — mat, rug, floor, sofa, roof. Step through each one's raw score, before anything reshapes it:

"mat" · raw score 4.0

Notice "mat" already leads on the raw number alone, 4.0 against the field. Softmax turns that list into probabilities that sum to one; what changes the shape of the distribution — not which one is winning — is temperature. The winning candidate is redrawn every time you move the slider. At the resting T = 1, "mat" already owns 68.7%:

T = 1.00 · "mat" 68.7%. drag left and right; the arrow keys step it and Home restores the opening state
T = 1.00 · "mat" 68.7%

Watch the bars flatten as T climbs past 1, and watch "mat" swallow almost everything as T drops toward 0. At it already owns essentially all of it — this is what argmax means in the limit: always the single highest logit, every time.

Argmax is one way to cut the tail off; keeping exactly the top k candidates is another. Drag k down and watch the discarded mass grow:

k = 5 · 0.0% of the mass cut

At this is argmax — one candidate, zero alternatives. Notice k never adapts to the shape of the distribution: it cuts to a fixed count whether the model is confident or not.

Top-p is the adaptive version — keep the smallest set whose probability reaches p, so a peaked distribution keeps few candidates and a flat one keeps many:

p = 1.00 · 5 candidates kept · 0.0% cut

At exactly two candidates survive — mat and rug clear 90% between them, the rest don't. Push p down and the set can shrink to one; push it toward 1 and every candidate with any probability at all stays in play.

Greedy decoding — argmax, every step, forever — has a real failure mode: drive it and watch it happen:

step 0 · emits "the"

Once the toy model's argmax path loops back on itself, greedy decoding repeats that token forever — it has no mechanism to notice and no way out. A single sampled step at the right moment breaks the cycle; production systems call this exact pathology repetition collapse. §04 takes the token this section just chose and asks what happens the moment it's fed back in.

04

Generating, one token at a time

A model that scores one next word isn't a generator yet. Feed the winner back in as if it had always been part of the prompt, and it is.

Every pass through the stack grows the sequence by exactly one token: run the whole thing forward, sample, append, repeat. There's no separate "generation mode" — it's the same forward pass §01 through §03 already built, called again and again on a longer input each time.

The prompt is just given, grey from the start; each generated token settles in behind it; the one under the slider is whatever the model is deciding right now. Step through it:

step 4 · prompt · "the". drag left and right; the arrow keys step it and Home restores the opening state
step 4 · prompt · "the"

Notice every settled token becomes part of the input for the next one — that's the whole definition of autoregressive: the output at step t is an ingredient of the input at step t + 1.

Here's what that costs if you're not careful. The naive approach reruns every existing token through every layer, every single step; the smarter one only processes what's new. Compare them on the same step:

n = 5 tokens so far · naive · reprocesses 6

Under "naive," the whole prefix relights every time you advance — n tokens of work to produce token n + 1. Under "cached," only the newest cell does.

Run that gap out over an actual generation — four toy layers, d = 64 — and watch the two totals pull apart:

T = 12 · naive 51.05M · cached 4.85M

At the first token the two costs are close — 5× apart. By , the naive total is over ten times the cached one, and the gap is still opening: naive cost grows with the square of how much you've generated, cached cost grows linearly.

Scale that up to GPT-2 small's own 12 layers and 768 dimensions, starting from a 20-token prompt:

1 generated · naive 3.4e+09 FLOPs · cached 1.7e+08 FLOPs · 20.0×

Drag to and the ratio lands past 69×. That's not a rounding error — it's the entire reason production systems never actually rerun the prefix. What they store instead, and exactly how much of it, is §05.

05

What the cache remembers

What "cached" stores, exactly, is every key and value vector attention has already computed — one pair per token, per head, per layer, never touched again.

Attention scores a new token's query against every key that came before it. Those old keys and values don't change once written — the block that produced them never runs on them again — so a cache is just somewhere honest to keep them.

Every generated token appends exactly one pair. Watch the filled slots accumulate:

1 pairs cached · newest "the"

That's the operational invariant this whole section rests on: after emitting token t the cache holds exactly t pairs, and producing token t + 1 appends one more — it never recomputes, never rewrites, the ones already there.

Scale that bookkeeping to GPT-3's own published shape — 96 layers, d_model = 12,288 — and it stops being free. Drag the context out:

context 2,048 tokens · 9 GiB

At GPT-3's own 2,048-token context the cache for one sequence alone reaches 9 GiB, in fp16 — before a single weight is loaded. Pull back to and it's exactly a quarter of that: the formula is linear, two bytes per number, one key vector and one value vector, one pair per layer, times however many tokens are in the window.

A cache is usually a fixed-size buffer, not an unbounded list. Push generation past its own capacity and watch what a naive ring buffer does:

position 0 → slot 0 · free. drag left and right; the arrow keys step it and Home restores the opening state
position 0 → slot 0 · free

Past slot 7 the write wraps around and lands on a slot a still-valid earlier token needs. If nothing guards the buffer's edge, this fails silently: no error, no crash, just an older token's key quietly replaced by a newer one's, and every score computed against it from then on is wrong.

A second thing has to travel with every cached key: which position it sits at. Toggle a real implementation bug and watch it drift:

wrong — always 0 · position written = 0 · real position = 0

Get this wrong — write position 0 for every new token instead of the cache's own current length — and the position encoding this whole page has assumed since §01 stops matching reality: the model scores every new token as if it were the very first one, and the output degrades without a single exception being thrown. §06 puts a number on what all of this actually costs to run.

06

The budget

Two different questions, both worth a real number: what does one pass cost, and where does the memory actually end up.

Every step reads the whole weight matrix from memory once and spends it on however many tokens are in that pass. The ratio — FLOPs earned per byte read — is what decides whether a step is limited by the chip's arithmetic or by how fast memory can feed it.

Prefill scores every prompt token in one pass, so one weight-read serves all of them; decode serves exactly one. Watch the two bars pull apart as the prompt grows:

P = 2 · prefill 2.0 FLOPs/byte · decode 1.0 FLOPs/byte

Prefill's intensity climbs with the prompt — push to and it earns 16 FLOPs a byte, straightforwardly compute-bound. Decode's stays pinned at exactly one, whatever the model's size — that's memory-bandwidth-bound, and no amount of extra compute fixes it.

Now the other question. At GPT-3's own scale, weigh the fixed weights against one sequence's own cache:

context 2,048 · weights 325 GiB · cache 9 GiB

Even at the full 2,048-token context, one sequence's cache is a sliver next to 325 GiB of weights. One request is cheap. The budget changes the moment there's more than one.

The weights are paid for once and shared by every request in flight; the cache is not — it's per sequence. Drag the number of concurrent sequences:

1 sequence · weights 325 GiB · cache 9 GiB. drag left and right; the arrow keys step it and Home restores the opening state
1 sequence · weights 325 GiB · cache 9 GiB

Past held open at once, their combined cache costs more memory than the weights that serve every one of them — the exact pressure that made KV-cache memory management, not raw FLOPs, the thing production serving systems are mostly built around.

One last number, closing the loop back to §01 and §02: of GPT-2 small's own 124M parameters:

roughly two-thirds sit in the twelve blocks — mostly the FFN, per §02 — and the rest is the embedding table this whole page started by looking up. §07 collects the loop into runnable code, the costs into one table, and points at what to read next.

07

Reference

The whole loop in nine lines, the costs in one table, and the six stages once more with the arrow that makes them a generator.

# grow the sequence one token at a time
cache = None
while len(tokens) < max_len:
    x = tokens[-1:] if cache else tokens
    logits, cache = model(x, cache)
    probs = softmax(logits[-1] / temperature)
    next_id = sample(probs, top_k, top_p)
    tokens.append(next_id)
    if next_id == EOS_ID: break

Every line of that loop is something a section on this page already put a number on — step through it and see which:

model(x, cache) — §01–02

What all four of those lines cost, together:

Time
cached: O(n) per new token · uncached: O(n²)
n is however many tokens exist so far. Attention itself stays O(n) either way — a cache removes the wasted O(n·d²) of rerunning old tokens through every projection and the FFN, not the score lookup against them.
Space
O(layers · d_model · n) for the cache, 2 bytes a number
Independent of vocabulary size — it grows with context length and model width alone, stacked on top of a fixed O(params) for the weights that never changes with n.

Why this bound is tight

The ceiling is rarely raw FLOPs. Prefill is compute-bound; decode is pinned near one FLOP per byte read, whatever the model's size, which makes it memory-bandwidth-bound by construction — and past a few dozen concurrent sequences the KV cache outweighs the weights serving every one of them.

Variants

prefillO(P) tokens, one passO(P·d·layers) cache built

Toggle the loop back on:

logits feeds back into tokenize — this is what makes it a generator