Positional Encoding Primer
Attention reads a sequence as a set: shuffle the tokens and every output comes back unchanged. Three things, proved with figures you can drive — a good encoding has to be bounded; a schedule of sines makes the dot product a function of the gap alone; and rotating the query and the key turns that from a hint into an invariant of the score itself.
Attention cannot see where anything is
Shuffle the tokens and every output comes back identical, just rearranged. The repair is one addition — and the addition is not free.
Look at what is actually in the score. Q[i]·K[j] reads the contents of two tokens and nothing else; there is no index anywhere in softmax(QKᵀ/√d)·V. Attention is permutation-equivariant: permute the input rows and the output rows permute with them.
Here are six orderings of three tokens, each with its attention output drawn underneath it. Drag the order and watch the outputs ride along the wires with the words they belong to:
Notice the readout never leaves 0.00. Dog bites man and man bites dog hand attention the same three vectors in a different arrangement — for language that is the wrong answer, because nothing downstream can recover which word was the agent.
So we put position into the vector before attention ever sees it. Every token's embedding gets a vector that depends only on where it sits added to it — slide the position and watch the middle row move while the top row holds still:
The bottom row is all attention ever reads. Watch ‖PE(p)‖ hold at 4.00 wherever the slider goes: each pair of slots is a point on a unit circle, so sixteen pairs always give length 4. Bounded at every position is the first thing an encoding has to get right.
But one vector is now carrying two facts, and they interfere. Turn the position weight up and compare two similarities — the same word five apart against two different words in the same slot:
The curves cross at . Past that, two different words in the same slot look more alike than the same word five apart, and every layer above has to unpick a sum that has already lost the distinction. The 2017 paper adds at weight 1, comfortably to the left of the crossing.
So the list is: bounded, distinct at every position, quiet enough that content survives the sum, and ideally saying something about distance. The next section tries the two schemes everyone thinks of first.
Counting is harder than it looks
Two obvious encodings and the specific way each one breaks. What survives is the scheme that counts at several scales at once.
The simplest encoding is the index itself: position 0 gets 0, position 1 gets 1, for ever. Nothing to train, nothing to store, no maximum length — and the one that dies fastest.
The vector it is added to has a length of about 1. Walk the marker up the ramp and watch how fast it clears the band the embedding lives in:
By the encoding is sixty-three times the size of everything it is added to. A layer reading that sum sees almost nothing but position, and the gradient reaching the token embedding is swamped. Bounded values are not a nicety.
The obvious repair is to divide by the sequence length, which does bound it. Stretch the second sequence and watch what slot 3 is worth in each:
Notice that slot 3 means 0.43 in an eight-token sequence and 0.05 in a sixty-four-token one. The encoding has stopped being a property of the position and become a property of the batch — “three tokens back” now has no fixed representation for the model to learn.
What we want is bounded, absolute, and still informative about distance. Binary already does all three: drag the position and watch each bit flip at its own rate:
Bit 0 flips every 2 tokens and bit 3 every 16. Read the column top to bottom: the fast bits separate neighbours, the slow ones separate whole regions, and every value stays in {0, 1} however long the sequence gets.
The one thing wrong with the odometer is that it is made of steps: 01111 to 10000 flips five bits at once, with no smooth direction for a gradient to follow. Round the corners off and you have the sinusoidal encoding.
A smooth odometer
Sines and cosines on a geometric ladder of frequencies. No parameters, and a wavelength for every scale.
The 2017 formula gives slot 2i of position p the value sin(p·θᵢ) and slot 2i+1 the value cos(p·θᵢ), with θᵢ = base^(−2i/d). Every pair of slots gets its own frequency, falling geometrically.
That is the odometer with the corners rounded off. Pick a pair and read its frequency off the readout — the top rows turn over in a handful of tokens, the bottom ones barely bend:
Pair 0 turns 1.000 radian per token, so it comes back round every 6.3. takes 353. The ladder is geometric on purpose: each rung costs the same two slots and buys a scale an order of magnitude coarser than the last.
The encoding of one position is a vertical cut through all of them. Slide the cut and watch the eight dots find their heights:
Those eight sine values, with their eight cosine partners, are the sixteen numbers of PE(p). No two positions share a column, and neighbouring columns differ only in the fast rows — which is exactly the distance-awareness the bits had, now continuous.
How far the ladder reaches is set by the base. Pick a pair, then change the base underneath it:
At the paper's base of 10,000 this ladder runs from 6.3 tokens at pair 0 to 35,333 at . Five of the sixteen rungs sit above the trained window and never complete a turn inside it. Raise the base and the whole ladder stretches, which is the lever Llama 3 pulls.
So the encoding is bounded, distinct at every position, multi-scale and free. What is left is the property everything after this turns on: whether it says anything about the distance between two positions.
The gap, not the place
Each pair of slots is a point on a circle, so moving forward is turning. That is what makes a sine schedule more than a hash.
Take one pair on its own. (sin p·θᵢ, cos p·θᵢ) is a point on the unit circle at angle p·θᵢ, and going from p to p+k is a rotation by k·θᵢ — a turn that does not depend on p at all.
Move the position and watch its arc grow while the step forward keeps exactly its size:
Because the step is the same turn wherever the position ends, PE(p+k) is a linear function of PE(p) whose matrix depends only on k. One weight matrix could implement “look three tokens back” for every position at once.
Do that in all sixteen pairs at once and something collapses. Take two cuts through the wave stack and read the dot product of the two columns they pick out:
Slide both cuts together and the number holds. At a gap of eight it reads 0.66 wherever the pair of cuts sits, and it peaks at 1.00 only when they coincide.
The same fact drawn as a curve. Fix the first position and score it against every other one — the whole curve travels with it, shape intact:
Because sin a sin b + cos a cos b is cos(a−b) exactly, PE(p)·PE(q) is Σᵢ cos((p−q)·θᵢ) — a function of the gap and nothing else. The value eight tokens along stays at 0.66 whether p is 0 or .
Here is the part summaries skip. Attention does not score PE against PE; it scores (x+PE)W_Q against (x+PE)W_K. Push the projections off the identity:
The eight curves sit exactly on the exact one at rest and fan out as soon as the projections move: by mix 1.00 the spread at a gap of eight is 0.48, nearly a quarter of the whole range. Translation invariance was a basis the model could use, never a guarantee it gets.
Worse, expanding (x+PE)W_Q · (x+PE)W_K gives four terms and only one is position against position. The gap between what the encoding promises and what the score delivers is the opening RoPE walks through.
A row for every position
BERT and GPT-2 skipped the arithmetic: allocate one row of parameters per position and let the optimiser decide what goes in it.
nn.Embedding(max_len, d_model), looked up by index and added exactly like the sinusoidal vector. Same code path as the token table, different vocabulary — anyone who has written a token embedding has already written this one.
Nothing constrains what a row holds. Pick a row and slide the smoothness — both ends are states gradient descent can leave behind:
At the rows are independent draws, which is exactly what the parameterisation guarantees: nothing. A trained table usually ends up smoother than that, but the model never asked for it — the structure is an artefact of the data, not a property of the scheme.
That decides what neighbouring rows look like. Compare the learned similarity with the exact sinusoidal one at the same gaps:
Because the slow pairs have barely moved, the sinusoidal curve leaves a one-token gap at 0.96 and falls smoothly. The learned curve starts wherever training left it — 0.60 here at smoothness 0.60 — and every position pair the data never exercised is a coin flip.
The fatal problem is simpler than any of that. Push the sequence past the end of the table:
There is no row 1024. GPT-2 small allocates 1,024 × 768 = 786,432 parameters for position and stops; ask for and the lookup raises an IndexError. It is the rare failure that fails loudly.
That is the trade. A learned table fits its own data perfectly up to max_len and says nothing past it; a formula says something everywhere and fits nothing perfectly. RoPE keeps the formula and changes where it is applied.
Rotate instead of add
Leave the embedding alone. Turn the query and the key by their own positions, inside the score, and the gap is all that survives.
RoPE (Su et al., 2021) never touches the token vector. Inside each head it treats every pair of dimensions of Q and K as a 2-D vector and turns it by p·θᵢ, on the schedule §03 built.
A dot product of two vectors on a circle depends only on the angle between them. Move the query's position and the key's and watch the score:
Move both by the same amount and the score does not budge; move one and it changes with the gap alone. That is the invariant: ⟨R_m q, R_n k⟩ = ⟨q, R_(n−m) k⟩ holds for every m, and it is a property of the score itself rather than a basis the model has to find.
Every pair turns at its own rate, so one position is a reading across the whole ladder at once. Slide the position and watch the dials wind:
At pair 0 has turned 32.0 radians — five full turns — while pair 7 has managed 0.569. The fast dials resolve neighbours, the slow ones keep tokens hundreds apart distinguishable, and the head reads all eight at once.
Summed over the ladder, that gives a score that falls off with distance. Probe a gap and change the base under it:
For a query and key carrying the same content this is the same sum of cosines as §04: RoPE's decay and the sinusoidal dot product are one object. A gap of 64 scores 0.54 at base 10,000 and 0.67 at , which is the lever Llama 3 pulled.
And it is nearly free: with the sines precomputed a Llama-3-8B layer spends 3 × (4,096 + 1,024) = 15,360 flops against 83.9 million for its four projections — 0.018%, no parameters, and V untouched.
Longer than it was trained on
No table means no hard limit — not that the model has seen the angles you are about to hand it.
A model trained to 2,048 tokens has only scored pairs whose per-pair angle lay inside 2048·θᵢ. Nothing errors at position 8,192 — every pair simply arrives four times further round than anything in training.
That is the whole long-context problem, drawn. Drag the scale until the angles being asked for land back on the ones the model was trained on:
At pair 4 reads 204.8 radians instead of 819.2, exactly the ceiling training reached. That is Position Interpolation, and it bought LLaMA-7B 2,048 → 32,768 tokens with 1,000 fine-tuning steps. The bill is resolution: adjacent tokens now differ by a quarter of the angle they used to.
The mechanism itself is two lines inside a head, and it never touches V:
q, k, v = x @ Wq, x @ Wk, x @ Wv q, k = rope(q, pos), rope(k, pos) # v is not a = softmax(q @ k.mT / d_head**0.5) @ v
The mistake to know about. Two pairings are in the wild: the paper pairs slot 2i with 2i+1, every HuggingFace LlamaAttention pairs i with i + d/2. They are a permutation of each other, so each is correct alone.
# the paper pairs (2i, 2i+1) q = interleave_rotate(q, cos, sin) # HuggingFace pairs (i, i + d/2) q = q * cos + rotate_half(q) * sin
Trace one slot of the head through both. The paper's bracket and HuggingFace's agree on slot 0 and then stop agreeing:
By one convention hands it θ2 and the other θ5 — a frequency thirty times slower. Run a checkpoint under the wrong one and nothing raises: the loss stays finite, the text stays grammatical, and quality quietly drops.
What none of this costs is parameters. Slide the context a learned table would have to cover and compare it with the other two:
GPT-2's table is 786,432 parameters for 1,024 positions; at it would be 25,165,824, still useless at position 32,769. Sinusoidal and RoPE stay on the axis at zero for the whole sweep.
So: sinusoidal fails silently past its training length, a learned table fails loudly at max_len, and RoPE fails silently too — but it is the only one whose failure has a knob on it.