Text in LLMs Primer
A model is a function on a rectangle of floats. A string is not one. This primer is the trip between them, and the three properties text has that a rectangle does not: no fixed length, an order that carries the meaning, and tokens whose meaning is decided by their neighbours. Every number here is computed by the figure beside it.
A model never sees text
It sees a rectangle of numbers. Everything interesting happens on the way there.
Every operation in a Transformer — a matrix multiply, a softmax, a residual add — is arithmetic on floats in a fixed rectangle. A string is none of that. Five conversions sit between the two, and it is worth knowing which of them are reversible and which quietly throw something away.
So here is one short string, café ☕, carried the whole way, one stage per press of the scrubber, with the count for the stage you are on beside it:
Notice that the count changes at almost every stage and never means the same thing twice. Six characters become six code points, then nine bytes, then three tokens, then three ids, then three rows of floats. Only the last of those is what the model consumes, and by then nothing of the original string survives except the ids.
Start underneath. UTF-8 spends one to four bytes on a code point, and the drawing makes that literal: each character box is exactly as wide as the bytes beneath it. Drag the slider through six scripts:
Watch the boxes stop being equal. English costs 1.00 bytes a character, , , an emoji 4.00. Four to one, and nothing announces which you are paying. ASCII stayed one byte deliberately, so old text is valid UTF-8 untouched.
“Character” is doing a lot of work in that paragraph. What a reader calls one character can be several code points glued together, and the number your language reports is none of the three. Slide through six of them:
Because len() counts code points, the — one glyph, 18 bytes — reports 5. Slice a string at an arbitrary index and you cut through the middle of a character: s[:1] on that emoji yields a lone man, and slicing the bytes instead yields no character at all.
Worse, one glyph can have two spellings that are both correct. NFC precomposes the accent into the letter; NFD keeps the letter and hangs a combining mark off it. Flip the form and watch the rendering hold still while the bytes move:
Both forms are canonically equivalent, both are valid, and nothing raises. Composed, the word is 5 bytes; decomposed, 6, and they first differ at byte 3 — so "café" == "café" is False, and every hash, index and exact-match filter downstream now holds two byte strings where a reader sees one word.
That failure never announces itself. It is a lookup returning nothing while the key sits in plain sight on the page. Push normalisation up the write path and watch the bucket nobody can see drain:
Once every key is normalised on the way in there is one bucket, not two, and the query finds all eight. Until then the store holds duplicates no reader can tell apart, and the only symptom is a recall number a little lower than it should be. Normalise at the boundary: after tokenizing is too late, because by then the two forms are already different ids.
Length is not a shape
“OK.” is two tokens and a Wikipedia article is a million. The same matrix multiply has to take both.
A GPU wants a rectangle: B rows, T columns, every row the same length. Text has no natural maximum and no natural minimum, so something has to make it rectangular — and each way of doing that bills you differently.
The usual answer is to pad: take the longest document in the batch and fill every other row out to it with a padding cell. Drag the twelfth document longer:
Watch the whole rectangle grow because of one row. At its shortest the batch wastes 32% of its cells; drag the outlier out to and 69% of the tensor holds nothing. Every one of those cells is multiplied, softmaxed and added exactly like a real token.
Wasted arithmetic is survivable as long as the model knows to ignore the result. That is the mask's entire job, and the classic bug is a pooling step that never asks it. Shorten the document and switch what the mean divides by:
The unmasked mean is not noisy, it is wrong, and wrong in one direction: at five real tokens in a twelve-wide row it reads 0.300 against a true 0.720. Nothing raises. The model trains, the loss falls, and a metric sits a few points low for reasons nobody can find.
The waste is not a law of nature either — it is a consequence of who shares a rectangle with whom. Sort the corpus by length and cut it into buckets, each padded only to its own longest:
One bucket is the naive batch: 220 cells for 99 real tokens, 55% padding. bring that to 110 cells and 10%. The catch is at the far end — one bucket per document is zero waste and a batch of one, which is the parallelism you bought the rectangle for.
The other answer is to refuse the tail outright: pick a maximum length and drop everything past it. Drag the cut through a four-thousand-token document:
At a the model reads 512 of 4,096 and never learns the other 88% was there. It does not error; it answers. That is the shape of every failure in this section: the tensor is well-formed, the arithmetic runs, and the only thing wrong is what the numbers mean.
The bag that cannot count to two
Six orderings of three words, one vector. Anything reading only the vector cannot tell them apart.
The cheapest way to make text a fixed shape is to stop caring where the words are: count how often each vocabulary entry appears and hand over the counts. One pass, any length, and still a decent baseline for spam and topic classification.
It also throws away the thing that makes language language. Walk the slider through every ordering of a three-word sentence and watch the count vector underneath refuse to move:
All six orderings, one vector — and is not “dog bit cat”. This is not an approximation that improves with data: the map from sequence to bag is not injective, so no function of the bag alone can separate them, at any scale, ever.
The classic patch is to count short runs instead of single words. Widen the window and watch what the model gets to count:
At the model can finally tell “dog bit” from “bit dog”, and the tuples below carry that. But read the right-hand number: the feature space is the vocabulary raised to the power n, so one more word of window multiplies the possibilities by 50,257.
That trade is worth seeing on an axis. The vertical scale is logarithmic — each gridline is a hundred thousand times the one below it — so an exponential blow-up draws as a straight line:
Because the axis is logarithmic, that straight line is the explosion: at there are 1.27×10¹⁴ possible trigrams, more than any corpus could ever populate. Nearly every feature is zero, and most of the rest were seen once.
And it still does not reach far. Drag then away from if, and watch the windows that hold both go out one by one:
As soon as the distance reaches the window width, no window contains both words — so no feature anywhere in the model mentions the pair. Counting harder cannot help, because the pair is not in the feature space at all. That is the wall, and it is why the rest of this primer is about distance rather than about counting.
A token has no meaning on its own
The bank by the river and the bank that sets rates are the same id and the same row of the table.
Before any of the arithmetic, each id has to become a vector, and the cheapest way is a lookup table: id 31 means row 31, always, everywhere, whatever the sentence around it says.
Both occurrences of bank below are the same token, so both read the same row. Move the slider between them and watch where the arrow lands:
Only the arrow moves. Word2Vec and GloVe are exactly this table, and it was an enormous advance over one-hot — but the vector is chosen before the sentence is read, so whatever separates the two banks has to be recovered downstream by something else.
Recovering it means mixing the neighbours in. Raise α and watch one row become two:
At α = 0 the rows are identical and the readout says they are 0.00 apart. At they are 1.12 apart and the model can answer “which bank”. That mix is what attention computes; the weights are learned, per token, per layer.
A recurrent net does the mixing one step at a time, which puts a hard ceiling on how far back it can carry anything. Set the gate and drag the distance:
With a gate keeping 90% per step, a token is under 1% of the state. Raising γ buys reach and risks a state that saturates; lowering it to 0.80 puts twenty tokens at just over 1%. No setting both remembers a paragraph and stays stable — which is the vanishing gradient, seen forwards.
How far apart two tokens are
One number scores every architecture here: the shortest path a signal takes from one position to another.
Short paths learn long-range structure. Long paths lose it, because everything the signal passes through on the way is free to overwrite it.
Pick an architecture and drag the two apart. Every tread of the staircase is one hop a signal has to make:
A bag has no path at all. A 3-gram has one hop out to distance 2 and none beyond it. A convolution with kernel 3 needs ⌈d/2⌉ layers. A needs twelve sequential steps — the decay figure again, one multiplication per tread. Attention needs one hop at every distance.
One hop everywhere is not free. To have a path from every token to every other, you have to score every pair. Drag n and watch the square rather than the row:
Double the sequence and the row doubles while the square quadruples: at that is 144 scores for 12 tokens, every one of them computed, softmaxed and multiplied through. The path length is constant; what pays for it is the area.
What the rectangle costs
Attention is the cheaper layer until the sequence is longer than the model is wide. That crossover is not a mystery.
A self-attention layer of width d does 2n²d multiply-accumulates; a recurrent layer of the same width does 2nd². Divide one by the other and the ratio is n/d, so they cost the same at exactly n = d.
Both axes below are logarithmic, so a power law draws as a straight line and the crossover is simply where the two lines meet. Drag the sequence length through it:
At GPT-2's width of the attention and recurrent curves cross at 9.06×10⁸ operations each. Below it attention is both cheaper and shallower. Above it, it is buying its one-hop path with a line twice as steep.
Compute is the half people quote. The other half is that the scores have to exist somewhere while the softmax runs. Drag the context window and read the memory:
At a twelve-layer, twelve-head model is holding 6.04×10⁸ scores — 1.13 GiB at fp16, for one sequence, before a single weight. That number is why FlashAttention exists: it computes the same softmax in tiles and never writes the matrix down. The quadratic is a memory problem long before it is a compute problem.
Four that fail quietly
A shape error raises. None of these does.
Normalisation. Two byte strings, one glyph; the index holds both and recall drops by a fraction nobody can attribute. The missing mask. A mean over the padded width is finite, plausible, and biased toward zero in proportion to the padding. Truncation. The model answers about the first L tokens and says nothing about the rest.
The fourth is slicing before tokenizing, and it has a mechanism you can feel. Take the sentence §01 started with, cut its bytes at an arbitrary offset, and watch what decoding the prefix gives back:
Notice the rule the format is built on: a byte starts a character exactly when its top two bits are not 10, so a decoder always knows whether it is standing on a boundary. Three of these nine offsets land — and on pure ASCII none of them do, which is why a chunker written and tested in English ships, and then quietly returns a different token sequence for the same document the moment it meets an accent.