Text in LLMs Primer

A model is a function on a rectangle of floats. A string is not one. This primer is the trip between them, and the three properties text has that a rectangle does not: no fixed length, an order that carries the meaning, and tokens whose meaning is decided by their neighbours. Every number here is computed by the figure beside it.

01

A model never sees text

It sees a rectangle of numbers. Everything interesting happens on the way there.

Every operation in a Transformer — a matrix multiply, a softmax, a residual add — is arithmetic on floats in a fixed rectangle. A string is none of that. Five conversions sit between the two, and it is worth knowing which of them are reversible and which quietly throw something away.

So here is one short string, café ☕, carried the whole way, one stage per press of the scrubber, with the count for the stage you are on beside it:

what you typed — 6 characters

Notice that the count changes at almost every stage and never means the same thing twice. Six characters become six code points, then nine bytes, then three tokens, then three ids, then three rows of floats. Only the last of those is what the model consumes, and by then nothing of the original string survives except the ids.

Start underneath. UTF-8 spends one to four bytes on a code point, and the drawing makes that literal: each character box is exactly as wide as the bytes beneath it. Drag the slider through six scripts:

English — 5 characters · 5 bytes

Watch the boxes stop being equal. English costs 1.00 bytes a character, , , an emoji 4.00. Four to one, and nothing announces which you are paying. ASCII stayed one byte deliberately, so old text is valid UTF-8 untouched.

“Character” is doing a lot of work in that paragraph. What a reader calls one character can be several code points glued together, and the number your language reports is none of the three. Slide through six of them:

1 glyph · 1 code point · 1 byte

Because len() counts code points, the — one glyph, 18 bytes — reports 5. Slice a string at an arbitrary index and you cut through the middle of a character: s[:1] on that emoji yields a lone man, and slicing the bytes instead yields no character at all.

Worse, one glyph can have two spellings that are both correct. NFC precomposes the accent into the letter; NFD keeps the letter and hangs a combining mark off it. Flip the form and watch the rendering hold still while the bytes move:

NFC · 5 bytes

Both forms are canonically equivalent, both are valid, and nothing raises. Composed, the word is 5 bytes; decomposed, 6, and they first differ at byte 3 — so "café" == "café" is False, and every hash, index and exact-match filter downstream now holds two byte strings where a reader sees one word.

That failure never announces itself. It is a lookup returning nothing while the key sits in plain sight on the page. Push normalisation up the write path and watch the bucket nobody can see drain:

0 of 8 normalised — 2 distinct byte strings

Once every key is normalised on the way in there is one bucket, not two, and the query finds all eight. Until then the store holds duplicates no reader can tell apart, and the only symptom is a recall number a little lower than it should be. Normalise at the boundary: after tokenizing is too late, because by then the two forms are already different ids.

02

Length is not a shape

“OK.” is two tokens and a Wikipedia article is a million. The same matrix multiply has to take both.

A GPU wants a rectangle: B rows, T columns, every row the same length. Text has no natural maximum and no natural minimum, so something has to make it rectangular — and each way of doing that bills you differently.

The usual answer is to pad: take the longest document in the batch and fill every other row out to it with a padding cell. Drag the twelfth document longer:

the twelfth document is 9 tokens. drag left and right to change the longest document; the arrow keys step one token at a time and Home restores the opening state
the twelfth document is 9 tokens — 32% of the tensor is padding

Watch the whole rectangle grow because of one row. At its shortest the batch wastes 32% of its cells; drag the outlier out to and 69% of the tensor holds nothing. Every one of those cells is multiplied, softmaxed and added exactly like a real token.

Wasted arithmetic is survivable as long as the model knows to ignore the result. That is the mask's entire job, and the classic bug is a pooling step that never asks it. Shorten the document and switch what the mean divides by:

mean over every cell = 0.300

The unmasked mean is not noisy, it is wrong, and wrong in one direction: at five real tokens in a twelve-wide row it reads 0.300 against a true 0.720. Nothing raises. The model trains, the loss falls, and a metric sits a few points low for reasons nobody can find.

The waste is not a law of nature either — it is a consequence of who shares a rectangle with whom. Sort the corpus by length and cut it into buckets, each padded only to its own longest:

1 bucket · 55% padding

One bucket is the naive batch: 220 cells for 99 real tokens, 55% padding. bring that to 110 cells and 10%. The catch is at the far end — one bucket per document is zero waste and a batch of one, which is the parallelism you bought the rectangle for.

The other answer is to refuse the tail outright: pick a maximum length and drop everything past it. Drag the cut through a four-thousand-token document:

the cut is at 512 tokens. drag the cut left and right; the arrow keys move it 64 tokens at a time and Home restores the opening state
keeps 512 of 4,096 — drops 88%

At a the model reads 512 of 4,096 and never learns the other 88% was there. It does not error; it answers. That is the shape of every failure in this section: the tensor is well-formed, the arithmetic runs, and the only thing wrong is what the numbers mean.

03

The bag that cannot count to two

Six orderings of three words, one vector. Anything reading only the vector cannot tell them apart.

The cheapest way to make text a fixed shape is to stop caring where the words are: count how often each vocabulary entry appears and hand over the counts. One pass, any length, and still a decent baseline for spam and topic classification.

It also throws away the thing that makes language language. Walk the slider through every ordering of a three-word sentence and watch the count vector underneath refuse to move:

ordering 1 of 6 — the vector is unchanged

All six orderings, one vector — and is not “dog bit cat”. This is not an approximation that improves with data: the map from sequence to bag is not injective, so no function of the bag alone can separate them, at any scale, ever.

The classic patch is to count short runs instead of single words. Widen the window and watch what the model gets to count:

n = 1 · 5 windows

At the model can finally tell “dog bit” from “bit dog”, and the tuples below carry that. But read the right-hand number: the feature space is the vocabulary raised to the power n, so one more word of window multiplies the possibilities by 50,257.

That trade is worth seeing on an axis. The vertical scale is logarithmic — each gridline is a hundred thousand times the one below it — so an exponential blow-up draws as a straight line:

n = 1 · 50,257 features

Because the axis is logarithmic, that straight line is the explosion: at there are 1.27×10¹⁴ possible trigrams, more than any corpus could ever populate. Nearly every feature is zero, and most of the rest were seen once.

And it still does not reach far. Drag then away from if, and watch the windows that hold both go out one by one:

1 words apart. drag left and right to move the two words apart; the arrow keys step one word at a time and Home restores the opening state
distance 1 · window 3 — 2 windows hold both

As soon as the distance reaches the window width, no window contains both words — so no feature anywhere in the model mentions the pair. Counting harder cannot help, because the pair is not in the feature space at all. That is the wall, and it is why the rest of this primer is about distance rather than about counting.

04

A token has no meaning on its own

The bank by the river and the bank that sets rates are the same id and the same row of the table.

Before any of the arithmetic, each id has to become a vector, and the cheapest way is a lookup table: id 31 means row 31, always, everywhere, whatever the sentence around it says.

Both occurrences of bank below are the same token, so both read the same row. Move the slider between them and watch where the arrow lands:

position 1 and position 7 — both read row 31

Only the arrow moves. Word2Vec and GloVe are exactly this table, and it was an enormous advance over one-hot — but the vector is chosen before the sentence is read, so whatever separates the two banks has to be recovered downstream by something else.

Recovering it means mixing the neighbours in. Raise α and watch one row become two:

α = 0.00 — the two rows differ by 0.00

At α = 0 the rows are identical and the readout says they are 0.00 apart. At they are 1.12 apart and the model can answer “which bank”. That mix is what attention computes; the weights are learned, per token, per layer.

A recurrent net does the mixing one step at a time, which puts a hard ceiling on how far back it can carry anything. Set the gate and drag the distance:

5 back · 59.0% left

With a gate keeping 90% per step, a token is under 1% of the state. Raising γ buys reach and risks a state that saturates; lowering it to 0.80 puts twenty tokens at just over 1%. No setting both remembers a paragraph and stays stable — which is the vanishing gradient, seen forwards.

05

How far apart two tokens are

One number scores every architecture here: the shortest path a signal takes from one position to another.

Short paths learn long-range structure. Long paths lose it, because everything the signal passes through on the way is free to overwrite it.

Pick an architecture and drag the two apart. Every tread of the staircase is one hop a signal has to make:

recurrent · distance 12 — 12 hops

A bag has no path at all. A 3-gram has one hop out to distance 2 and none beyond it. A convolution with kernel 3 needs ⌈d/2⌉ layers. A needs twelve sequential steps — the decay figure again, one multiplication per tread. Attention needs one hop at every distance.

One hop everywhere is not free. To have a path from every token to every other, you have to score every pair. Drag n and watch the square rather than the row:

6 tokens. drag left and right to lengthen the sequence; the arrow keys step 64 tokens at a time and Home restores the opening state
n = 6 · 36 pairs

Double the sequence and the row doubles while the square quadruples: at that is 144 scores for 12 tokens, every one of them computed, softmaxed and multiplied through. The path length is constant; what pays for it is the area.

06

What the rectangle costs

Attention is the cheaper layer until the sequence is longer than the model is wide. That crossover is not a mystery.

A self-attention layer of width d does 2n²d multiply-accumulates; a recurrent layer of the same width does 2nd². Divide one by the other and the ratio is n/d, so they cost the same at exactly n = d.

Both axes below are logarithmic, so a power law draws as a straight line and the crossover is simply where the two lines meet. Drag the sequence length through it:

n = 256 · attention 1.01×10⁸

At GPT-2's width of the attention and recurrent curves cross at 9.06×10⁸ operations each. Below it attention is both cheaper and shallower. Above it, it is buying its one-hop path with a line twice as steep.

Compute is the half people quote. The other half is that the scores have to exist somewhere while the softmax runs. Drag the context window and read the memory:

n = 512 · 3.77×10⁷ scores — 72.0 MiB at fp16

At a twelve-layer, twelve-head model is holding 6.04×10⁸ scores — 1.13 GiB at fp16, for one sequence, before a single weight. That number is why FlashAttention exists: it computes the same softmax in tiles and never writes the matrix down. The quadratic is a memory problem long before it is a compute problem.

07

Four that fail quietly

A shape error raises. None of these does.

Normalisation. Two byte strings, one glyph; the index holds both and recall drops by a fraction nobody can attribute. The missing mask. A mean over the padded width is finite, plausible, and biased toward zero in proportion to the padding. Truncation. The model answers about the first L tokens and says nothing about the rest.

The fourth is slicing before tokenizing, and it has a mechanism you can feel. Take the sentence §01 started with, cut its bytes at an arbitrary offset, and watch what decoding the prefix gives back:

the cut is at byte 5. drag the cut left and right; the arrow keys move it one byte at a time and Home restores the opening state
cut at byte 5 · 4 characters back — lands on a character boundary

Notice the rule the format is built on: a byte starts a character exactly when its top two bits are not 10, so a decoder always knows whether it is standing on a boundary. Three of these nine offsets land — and on pure ASCII none of them do, which is why a chunker written and tested in English ships, and then quietly returns a different token sequence for the same document the moment it meets an accent.