Word Embeddings Primer
A word becomes a short row of numbers, and similarity becomes an angle. Every figure on this page reads from one skip-gram model that was really trained — forty-one tokens, eight numbers a word, sixty epochs — so the cosines, the analogy, the collapse when the negatives are switched off and the 28% of the space that the plane cannot draw are all measurements, not illustrations.
A word has to become numbers
A network multiplies and adds. A word is a symbol. The first bridge anyone builds is also the one that cannot carry any meaning across.
Give every word in the vocabulary an index, and write it as a vector that is 1 at that index and 0 everywhere else. That is one-hot encoding, and for a vocabulary of eight words the entire code fits in one picture. Slide along the vocabulary and watch the single 1 move down the diagonal:
Notice that the table is the code: eight words need eight slots, and every row is 87.5% zeros. The 1 carries the word's identity and nothing else — no length, no spelling, no company it keeps. Two rows differ in exactly two positions whatever the two words are.
That last sentence is the whole problem, and it is easier to feel than to read. Pick a first word and a second, and multiply their matching slots together the way a dot product does:
Watch the product row. Unless the two sliders land on the same word, every product is 0 × 1 or 1 × 0 or 0 × 0, so the sum is zero and the cosine is zero. cat and kitten sit at exactly the same similarity as cat and calculator: there is no near-miss in this code, only identical and unrelated.
The second problem is arithmetic. A real tokeniser does not have eight entries; GPT-2 has 50,257 and Llama-3 has 128,256. Below is one row drawn to scale, starting at the eight words we have been using, with the marker above the slot in use — raise the vocabulary and watch the row fill up:
At the slot is a hundredth of a pixel wide — the marker is pointing at it, not drawing it — and the row costs 201 KB of float32 of which four bytes do anything. Nobody materialises that vector: one_hot(i) @ W is row i of W, so every framework gathers the row and skips the multiply.
Which raises the question the rest of the page answers: if the row we look up is the real object, why should it be long and empty? Give each word a short row of numbers it is free to choose, and this is what training produces — drag through the vocabulary and watch the eight slots:
Every slot is in use, and none of them was assigned by hand. Thirty-two bytes per word instead of 201 KB, and — the part that matters — two rows can now be partly alike. The rest of this page is about what that buys, what it costs, and the three places where reading these numbers goes wrong.
Similarity is an angle
Once a word is a direction, “how related are these two?” has an answer with a number in it — and the number is not a distance.
Start with two words that have only two slots each, so the whole space fits on a page. The dot product multiplies matching slots and adds: v · w = v₁w₁ + v₂w₂. Swing v around w and watch the cosine:
Notice where the sign turns over. The cosine is positive while the two lean the same way, exactly zero on the perpendicular, and negative once v crosses it. The cosine is a measure of agreement in direction, and it lives in [−1, 1] no matter how big the vectors get.
The dot product does not. It is ‖v‖ ‖w‖ cos θ, so it answers a question about size as well as direction. Stretch v along its own ray and watch the two numbers come apart:
Watch the shadow grow while the cosine sits still. At the dot product has nearly tripled and the angle has not moved. This is why retrieval ranks with cosine and not with the raw inner product: without the division a long document beats a relevant one for being long, and the bug never raises — it quietly returns the wrong first result.
Now the real thing. Every figure from here on is the same trained space: a skip-gram model, 41 tokens, eight numbers per word, and no supervision beyond which words turn up in the same company. Here are eight of its words on a plane. Drag the query and watch which word it is nearest to:
Notice that the query is a genuine vector in the eight-dimensional space, not a point on a picture: the cosine in the readout is computed against the real rows. The people are on one side, the animals on the other, and no one wrote that down anywhere.
The useful operation is to ask a word for its own neighbours. Pick one and watch the three nearest light up with their cosines:
Because king and queen appear in the same frames — thrones, palaces, crowns — the model has no choice but to score them 0.78, while cat, the furthest of the eight, scores 0.20. That is the entire distributional hypothesis, and it is the only thing the training signal contains.
The picture is a shadow
Every embedding plot you have ever seen threw away most of the space to get onto the page. Ours throws away 28%, and it is enough to change the answers.
The plane in the last section is not the space. It is a projection: the eight vectors have eight numbers each, and the page has room for two. Which two is a choice we made. Switch between two different pairs of directions and watch the map redraw:
Notice that nothing moved. The vectors are identical in both frames; only the two directions we chose to look along changed, and with them 72% of the spread became 41%. The axes are not gender and royalty: they are just the directions these eight rows happen to vary along most.
The part that did not fit is not gone, it is behind the page. Draw it as a ring around each dot whose radius is exactly the length of the off-plane remainder, and ask each word for the neighbour the space gives it:
Watch the two links disagree. For all eight words, the nearest dot on the page is not the nearest word in the space: the page puts king next to man, the space says queen. The plane keeps only 60% of king, and a ring of radius 1.33 is more than the 0.58 that separates the two candidates.
So how many numbers does a word actually need? Every point on this curve is a separate training run of the same corpus at a different width. Drag d down and watch every pair collapse onto each other:
At every pair scores 1.00: one number can only encode a magnitude, so the twelve words become twelve points on one ray and the ranking is a coin flip. Two slots give 0.85, three give 0.53, and by eight the mean has settled near 0.44 and stops improving.
The floor is set by the corpus, not the width: our frames encode four distinctions, so four or five directions are enough and the curve is flat from there. A real corpus encodes vastly more, which is why production widths run 300 (word2vec), 768 (BERT, GPT-2) and 4,096 (Llama-3) rather than 12.
What training actually moves
Two rules, applied a few hundred thousand times: pull the pairs that occur together, push apart the pairs that do not.
Skip-gram takes a word and one of its neighbours in the text and asks a single yes/no question: did these two occur together? The model's answer is σ(w · c), and each step nudges it toward 1. Add gradient steps and watch the target and the context swing toward each other:
Notice that both arrows move: the target swings up and the context swings across. The gradient with respect to w is (σ − 1)·c and the one with respect to c is (σ − 1)·w, so each vector is pushed along the other. At the model scores this real pair 0.09, which is a large error and therefore a large step.
Pull alone has a trivial optimum: make every vector the same enormous vector and every dot product is huge. Something has to push. Skip-gram draws k random words per step and pushes those the other way. Raise k from zero:
Watch what happens at . The unrelated word ends up at 0.85 from the target while the genuine context word is at 0.90 — the model is confident about a pair it was never shown, and the loss went down the whole time. This is the failure that does not raise. word2vec uses 5 to 20 negatives; five is what our trainer draws.
That is the entire algorithm. Run it on the corpus — 41 tokens, 13 frames, 60 epochs — and at epoch 1 all eight words are still one blob near the origin. Each track is one word and the dot on it is where that word is now, so drag through the epochs and watch king and queen part company with the animals:
Watch the first ten epochs do almost all of the work. At epoch 1 every pair looks alike — king/queen 0.95, king/dog 0.99 — and by epoch 10 king/dog has fallen to 0.24 while king/queen holds at 0.81. What prised them apart is the negatives.
Here is the invariant, and it is the one thing to carry away from this section: every update touches only dot products between pairs. Rotate the whole space and every dot product is unchanged, so the loss is too. The solution is therefore defined only up to a rotation — which is why no individual coordinate means anything, in any embedding, ever.
The number being minimised is the negative log-likelihood of those yes/no questions. Drag along the loss:
Notice the curve is not monotone. Each epoch draws fresh negatives, so the loss it reports is a different random subsample of the same objective: 3.77 at the first epoch, 1.19 at the sixtieth, and a great deal of noise in between. A loss that never wobbles under sampled negatives is usually a loss computed on the training minibatch it just fitted.
The geometry nobody asked for
Nothing in the objective mentions directions between words. They show up anyway, and they are both more useful and more fragile than the famous example suggests.
Take the difference between two words — woman minus man — and add it to a third. That difference is just a vector, so it can start anywhere. Add more and more of it to king and watch where the sum lands:
At the sum scores 1.00 with queen and 0.44 with the next candidate. Nobody put a gender direction in this model. It exists because the frames that separate he from she apply the same displacement to every word they touch.
The famous version of this claim overstates it in two ways, and both are visible. First, the sum never lands on the winner. Here are three analogies with the leftover drawn between the sum and the word that wins:
Because the leftover is 0.247 long against a query of 2.90, the parallelogram misses by 8.5% — and by 13% on the animal pair. The answer is the nearest word, not the right one, and “nearest” here also means nearest after the three input words have been struck off the list. Leave them in and king comes back at 0.73, third.
Second, the direction needs room to exist. Every bar below is the cosine between the analogy query and one word, at one of the widths the trainer was actually run at. Drop d and watch the ranking come apart:
At the analogy answers prince at 0.999, with princess, queen and king inside two thousandths of it. The arithmetic is fine; there is no room for gender, rank and age to be separate directions at once, so they share, and the winner is decided by noise.
One more property of the whole space, and it is the one that surprises people with a working retrieval system. Take all 66 pairs of our twelve words and count their cosines into bins — the rule marks a right angle:
Notice there is nothing to the left of it. Every pair scores positive, the mean of the absolute values is 0.44, and the lowest of the 66 is 0.071 — the twelve words sit inside a cone, all within 43.8° to 48.7° of their own centroid. “Unrelated” does not mean orthogonal in a real embedding space.
Subtract the centroid and the cone opens: at the pairwise mean falls to 0.36 and the lowest pair reaches −0.74. That one line of arithmetic is why mean-centring, or whitening, is standard before a nearest neighbour search on embeddings.
Where one vector per word stops
A lookup table has exactly one row per token. Language does not have exactly one meaning per token, and the gap is where the Transformer starts.
Suppose one token had to carry two unrelated senses, the way bank carries a riverside and a financial one. Our corpus has no such word, so we build one: a token that means dog some of the time and king the rest. Every occurrence contributes a gradient, so its row is the frequency-weighted average. Slide the share up from zero:
At the row still scores 0.97 with the majority sense and 0.50 with the minority one: the rare reading is barely represented. At an even split it is 0.86 and 0.75 — near both, a good match for neither.
There is no setting of that slider where the row does the right thing, because the table has one row and the word has two meanings. This is not a bug in word2vec; it is the reason contextual models exist. A Transformer keeps the lookup as its first layer and then lets attention rewrite the vector using the sentence around it, so bank leaves layer 0 as an average and arrives at layer 12 as one sense or the other.
The last thing worth knowing is what the table costs. Every model still starts with one, and the shape is always V × d — pick a model and read the two numbers off its position:
holds 50,257 × 768 = 38.6 M parameters, 154 MB at float32, and it is the single largest matrix in the 124 M-parameter model — 31% of it. GPT-2 ties that table to its output projection and pays for it once; Llama-3 8B does not, so its 128,256 × 4,096 = 525 M sits on the books twice, 13% of an 8 B model.
That table is also why the objective in §04 is shaped the way it is. The correct one is a softmax over the whole vocabulary, which touches every row for every token; sampling touches k + 1 of them. Both bars below are that same table, at the same scale:
The softmax costs O(V·d) — 38.6 M multiply-adds per GPT-2 token — and sampling costs O((k+1)·d), or 4,608 at k = 5. That factor of 8,000 is why word2vec trained on a billion words in 2013, and why the amber bar looks empty.
The four that fail quietly
A shape error raises. These four return a number.
Unnormalised rank. q · v prefers the longest vector, and length tracks frequency. A meaningful axis. The loss is rotation-invariant, so a coordinate, a PCA axis and a t-SNE cluster belong to the plot. The wrong matrix. Skip-gram trains two tables; published results use the input one. The unshifted cone. Raw cosines never go below 0.07 here, so a threshold picked by intuition admits everything.