Tokens and embeddingsThe first move a transformer makes: every word becomes an arrow, and direction becomes meaning

Half of this library talks about transformers — the machinery inside systems like the one you may be reading about right now — and it leans on vocabulary nobody grew up with. This module builds all of it from nothing, one section per idea, because each idea is a simple physical picture wearing an engineering name. If you have taken the rotations primer, you already own the only math involved: arrows, and the dot product as a measure of agreement. This first section is about the raw material everything else acts on: the token.

1. A token is an arrow in meaning space

A transformer reads text in pieces called tokens — roughly, words. The first thing it does with a token is look up its meaning as a list of numbers. A list of numbers is an arrow: two numbers give an arrow on this page, three an arrow in the room, and a real model uses hundreds or thousands — an arrow in a space you cannot picture but can reason about exactly the same way. The space is organized so that direction is meaning: words that mean similar things point similar ways, and the dot product — the agreement measure from the rotations primer — tells you how related two meanings are.

Drag the arrowhead anywhere. The readout shows the two numbers that are the arrow, and which word's direction it currently agrees with most. A real model does exactly this with thousands of numbers per token instead of two. arrow = —

2. Neighborhoods: meaning comes in clusters

The lookup table that assigns each word its arrow is called the embedding, and it is not arbitrary — it is learned, and what it learns is geography. Words that behave alike in text end up living near each other: a warm district, a cold district, a night-sky district. Once meaning is geography, "what does this word mean?" gets a mechanical answer: look at what its arrow points along with. Below is a toy embedding with three districts. Drag the probe arrow around and watch the neighborhood ranking reorder itself — this ranking, by dot product, is the primitive that everything smarter is built from.

Twelve words, three districts, one draggable probe (orange). The panel ranks every word by its agreement with the probe, live. Point into the cold district and watch ice, snow, winter rise together — clusters move as clusters, which is exactly what makes direction usable as meaning. top = —

3. One word, many contexts

There is a problem, and it is the problem the rest of the transformer exists to solve. Words do not have one meaning. Cool is a temperature in one sentence and a compliment in another — but the embedding table has exactly one arrow per token. The stored arrow for an ambiguous word sits between its meanings, committed to neither.

What resolves it is context: the other words in the sentence pull the token's working copy toward the meaning that fits. Press the two context buttons below and watch the stored arrow bend.

The stored embedding for cool (dashed) sits between its two meanings. Each context pulls the working copy (orange) toward a different district, and the nearest-neighbor readout flips. The machinery that does this pulling in a real model is attention — the next section. context = none

Hold onto both pictures from this section: meaning is a direction, and context is a force that moves it. The next section opens the machine that applies the force.