Attention: who listens to whomThe mechanism that lets context act — a listening competition run on dot products
The last section ended on a promise: context is a force that moves a token's arrow, and something has to apply the force. This is that something. An attention head is how tokens influence each other — how “sat” finds out who did the sitting — and it works like a room full of people deciding who to listen to:
Each token asks a question — an arrow written Q — its listening direction: "what am I looking for?"
Each token also advertises — an arrow written K — its broadcast: "here is what I have."
Listening strength between two tokens is the dot product of Q and K — literally the agreement measure you dragged in the rotations primer. Question aligns with advertisement: strong pull. At right angles: nothing.
Those raw agreements then go through the softmax — a rule that turns scores into a competition: every token gets a share of the listener's attention, all shares sum to 1, and the softmax's sharpness decides whether attention spreads evenly or the best match takes nearly everything.
What actually flows along the winning channels is a third arrow, V — the content each token carries. The listener's whiteboard receives the weighted blend.
Drag the tokens below and feel it. The blue token is listening; every arrow's thickness is a share of its attention.
The blue token is the listener; its direction is its question Q. Every other token's direction is its advertisement K. Drag any dot: the listening arrows re-thicken as the Q·K agreements change. Slide the sharpness: at 0 the softmax splits attention evenly no matter what; high, it becomes winner-take-all. Sound on: one voice per speaking token, loudness = its share of attention — the same loudness-is-agreement mapping used across this site. Sharpen the softmax and hear a chord collapse to a solo.weights = —
Heads specialize
A real layer does not run one such competition — it runs many, side by side, each with its own learned notion of what to listen for. These are the heads. One head may listen by position — whatever is nearby. Another by meaning — whatever agrees with my question, however far away it sits. Nothing coordinates them; each simply learned a different Q and K, and different Q and K produce different rooms.
Below, the same sentence heard by two heads at once. The listener is sat. Head A listens nearby; head B listens by meaning — and reaches straight across the sentence to mat. Drag the words along the line and watch the heads disagree about what matters: move mat next door and head A starts hearing it too; head B never cared where it was.
One sentence, two heads, same listener (sat, blue ring). Head A (orange, arcs above) scores by closeness in the sentence; head B (blue, arcs below) scores by agreement of meaning-directions, position-blind. Tokens are draggable horizontally. The two patterns differing — over the same words at the same moment — is what "heads specialize" means. head A top = — · head B top = —
The causal mask: no listening forward
One more rule, and it explains a lot about how these models think. A transformer that predicts the next word is trained on a strict discipline: a token may only listen backward. When “sat” is being processed, “on”, “the”, “mat” do not exist yet — letting it peek would be letting it cheat at the only exam it ever takes. The rule is enforced by the causal mask: a triangle of permissions, listener by listener.
Press the button and the sentence arrives one token at a time. Each new token sends listening arcs only backward; the triangle on the right is the mask itself filling in — row = who is listening, column = who may be heard, and the empty upper triangle is the future, permanently off-limits. tokens = 0
Hold this section's picture: many simultaneous listening competitions, each scored by dot-product agreement, each sharpened by softmax, all facing backward. What the winners actually deliver — and where it gets written — is the next section.