Attention: who listens to whomThe mechanism that lets context act — a listening competition run on dot products

The last section ended on a promise: context is a force that moves a token's arrow, and something has to apply the force. This is that something. An attention head is how tokens influence each other — how “sat” finds out who did the sitting — and it works like a room full of people deciding who to listen to:

Drag the tokens below and feel it. The blue token is listening; every arrow's thickness is a share of its attention.

The blue token is the listener; its direction is its question Q. Every other token's direction is its advertisement K. Drag any dot: the listening arrows re-thicken as the Q·K agreements change. Slide the sharpness: at 0 the softmax splits attention evenly no matter what; high, it becomes winner-take-all. Sound on: one voice per speaking token, loudness = its share of attention — the same loudness-is-agreement mapping used across this site. Sharpen the softmax and hear a chord collapse to a solo. weights = —

Heads specialize

A real layer does not run one such competition — it runs many, side by side, each with its own learned notion of what to listen for. These are the heads. One head may listen by position — whatever is nearby. Another by meaning — whatever agrees with my question, however far away it sits. Nothing coordinates them; each simply learned a different Q and K, and different Q and K produce different rooms.

Below, the same sentence heard by two heads at once. The listener is sat. Head A listens nearby; head B listens by meaning — and reaches straight across the sentence to mat. Drag the words along the line and watch the heads disagree about what matters: move mat next door and head A starts hearing it too; head B never cared where it was.

One sentence, two heads, same listener (sat, blue ring). Head A (orange, arcs above) scores by closeness in the sentence; head B (blue, arcs below) scores by agreement of meaning-directions, position-blind. Tokens are draggable horizontally. The two patterns differing — over the same words at the same moment — is what "heads specialize" means. head A top = — · head B top = —

The causal mask: no listening forward

One more rule, and it explains a lot about how these models think. A transformer that predicts the next word is trained on a strict discipline: a token may only listen backward. When “sat” is being processed, “on”, “the”, “mat” do not exist yet — letting it peek would be letting it cheat at the only exam it ever takes. The rule is enforced by the causal mask: a triangle of permissions, listener by listener.

Press the button and the sentence arrives one token at a time. Each new token sends listening arcs only backward; the triangle on the right is the mask itself filling in — row = who is listening, column = who may be heard, and the empty upper triangle is the future, permanently off-limits. tokens = 0

Hold this section's picture: many simultaneous listening competitions, each scored by dot-product agreement, each sharpened by softmax, all facing backward. What the winners actually deliver — and where it gets written — is the next section.