The residual streamA whiteboard that rides through the machine — and why nothing is ever overwritten

You now know what a token is (an arrow) and how tokens influence each other (attention's listening competition). This section is about where all of that happens — the one design decision that makes a transformer's insides legible at all.

1. The riding whiteboard

A transformer is built as a stack of identical layers — a few dozen of them. Each token's arrow is written on a kind of whiteboard that rides through the whole stack, and every layer only adds small nudges to what is written there. Nothing is erased. Nothing is overwritten. The arrow that leaves layer 30 is the original meaning plus every adjustment every layer chose to make along the way.

That riding whiteboard is the residual stream. When the essays call it "the shared field" or "the coupling medium," this is all they mean: the one persistent state that every part of the machine reads from and writes into, the way every voice in a room moves the same air.

Three tokens ride left to right through the layers of the stack. The arrow inside each dot is that token's current meaning; at each layer, small writes nudge it — some from the token's own lane, some (the faint arcs) reaching over from other tokens' lanes. Those arcs are the attention you met last section. The lane itself is the residual stream: a running sum that carries everything forward.

2. The state is a sum — take it apart

Because layers only ever add, a token's state at the top of the stack is literally a sum: the original embedding, plus layer 1's contribution, plus layer 2's, and so on. That is not a metaphor you are being sold — it is arithmetic you can perform. Below, one token's journey is laid out as its pieces, head to tail. Toggle any layer's contribution off and the final arrow recomputes: the whole state is nothing but its parts.

The story in the numbers: the token is the word it — a pronoun, almost meaningless on its own — in a sentence about a cat. Watch what the layers add.

contributions:
The faint gray arrow is the stored embedding of it; each colored segment is one layer's added contribution, drawn head to tail; the bold orange arrow is their sum — the state. Un-tick a layer and its segment vanishes from the chain, and the sum moves. Nothing is hidden anywhere else: the state is the sum of what was written. state = —

3. Read the board at any depth

One consequence deserves its own figure. Since the stream carries a plain running sum, anyone can read it at any point — you do not have to wait for the top of the stack. Slide the probe below through the layers and watch the same token's meaning sharpen: at depth 0 the board just says "it" — a pronoun, pointing nowhere in particular; each layer's write bends it further; by the top it reads, unambiguously, cat. Interpretability researchers do exactly this to real models — insert a probe at layer k and ask what the board says so far.

The same token, probed at increasing depth. The readout is what a probe at that layer would report: the nearest meaning-direction and how strongly the state agrees with it. Ambiguity at the bottom, commitment at the top — and every intermediate stage fully readable, because the stream never hides its working. depth 0 — reads: —

The board, the writes, the running sum: that is the residual stream. One question remains — what keeps forty layers of enthusiastic addition from blowing the arrows up entirely? The answer puts the whole computation on a sphere, and hands this module over to the rest of the library.