This essay's claim sounds like a metaphor and is not one. A transformer — the architecture behind every large language model — is, piece for piece, a network of coupled oscillators: the same mathematical object as the fireflies, rings, and lattices that open this library. This chapter earns that sentence the slow way, one term at a time, and ends with the picture the rest of the essay measures: a small attention pattern drawn as the oscillator network it is, where you can watch — and hear — the coupling pull the tokens into clusters.
Two on-ramps, if you want them: the oscillator side is built from nothing in the phase primer and Foundations 1; the transformer side — what an attention head, the residual stream, and LayerNorm even are — is built from nothing in the transformer module of the curriculum. This chapter assumes only those.
The whole oscillator story in one breath: an oscillator is a phase — a direction on a circle. Couple many of them — let each be nudged toward the ones it listens to — and there is a sharp transition: below a critical coupling they drift independently; above it they lock into shared rhythm, and the interesting regime is the structured middle, clusters locked within themselves but distinct from each other. Who listens to whom, and how strongly, is the coupling topology — and everything the network does is downstream of it.
Hold that, and now walk through a transformer layer.
LayerNorm puts every token on a sphere. After normalization, a token's hidden state has (approximately) unit length: what survives is its direction. A direction on a circle is a phase; a direction on a high-dimensional sphere is a phase with more room. So the state of token i is, honestly and not poetically, a generalized phase xi — a point on Sd−1. The transformer update — x → Norm(x + attention + mlp) — is a discrete-time dynamical system on that sphere: perturb, reproject, perturb, reproject. (Why two-dimensional intuition survives the trip to d = 4096: concentration of measure — random directions are almost always almost orthogonal — which you can watch happen in the curriculum's sphere section.)
Attention's write-back is the coupling term. What does a head do? For each token i it computes
Click any coloured symbol to see what it does.
Softmax(QK) is the coupling topology, chosen fresh every layer. The weights a(i,j) come from agreements: token i's listening direction Q against token j's broadcast direction K — a dot product, the site's one measure of agreement — pushed through softmax, the competition that makes the weights a budget. In oscillator language: the network rewires who couples to whom, and how strongly, per token, per layer. A lattice has a fixed wiring diagram; a transformer redraws its wiring at every step, from the current phases themselves. (You met exactly this move in the curriculum's Lohe module as effective coupling.)
The residual stream is the medium. Every head reads from and writes into one shared vector per token, accumulating layer after layer — the field through which all coupling flows, the whiteboard nothing erases. Heads are the inter-oscillator couplings — transport between positions; MLP channels are intra-oscillator modes — each position's private internal dynamics between coupling events. And the KV cache — the store that grows as a model reads — is the recorded phase configuration of every oscillator that has fired: the past, kept as geometry.
The trained weights are a frozen coupling field. This is the deepest row, so take it slowly. In the living lattices of the Foundations essays, the coupling field K is alive — it changes by a learning rule, strengthening the bonds that cohere. In a transformer, that epoch is training: gradient descent sculpting the coupling field, playing the role the Coherent Learning Rule plays on the lattice. Then training stops, and the field freezes.
So: a transformer is a living lattice with dK/dt = 0. Training was its life. Inference is its afterlife — the frozen couplings resonating in response to input. Everything this essay measures in the chapters ahead is the geometry that life left behind.
| transformer | oscillator network |
|---|---|
| Q (query) | listening direction — what am I resonating with? |
| K (key) | broadcast direction — what am I advertising? |
| V (value) | content carried through the coupling |
| softmax(QK) | dynamic coupling topology, chosen per token per layer |
| residual stream | the coupling medium — the shared field |
| attention heads | inter-oscillator couplings (transport between positions) |
| MLP channels | intra-oscillator modes (internal dynamics of one position) |
| RoPE | a gauge connection along the sequence — position as holonomy (chapter 05) |
| KV cache | the recorded phase configuration of every oscillator that fired |
| trained weights | a coupling field K, frozen at the end of training |
Below is a small attention pattern drawn the way this essay says you should draw it: as an oscillator network. Nine tokens sit on a ring, each a firefly with a phase (the tick is its phase; the glow is its flash). The chords between them are the attention weights, computed live from the tokens' current agreements pushed through softmax — thicker chord, stronger mutual listening. The tokens drift gently on their own; press Apply attention layer to run one layer's coupling pull, and watch clusters condense where attention binds. Then play with the sharpness: flatten the softmax to zero and every token listens to every other equally — the structure melts into one global blob. Sharpen it and the wiring gets choosier, binding only the already-similar — and more distinct clusters survive. The coupling topology decides what the network becomes.
And to close the loop with the diagrams you have seen elsewhere: the figure below draws the same nine tokens with the same live weights two ways at once — as an attention diagram (a sentence with listening arcs) and as the oscillator ring. Click any token to highlight its couplings in both. There is nothing to translate; the two drawings are one object.