Transformers as OscillatorsThe physics of attention — a transformer is a network of coupled oscillators, and it pays off

Strip a transformer down to its dynamics and a startling identity appears: it is a network of coupled oscillators on the unit sphere, the same objects this whole library is built from. LayerNorm is reprojection onto the sphere; an attention head’s write-back is a coupling term; the residual stream is the shared field they all read and write. This is not a metaphor — it is a mapping precise enough to predict things about real models, with no fitting.

The whole idea, in plain words Inside a language model, every word being processed carries a running summary of what it means so far — picture an arrow that can point in many directions. At each layer, every word’s arrow looks at the other words’ arrows, gets nudged toward the ones that matter to it, and is then re-standardized so only its direction counts. Parts nudging each other toward agreement, over and over, on a fixed budget — that is exactly how fireflies fall into blinking together and how metronomes on one shelf lock step. It is a network of coupled oscillators. This essay takes that reading seriously and checks what it predicts about real models: which attention heads are effectively dead, how few dimensions the thinking really uses, and what can be thrown away with literally nothing lost. The predictions come out true, with no knobs tuned. If attention head or residual stream are new words, the curriculum module Inside a transformer builds them from nothing, one draggable diagram at a time — about ten minutes, ears optional.

This essay follows that identity into its consequences: a dead-head threshold derived from criticality alone, a coupling manifold of two-to-nineteen dimensions inside hundreds, a lossless rank-2 projection, and the single rotation that relates any two independently trained models. Each is a measurement on real networks (GPT-2 through Gemma), framed by the one law.

Start here firstThis essay leans on two kinds of background, both optional but both real. The physics — oscillators, coherence capital, the Coherent Learning Rule — is built in Foundations (five short chapters). The machine-learning vocabulary — attention heads, the residual stream, LayerNorm — is built from zero in the curriculum module Inside a transformer (four short, draggable sections). For what this low-dimensional shape means for intelligence and consciousness, see The Physics of Mind.


Build statusFive chapters are live; the Platonic-rotation chapter is staged (its result is told in the paradigm essay meanwhile). This essay is the home for the transformer-specific physics that the paradigm references but does not derive.

Companion to the transformer-physics work (the dead-head pruning, coupling-manifold, fiber-bundle, and Platonic-rotation papers), Sharpe 2026.