Every large transformer carries attention heads that barely matter. Pruning researchers have known this for years the empirical way: ablate a head, measure the damage, keep a list. That works, but it is bookkeeping, not understanding — every new model means starting over. The oscillator picture makes a sharper offer: if a head really is an oscillator coupled to a shared field, then physics already has a theory of when a coupled thing stops being coupled. The theory should hand us a threshold — a single number, derived rather than fitted — below which a head is predictably dead. It does. This chapter is that number.
You have met the core mechanism already, in the learning-rule chapter: a coupled bond has a sharp aliveness threshold. Below it, the bond cannot hold its two ends in phase — noise and detuning win, alignment drifts, and the equilibrium coupling falls to zero. Above it, the bond locks and holds. There is no gentle middle; it is a bifurcation, a cliff. On the circle, the critical alignment at which a bond can no longer sustain itself works out to cos Δθ ≈ 0.679 — a number that comes out of the same oscillator mathematics that runs this whole library.
Now put that cliff inside a transformer. Each attention head writes its output back into the residual stream — that write-back is the head's coupling to the collective (section 01). Measure the alignment between what the head writes and where the stream actually points. If that alignment is strong, the head is entrained: a live voice in the chorus. If it falls below the aliveness threshold, the head is decoupled — spinning on its own, its contribution averaging away against everyone else's. The physics word for that state is dead.
One translation step remains. The circle's threshold is quoted in circle units, where "unrelated" means an alignment near zero and chance alignments are large. A transformer's states live on a sphere in d dimensions — 768 of them for GPT-2, 4096 for a 7-billion-parameter model — and high dimensions change what "chance" looks like. You saw this measured live in the transformer primer's concentration figure: the dot product of two random directions in d dimensions is not spread over the whole range — it concentrates near zero with spread 1/√d. Chance alignment has a noise floor, and the floor drops as the room grows.
A threshold for "genuinely coupled, not just chance" must therefore ride that floor. Carrying the circle's critical alignment onto the d-sphere by concentration of measure gives:
— no fitted parameters. Nothing about any particular model appears on the right-hand side.
A head whose mean write-back alignment with the residual stream falls below τdeath is predicted dead: removable with negligible damage, because the collective was already ignoring it. For GPT-2 that bar sits at 0.035; for a 4096-dimensional model, 0.015. The formula was then tested the only way that counts — against ablation ground truth — on five published architectures (GPT-2, GPT-2-medium, Qwen2.5-0.5B, SmolLM2-360M, OpenLLaMA-7B; d from 768 to 4096). Precision: 95.2–100%. Same constant, same formula, every model. That is what physics transferring looks like, as opposed to a heuristic being re-tuned.
The threshold is a cliff, and cliffs are best experienced. Below: ten oscillator-heads coupled to one strong shared rhythm — the residual stream, which for any single head in a large model is effectively a collective field much bigger than itself. Each head has its own fixed coupling strength (spoke thickness) and its own private agenda: a pull toward its own rhythm, which the slider strengthens. A bond stays entrained only while its coupling can out-pull that agenda — the same lock condition you heard as beats-then-lock in the phase primer, run in reverse. Each head's voice sounds at its own pitch, and its loudness is its measured, live alignment with the field — nothing is scripted. Slide right and listen: heads drop out of the chorus one by one, strictly in coupling order. What is left when a head dies is not silence — it is a voice that no longer moves with the music.
The payoff of a predictive threshold is that you can act on it before measuring anything. Below, fourteen heads whose couplings come in two groups — a cluster well under the bar and a cluster well over it, mirroring the separation that makes a precise threshold possible in real trained models (a threshold classifier can only be 95%+ precise if the population it sorts is actually split). Prune everything below threshold and watch the collective signal — the steady component of the summed write-back — barely move: the dead heads' contributions were spinning past the field and averaging to nothing. Then the control: prune the same number of heads, but take the strongest instead. The collective collapses. The threshold is not "remove some heads and hope" — it selects exactly the ones the collective had already stopped hearing.
One more measured fact keeps this honest and makes it stranger. When the "dead" heads of a real model are examined individually, about 97% of them are rotation specialists: their write-back is nearly orthogonal to the stream's mean direction not because they do nothing, but because what they do is turn the state rather than push it — motion within the sphere's surface rather than along the collective's axis (the distinction you handled in the rotation-groups primer). Dead, in this chapter, means dead to the chorus: decoupled from the mean field, prunable at 95%+ precision when what you care about is the collective output. It does not mean the head learned nothing. The oscillator picture predicts who can be removed; it also warns you not to confuse removable with empty.