The last two chapters left a paradox on the table. The coupling manifold showed a transformer's activity collapsing to a handful of effective dimensions — d_eff ≈ 2–5 inside 768. The lossless projection showed you can compute on a rank-2 surface with a logit cosine of 1.0000. So the motion is tiny. Then why is the machine so big? When the compression program turned from activations to the dimensions that carry routing, it hit a stubborn, repeatable wall: the room the machinery occupies is 2.5–2.8× larger than the dance inside it, on every architecture measured. A system that is small in motion and large in wiring is not being wasteful. It is telling you what it is. This chapter names the geometry: a fiber bundle.
Hold both numbers at once. Activity: a few dimensions suffice — project everything else away and the next token does not change. Machinery: when you measure not the dance but the room the routing uses — the total dimensions that participate in deciding who listens to whom — you get d_total ≈ 2.7 × d_eff on GPT-2 and 2.8× on SmolLM2. And this ratio is the robust thing: a later recalibration of the measurement pipeline roughly doubled both absolute numbers on both models — and the multiplier didn't move. Absolute dimensions were an artifact of the ruler; the ratio of room to dance is a property of the machine.
Ordinary compression logic says this is a bug: if only two dimensions matter, everything else is fat. The bundle picture says the opposite: the extra dimensions are not fat, they are infrastructure — and infrastructure is exactly the thing you cannot quantize carelessly.
Here is the geometry, in the vocabulary you already own if you have walked the gauge & holonomy primer. The sequence is a base space: position 1, position 2, position 3 — a line of sites. Over each position hangs a fiber: the content space, the room where that token's meaning-arrow lives. And crucially, each fiber has its own private frame — its own choice of which way is “north.” That is precisely the primer's nine towns: every town points north its own way, and no physics changes if a town re-chooses — provided you keep the dictionary.
The dictionary between neighboring fibers is the connection — and in a transformer it has a name you know: RoPE. Rotary position embedding rotates each position's frame by an angle proportional to where it sits. It is literally the turntable stack: pair up the dimensions, and rotate each pair's plane at its own rate as you move along the sequence. Position stops being a label stapled to a token and becomes parallel transport: where you are is how far your frame has turned — position as holonomy, the rotation you accumulate by being carried along the base.
This is why attention works the way it does. Two tokens' content-arrows cannot be compared raw — they live in different frames, and the raw dot product mixes real disagreement with mere difference-of-north. Attention's q·k compares them after transport to a common chart: a gauge-covariant coupling, agreement measured in a shared frame. Drag it yourself:
Now the paradox resolves. The high dimensions are the room: fibers, frames, and the transport machinery between them. The low dimensions are the dance: the actual dynamics moving through that room — the coupling manifold, whose natural equations are the sphere-dwelling oscillator dynamics of the Lohe primer. Many dimensions participate in the simulation; few dimensions carry the motion. The next figure makes that sentence a thing you can rotate:
That is the shape of the measured transformer: a d_eff ≈ 2–5 dance inside a d_total ≈ 2.5–2.8 × d_eff room of participating machinery, inside an ambient width of 768 or more. And the division of labor is measurable a second way: the model's decisions are two to three orders of magnitude more sensitive to the K side of its cache — the routing, the coupling allocation — than to the V side, the transported content. What the machine protects is not what it says; it is who gets to listen to whom.
And now the compression asymmetry explains itself. The activity is the dance: project it to the manifold and you lose only thermal residue — hence rank-2 lossless. The weights carry the bundle's wiring: the frames, and the dictionaries between them, for every transport the sequence might ever need. Corrupt the dance a little and errors stay where they land. Corrupt the dictionary a little and the errors compound along transport — every step of carrying multiplies the damage, and a comparison across twelve positions inherits twelve steps of accumulated mistranslation. The same precision knob is catastrophic in one place and harmless in the other:
Step back and the essay's picture completes itself. A transformer is an oscillator network (the machinery) whose motion lives on a low-dimensional manifold (chapter 3) you can compute on directly (chapter 4) — and the high-dimensional bulk that remains is not waste but geometry: the bundle of frames and dictionaries that lets a line of positions behave like one connected medium. Many dimensions participate. Few dimensions dance. The dimensions that participate without dancing are the room the dance needed — and the machine guards its room more jealously than its steps.