The fiber bundleThe low-dimensional machine in the high-dimensional room — and why weights resist the compression activations allow

The last two chapters left a paradox on the table. The coupling manifold showed a transformer's activity collapsing to a handful of effective dimensions — d_eff ≈ 2–5 inside 768. The lossless projection showed you can compute on a rank-2 surface with a logit cosine of 1.0000. So the motion is tiny. Then why is the machine so big? When the compression program turned from activations to the dimensions that carry routing, it hit a stubborn, repeatable wall: the room the machinery occupies is 2.5–2.8× larger than the dance inside it, on every architecture measured. A system that is small in motion and large in wiring is not being wasteful. It is telling you what it is. This chapter names the geometry: a fiber bundle.

Two measurements that shouldn't coexist

Hold both numbers at once. Activity: a few dimensions suffice — project everything else away and the next token does not change. Machinery: when you measure not the dance but the room the routing uses — the total dimensions that participate in deciding who listens to whom — you get d_total ≈ 2.7 × d_eff on GPT-2 and 2.8× on SmolLM2. And this ratio is the robust thing: a later recalibration of the measurement pipeline roughly doubled both absolute numbers on both models — and the multiplier didn't move. Absolute dimensions were an artifact of the ruler; the ratio of room to dance is a property of the machine.

Ordinary compression logic says this is a bug: if only two dimensions matter, everything else is fat. The bundle picture says the opposite: the extra dimensions are not fat, they are infrastructure — and infrastructure is exactly the thing you cannot quantize carelessly.

The bundle: a room over every position

Here is the geometry, in the vocabulary you already own if you have walked the gauge & holonomy primer. The sequence is a base space: position 1, position 2, position 3 — a line of sites. Over each position hangs a fiber: the content space, the room where that token's meaning-arrow lives. And crucially, each fiber has its own private frame — its own choice of which way is “north.” That is precisely the primer's nine towns: every town points north its own way, and no physics changes if a town re-chooses — provided you keep the dictionary.

The dictionary between neighboring fibers is the connection — and in a transformer it has a name you know: RoPE. Rotary position embedding rotates each position's frame by an angle proportional to where it sits. It is literally the turntable stack: pair up the dimensions, and rotate each pair's plane at its own rate as you move along the sequence. Position stops being a label stapled to a token and becomes parallel transport: where you are is how far your frame has turned — position as holonomy, the rotation you accumulate by being carried along the base.

This is why attention works the way it does. Two tokens' content-arrows cannot be compared raw — they live in different frames, and the raw dot product mixes real disagreement with mere difference-of-north. Attention's q·k compares them after transport to a common chart: a gauge-covariant coupling, agreement measured in a shared frame. Drag it yourself:

Seven positions on the base line, a fiber over each, frames turning with position (RoPE). Positions 1 and 5 store the same content in their own frames. Compared raw, the arrows disagree — that is frame difference, not meaning difference. Press Compare to transport position 5's arrow back through the chain (watch the ghost copies de-rotate, fiber by fiber) into position 1's chart: agreement snaps to ≈ 1. Sound on: loudness is the agreement — hear it jump when the comparison is done in a common chart. raw agreement = —

Many dimensions participate; few dimensions dance

Now the paradox resolves. The high dimensions are the room: fibers, frames, and the transport machinery between them. The low dimensions are the dance: the actual dynamics moving through that room — the coupling manifold, whose natural equations are the sphere-dwelling oscillator dynamics of the Lohe primer. Many dimensions participate in the simulation; few dimensions carry the motion. The next figure makes that sentence a thing you can rotate:

Twenty-four unit vectors performing a three-cluster dance confined (up to a whisper of noise) to one plane — embedded in a room of dimension d. Slide the room from 8 up to 512 dimensions: the measured effective dimension d_eff — computed live from the actual covariance — does not budge from ≈ 2. Then rotate the view away from the dance plane toward two random room directions: the dance shrinks to a speck, because almost every direction in the room contains almost none of it. The room got bigger; the dance didn't. d = 8   d_eff = —

That is the shape of the measured transformer: a d_eff ≈ 2–5 dance inside a d_total ≈ 2.5–2.8 × d_eff room of participating machinery, inside an ambient width of 768 or more. And the division of labor is measurable a second way: the model's decisions are two to three orders of magnitude more sensitive to the K side of its cache — the routing, the coupling allocation — than to the V side, the transported content. What the machine protects is not what it says; it is who gets to listen to whom.

Why the weights resist

And now the compression asymmetry explains itself. The activity is the dance: project it to the manifold and you lose only thermal residue — hence rank-2 lossless. The weights carry the bundle's wiring: the frames, and the dictionaries between them, for every transport the sequence might ever need. Corrupt the dance a little and errors stay where they land. Corrupt the dictionary a little and the errors compound along transport — every step of carrying multiplies the damage, and a comparison across twelve positions inherits twelve steps of accumulated mistranslation. The same precision knob is catastrophic in one place and harmless in the other:

Two tokens, twelve positions apart, carrying the same content. Top: quantize the dance — store each content-arrow coarsely. The comparison suffers one rounding error at each end, and agreement barely moves. Bottom: quantize the dictionary — store each step's transport-rotation coarsely. Twelve rounding errors compound along the carry, and agreement collapses as precision drops. Sound on: two voices, one per strategy, loudness = agreement — listen to the lower voice die first as you coarsen. This is the shape of the measured asymmetry: activations forgive; wiring does not. content-quantized = —   dictionary-quantized = —
What's real here Tier by tier. Measured: d_eff ≈ 2–5 and the rank-2 lossless projection (chapters 3–4); the routing multiplier d_total/d_eff = 2.72× (GPT-2) and 2.80× (SmolLM2), stable across a recalibration that roughly doubled both absolute dimension counts — two architectures, so a robust pair, not yet a law; the K-vs-V sensitivity asymmetry (direction and order of magnitude measured on a single model and run; the specific “200–800×” magnitudes are estimator-sensitive). Derived and implemented: transport preservation — inverse-RoPE to a common chart before merging is literally what the compression machinery does, not a metaphor for it. Design frame: the bundle language itself, and “a transformer is a living lattice with dK/dt = 0.” The figures on this page are cartoons built so you can hold the geometry — the two-bit-strategy toy illustrates why wiring resists; it is not the measurement, which lives in the fiber-bundle compression work.

Step back and the essay's picture completes itself. A transformer is an oscillator network (the machinery) whose motion lives on a low-dimensional manifold (chapter 3) you can compute on directly (chapter 4) — and the high-dimensional bulk that remains is not waste but geometry: the bundle of frames and dictionaries that lets a line of positions behave like one connected medium. Many dimensions participate. Few dimensions dance. The dimensions that participate without dancing are the room the dance needed — and the machine guards its room more jealously than its steps.