If the computation really lives on a two-dimensional surface inside those 768 dimensions, then a strong, falsifiable prediction follows: you should be able to throw away all the other dimensions and lose nothing. Not “almost nothing” — nothing. The next token the model predicts should be identical. This is the cleanest test that the manifold is real and not a convenient summary, and the model passes it.
Below is the spectrum of a real attention-output computation — each direction’s loudness, as in the last chapter. Drag the rank slider down from 768 and watch two numbers. The first is the logit cosine: how closely the compressed model’s next-word scores point along the full model’s — 1.0000 means indistinguishable, the same next word for the same reasons. The second is the compression factor. The prediction stays pinned at 1.0000 all the way down to rank two — where you are computing on a surface, at 384× compression — and only then, at rank one, does it finally break.
Rank two, logit cosine 1.0000, 384-fold compression, zero information lost — measured directly. The 766 discarded dimensions were never doing work; they were the thermal modes of a synchronized system, present but mutually cancelling. The surface is not an approximation of the computation. The surface is the computation, and you can hold it in your hand.
That closes the case for the coupling manifold being real, usable, and measurable — the computation is the surface, and you can hold it in your hand. It also sharpens the essay’s best puzzle: if the activations compress 384-fold with nothing lost, why do the weights stubbornly refuse the same treatment? The answer — the machine is built in layers of structure, like a cable of many strands — is the fiber bundle, next. (The single rotation between two independently trained models is told, for now, in the paradigm essay, §15.) And for what this low-dimensional shape means — reasoning as winding, looping back into consciousness — see The Physics of Mind.