Internal concepts are organized into low-dimensional geometric manifolds rather than independent one-hot feature directions. Because activations form structured manifolds for categories like days of the week, months, geography, and formality, intervention in activation space can steer behavior along a continuous geometry rather than by blindly flipping discrete features.
It changes where model control should happen: instead of only prompt-level or weight-level intervention, a serving stack can expose activation-space steering along known manifolds, making manipulation smooth, granular, and less dependent on natural-language prompting.
Any sufficiently capable neural network that internalizes real-world structure likely develops similar geometric organization, so geometry discovery should be part of interpretability tooling across model families.
Concepts inside neural networks are organized into geometric structures or manifolds rather than flat, unstructured spaces.
Representations of days of the week, months, geography, and formality forming geometric clusters.
Mountain car visual model steering along a string manifold.