Training data for LLMs is made up of many sources. Given that, what structure should we expect in the activations?
New work from Simplex shows how the belief geometry over this type of data forms telescoping cones, and transformers represent them! ๐งต๐