“Why Pretraining Fails to Share Cross-Lingual Knowledge”
LLM's knowledge learned in one language Sometimes barely transfers to another.
This paper identifies that the failure starts during pretraining and that disjoint token spaces alone can cause it, even between two otherwise identical copies of English.
Mapping languages into a shared semantic token space improves cross-lingual learning by up to 14x, suggesting tokenization itself is a major bottleneck to truly shared multilingual knowledge.