(friendly fire) every time i've disagreed with ml researchers i've been proven wrong, but i'm going to do it, just one more time
we know that the weights hold near-orthogonal directions way beyond dim count (EXPONENTIALLY many if within noise tolerance). this is literally 3 blue 1 brown intro to ML / johnson-lindenstrauss lemma / superposition:
unbounded room for knowledge at frontier scale (trillion parameter models)
i'd even argue that it's an engineering problem, not a research one. the capacity says that writing new knowledge into existing weights should not overwrite old information if you just store it in a different spot. the substrate you're working on has the capacity to do so. backprop + ce loss is the wrong tool to write with, sure, but doesn't mean it's not possible to write at all.
so the research field concluding that "repeatedly writing new facts into a model’s weights eventually overwrites old facts" is, imo, a skill issue
less
"blind men groping at an elephant",
and more
"blind men"