註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

ali
@waterloo_intern
ml research, kernels, and the occasional peer-reviewed shitpost inference @baseten || eng @uwaterloo
加入 October 2024
117 正在關注    35.6K 粉絲
(friendly fire) every time i've disagreed with ml researchers i've been proven wrong, but i'm going to do it, just one more time we know that the weights hold near-orthogonal directions way beyond dim count (EXPONENTIALLY many if within noise tolerance). this is literally 3 blue 1 brown intro to ML / johnson-lindenstrauss lemma / superposition: unbounded room for knowledge at frontier scale (trillion parameter models) i'd even argue that it's an engineering problem, not a research one. the capacity says that writing new knowledge into existing weights should not overwrite old information if you just store it in a different spot. the substrate you're working on has the capacity to do so. backprop + ce loss is the wrong tool to write with, sure, but doesn't mean it's not possible to write at all. so the research field concluding that "repeatedly writing new facts into a model’s weights eventually overwrites old facts" is, imo, a skill issue less "blind men groping at an elephant", and more "blind men"
顯示更多
0
8
133
7
轉發到社區