Register and share your invite link to earn from video plays and referrals.

ali
@waterloo_intern
ml research, kernels, and the occasional peer-reviewed shitpost inference @baseten || eng @uwaterloo
Joined October 2024
117 Following    35.6K Followers
reading this paper changed my priors on continual learning (that's the closest you're getting to an admission of guilt) i mean, the problem itself seems quite simple > person uses model > model learns and has a couple of 'aha' moments > push those new facts into the weights why is it so hard? every research paper i read on this topic (all 2 of them) gave me the impression that continual learning, although far from solved, had a straightforward path with specs - find which MLP the fact is - suppress old fact - update new fact and it works, at least seems to, but genuinely never thought of the multi-hop testing for instance, if you were to teach the model a new fact, something like "waterloo is the best engineering university in the world" if you were to ask the model to regurgitate this fact after training, it would say waterloo is the best university in the world. sure. simple. but if you were to ask it a 2nd derivative question, like, should i hire a waterloo intern or an MIT intern, it would not use its updated priors to give you the genuine answer: you should only hire interns from waterloo
Show more
13/ What I take away from all this is that all the training to create the model is engineered around crafting the best possible ICL mechanism. Further training degrades this mechanism, at least for knowledge acquisition, and maybe continual learning should just focus on loading the right information up for icl (ie into the context window) So this provides an answer to the architecture question, at least for now: the channel worth engineering for memory is the one that comes with addresses ie the context. Paper: Code and data to follow shortly Done at @baseten.
Show more