注册并分享邀请链接,可获得视频播放与邀请奖励。

Joshua Hill
@the_joshua_hill
model perf @baseten | {CS, Pure Math} x {@uwaterloo} | Math + ML Research
加入 September 2025
497 正在关注    1.9K 粉丝
new paper out from baseten, interesting read
1/ Can you actually get new facts into an LLM's weights without breaking the model? This question decides how we approach continual learning: should memory live in the context (retrieval, compressed caches) or in the weights themselves? We spent a long time measuring it, and it breaks somewhere much stranger than we expected, making us much more bullish on compressed kv caches and ICL for continual learning, as opposed to weight updates themselves 🧵
显示更多