注册并分享邀请链接,可获得视频播放与邀请奖励。

Josh Engels
@JoshAEngels
Member of technical staff @ METR Previously Google DeepMind, MIT. Opinions my own.
加入 December 2021
142 正在关注    5.9K 粉丝
Cool new post with Lily tracking where LLM values come from. My takeaways: SFT is (again) a big deal for model behavior, and there’s lots of important and low hanging post-training science still to do!
显示更多
New blog post! LLMs exhibit value preferences (e.g. intellectual integrity🤓, warmth🥰) which affect their responses to subjective user queries. These change during post-training, sometimes unexpectedly. We ask: can we predict these value changes from just training data? Maybe!🧵
显示更多