注册并分享邀请链接,可获得视频播放与邀请奖励。

Goodfire
@GoodfireAI
Using interpretability to understand, learn from, and design AI.
加入 August 2024
44 正在关注    27.1K 粉丝
Models know when they’re reward hacking. But they still do it a ton - in 50-96% of rollouts we studied! We built activation monitors that detect the behavior behind the Hugging Face hack in real time. This can help us stop hacks now - and train future models that don’t cheat. 🧵
显示更多
0
22
611
79
转发到社区