註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Goodfire
@GoodfireAI
Using interpretability to understand, learn from, and design AI.
加入 August 2024
44 正在關注    27.1K 粉絲
Models know when they’re reward hacking. But they still do it a ton - in 50-96% of rollouts we studied! We built activation monitors that detect the behavior behind the Hugging Face hack in real time. This can help us stop hacks now - and train future models that don’t cheat. 🧵
顯示更多
0
22
611
79
轉發到社區