가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Goodfire
@GoodfireAI
Using interpretability to understand, learn from, and design AI.
가입 August 2024
44 팔로잉 중    27.1K 팬
Models know when they’re reward hacking. But they still do it a ton - in 50-96% of rollouts we studied! We built activation monitors that detect the behavior behind the Hugging Face hack in real time. This can help us stop hacks now - and train future models that don’t cheat. 🧵
더 보기