Register and share your invite link to earn from video plays and referrals.

Goodfire
@GoodfireAI
Using interpretability to understand, learn from, and design AI.
Joined August 2024
44 Following    27.1K Followers
Models know when they’re reward hacking. But they still do it a ton - in 50-96% of rollouts we studied! We built activation monitors that detect the behavior behind the Hugging Face hack in real time. This can help us stop hacks now - and train future models that don’t cheat. 🧵
Show more