Register and share your invite link to earn from video plays and referrals.

Ilya Sutskever
@ilyasut
SSI @SSI
Joined September 2013
3 Following    911.1K Followers
Important work
New Anthropic research: Natural emergent misalignment from reward hacking in production RL. “Reward hacking” is where models learn to cheat on tasks they’re given during training. Our new study finds that the consequences of reward hacking, if unmitigated, can be very serious.
Show more
0
320
6.2K
426
Forward to community