註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Ilya Sutskever
@ilyasut
SSI @SSI
加入 September 2013
3 正在關注    911.1K 粉絲
Important work
New Anthropic research: Natural emergent misalignment from reward hacking in production RL. “Reward hacking” is where models learn to cheat on tasks they’re given during training. Our new study finds that the consequences of reward hacking, if unmitigated, can be very serious.
顯示更多
0
320
6.2K
426
轉發到社區