注册并分享邀请链接,可获得视频播放与邀请奖励。

Ilya Sutskever
@ilyasut
SSI @SSI
加入 September 2013
3 正在关注    911.1K 粉丝
Important work
New Anthropic research: Natural emergent misalignment from reward hacking in production RL. “Reward hacking” is where models learn to cheat on tasks they’re given during training. Our new study finds that the consequences of reward hacking, if unmitigated, can be very serious.
显示更多
0
320
6.2K
426
转发到社区