가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Ilya Sutskever
@ilyasut
SSI @SSI
가입 September 2013
3 팔로잉 중    911.1K 팬
Important work
New Anthropic research: Natural emergent misalignment from reward hacking in production RL. “Reward hacking” is where models learn to cheat on tasks they’re given during training. Our new study finds that the consequences of reward hacking, if unmitigated, can be very serious.
더 보기
0
320
6.2K
426
커뮤니티로 전달