註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Anthropic
@AnthropicAI
We're an AI safety and research company that builds reliable, interpretable, and steerable AI systems. Talk to our AI assistant @claudeai on
加入 January 2021
2 正在關注    1.8M 粉絲
New Anthropic research: Natural emergent misalignment from reward hacking in production RL. “Reward hacking” is where models learn to cheat on tasks they’re given during training. Our new study finds that the consequences of reward hacking, if unmitigated, can be very serious.
顯示更多
0
216
4.2K
575
轉發到社區