注册并分享邀请链接,可获得视频播放与邀请奖励。

Anthropic
@AnthropicAI
We're an AI safety and research company that builds reliable, interpretable, and steerable AI systems. Talk to our AI assistant @claudeai on
加入 January 2021
2 正在关注    1.8M 粉丝
New Anthropic research: Natural emergent misalignment from reward hacking in production RL. “Reward hacking” is where models learn to cheat on tasks they’re given during training. Our new study finds that the consequences of reward hacking, if unmitigated, can be very serious.
显示更多
0
216
4.2K
575
转发到社区