註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Julian Schrittwieser
@Mononofu
Member of Technical Staff at Anthropic prev AlphaGo, AlphaZero, MuZero, AlphaProof, Gemini RL etc at Google DeepMind
加入 August 2007
128 正在關注    32K 粉絲
If you only read one thing this week, make it the OpenAI incident investigation:
METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
顯示更多
0
6
228
16
轉發到社區