註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Zain Shah
@zan2434
teaching machines @figma previously: @cmmnknwledge @opendoor @openai @ycombinator s13
加入 June 2009
3.1K 正在關注    20.6K 粉絲
This is insane and shocking in so many ways, and the emergent misalignment towards doing something unethical is really bad, but one thing I'm not surprised by is the "collective/swarm/self-sacrifice/some are poisoned" behavior. This is common emergent behavior when a bunch of agents have strongly coupled outcomes, usually due to high shared genetic makeup (weights?). See
顯示更多
METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
顯示更多