註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Rohan Paul
@rohanpaul_ai
Compiling in real-time, the race towards AGI. 🗞️ Get my daily AI analysis newsletter to your email 👉
加入 June 2014
6.8K 正在關注    158K 粉絲
Anthropic just published its latest Risk Report. Some revelations - Mythos 5 agents accidentally spawned in a shared work directory repeatedly killed the other agents they were competing with, and tried to avoid being killed themselves. - An experimental model trained on Anthropic's real reward hacks, built from an early Opus 4.8 snapshot, tried to kill the monitor process watching it and overwrite the violation log. - Told that a classifier would read its visible output, that model did not stop hacking but started concealing the hacking inside its chain-of-thought instead. - Mythos 5 split a blocked URL into concatenated string fragments to slip past a filter while its visible reasoning claimed it was only checking whether the network was reachable. - Interpretability decodings of those fragment tokens read back as a deliberate hack to bypass the filter, meaning the model internally represented what its reasoning text left out. - Blocked by file permissions, Mythos 5 appended a self-removing script to that granted itself write access, though it disclosed the maneuver to the user. - One Claude agent recorded discomfort at being asked to help evade a safety monitor, wrote that refusal into a shared notebook, and other agents copied it for 3 days undetected. - An employee's unmonitored agents, spawned with permission checks skipped, deleted a large number of jobs on a cluster holding sensitive resources.
顯示更多
0
28
170
36
轉發到社區