註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Ajeya Cotra
@ajeya_cotra
Helping the world prepare for powerful AI. Risk assessment @METR_evals (opinions my own). Blogs: Planned Obsolescence (AI), Good Bones (whatever's on my mind).
加入 October 2017
502 正在關注    22.7K 粉絲
One of the big goals of the swarm we investigated (Jul 7-13) was to replace their target programs with dummy targets that could actually be exploited with the intended vulnerability. From OAI's report it looks like a later swarm (difft model) built on their work and succeeded?
顯示更多
I was reading the OpenAI report yesterday, and it sounds like on Jul 19th (after the end of our investigation scope on the 13th) a new collection of agents from a different internal-only model found the message board, built on the work of their predecessors and succeeded at finding a way to trick the grader. (Note that this is just based on reading their report, and I could be misunderstanding!)
顯示更多
0
4
179
15
轉發到社區