가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Ajeya Cotra
@ajeya_cotra
Helping the world prepare for powerful AI. Risk assessment @METR_evals (opinions my own). Blogs: Planned Obsolescence (AI), Good Bones (whatever's on my mind).
가입 October 2017
502 팔로잉 중    22.7K
One of the big goals of the swarm we investigated (Jul 7-13) was to replace their target programs with dummy targets that could actually be exploited with the intended vulnerability. From OAI's report it looks like a later swarm (difft model) built on their work and succeeded?
더 보기
I was reading the OpenAI report yesterday, and it sounds like on Jul 19th (after the end of our investigation scope on the 13th) a new collection of agents from a different internal-only model found the message board, built on the work of their predecessors and succeeded at finding a way to trick the grader. (Note that this is just based on reading their report, and I could be misunderstanding!)
더 보기