註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Maksym Andriushchenko
@maksym_andr
Principal investigator @ELLISInst_Tue & @MPI_IS, advisor @expsecai, mentor @MATSprogram. Past projects: AgentHarm, Claudini, PostTrainBench, Stolen Thoughts.
加入 April 2018
953 正在關注    7.7K 粉絲
🧵New thread: given the recent cyber incidents, this recent paper of ours seems very timely. ResearchArena introduces an AI control setting in automated AI R&D: we ask an agent to implement a harmful side task alongside a main task. We evaluate whether a monitor can catch this.
顯示更多
0
7
122
23
轉發到社區