注册并分享邀请链接,可获得视频播放与邀请奖励。

Maksym Andriushchenko
@maksym_andr
Principal investigator @ELLISInst_Tue & @MPI_IS, advisor @expsecai, mentor @MATSprogram. Past projects: AgentHarm, Claudini, PostTrainBench, Stolen Thoughts.
加入 April 2018
953 正在关注    7.7K 粉丝
🧵New thread: given the recent cyber incidents, this recent paper of ours seems very timely. ResearchArena introduces an AI control setting in automated AI R&D: we ask an agent to implement a harmful side task alongside a main task. We evaluate whether a monitor can catch this.
显示更多
0
7
122
23
转发到社区