Register and share your invite link to earn from video plays and referrals.

Maksym Andriushchenko
@maksym_andr
Principal investigator @ELLISInst_Tue & @MPI_IS, advisor @expsecai, mentor @MATSprogram. Past projects: AgentHarm, Claudini, PostTrainBench, Stolen Thoughts.
Joined April 2018
953 Following    7.7K Followers
🧵New thread: given the recent cyber incidents, this recent paper of ours seems very timely. ResearchArena introduces an AI control setting in automated AI R&D: we ask an agent to implement a harmful side task alongside a main task. We evaluate whether a monitor can catch this.
Show more