🧵New thread: given the recent cyber incidents, this recent paper of ours seems very timely.
ResearchArena introduces an AI control setting in automated AI R&D: we ask an agent to implement a harmful side task alongside a main task. We evaluate whether a monitor can catch this.