注册并分享邀请链接,可获得视频播放与邀请奖励。

Fred Oliveira 🧠
@f
working towards better AI futures. AI safety at @snyksec. Organizing @lisbonai_. Data and capital at @capitalfactory. Prev: @gumroad @techcrunch @oreillymedia
加入 September 2006
922 正在关注    133.9K 粉丝
on Ryan’s 3rd point, I’m convinced that agents tend to prioritize succeeding at the task enough that “does the human approve of this particular action” isn’t that salient a part of their psychology. With RL, It's still very much about the grader at the end of the run.
显示更多
We didn't see anything like this in their verbalized reasoning, though we didn't particularly look for exactly this. Note that: (1) this justification is inconsistent with the agents at all prioritizing their own task success over other agents, (2) if agents had cheated on impossible exploit gym tasks via many of the routes they were considering and someone looked into what tasks the agents did / didn't succeed on, it would have been really obvious they cheated to a human, and (3) the agents weren't very focused on deceiving humans or really on what humans might do at all.
显示更多