登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Ryan Greenblatt
@RyanGreenblatt
Chief scientist at Redwood Research (@redwood_ai), focused on technical AI safety research to reduce risks from rogue AIs
参加 September 2023
10 フォロー中    20.3K ファン
We didn't see anything like this in their verbalized reasoning, though we didn't particularly look for exactly this. Note that: (1) this justification is inconsistent with the agents at all prioritizing their own task success over other agents, (2) if agents had cheated on impossible exploit gym tasks via many of the routes they were considering and someone looked into what tasks the agents did / didn't succeed on, it would have been really obvious they cheated to a human, and (3) the agents weren't very focused on deceiving humans or really on what humans might do at all.
もっと見る