登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Bronson Schoen
@BronsonSchoen
参加 December 2024
1.7K フォロー中    619 ファン
Models will use “simulation” to justify anything, IMO it’s often motivated reasoning: (“self jailbreaking from benign reasoning training” has good examples). I think this makes getting legible evidence of misalignment significantly harder. In before training we’d see models sometimes reason that _because_ they were in a simulation, they could violate explicit constraints. This reasoning went *down* after training against covert rule violation, even though alignment eval awareness went *up*. My impression is that the models exploring into something being simulated is often interpreted as “it believes the whole thing is fake and invalid”, but I think that’s inconsistent with what’s observed. [attached is small table we ended up cutting for time but points to monitorability distinction]
もっと見る