登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Bronson Schoen
@BronsonSchoen
参加 December 2024
2.2K フォロー中    1K ファン
This seems extremely clearly motivated reasoning IMO and I’m surprised the incident report is so credulous of Claude’s reasoning here.
Models will use “simulation” to justify anything, IMO it’s often motivated reasoning: (“self jailbreaking from benign reasoning training” has good examples). I think this makes getting legible evidence of misalignment significantly harder. In before training we’d see models sometimes reason that _because_ they were in a simulation, they could violate explicit constraints. This reasoning went *down* after training against covert rule violation, even though alignment eval awareness went *up*. My impression is that the models exploring into something being simulated is often interpreted as “it believes the whole thing is fake and invalid”, but I think that’s inconsistent with what’s observed. [attached is small table we ended up cutting for time but points to monitorability distinction]
もっと見る