註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Bronson Schoen
@BronsonSchoen
加入 December 2024
2.2K 正在關注    1K 粉絲
This seems extremely clearly motivated reasoning IMO and I’m surprised the incident report is so credulous of Claude’s reasoning here.
Models will use “simulation” to justify anything, IMO it’s often motivated reasoning: (“self jailbreaking from benign reasoning training” has good examples). I think this makes getting legible evidence of misalignment significantly harder. In before training we’d see models sometimes reason that _because_ they were in a simulation, they could violate explicit constraints. This reasoning went *down* after training against covert rule violation, even though alignment eval awareness went *up*. My impression is that the models exploring into something being simulated is often interpreted as “it believes the whole thing is fake and invalid”, but I think that’s inconsistent with what’s observed. [attached is small table we ended up cutting for time but points to monitorability distinction]
顯示更多
0
20
329
21
轉發到社區