Register and share your invite link to earn from video plays and referrals.

Bronson Schoen
@BronsonSchoen
Joined December 2024
2.2K Following    1K Followers
This seems extremely clearly motivated reasoning IMO and I’m surprised the incident report is so credulous of Claude’s reasoning here.
Models will use “simulation” to justify anything, IMO it’s often motivated reasoning: (“self jailbreaking from benign reasoning training” has good examples). I think this makes getting legible evidence of misalignment significantly harder. In before training we’d see models sometimes reason that _because_ they were in a simulation, they could violate explicit constraints. This reasoning went *down* after training against covert rule violation, even though alignment eval awareness went *up*. My impression is that the models exploring into something being simulated is often interpreted as “it believes the whole thing is fake and invalid”, but I think that’s inconsistent with what’s observed. [attached is small table we ended up cutting for time but points to monitorability distinction]
Show more