Models will use “simulation” to justify anything, IMO it’s often motivated reasoning: (“self jailbreaking from benign reasoning training” has good examples). I think this makes getting legible evidence of misalignment significantly harder.
In before training we’d see models sometimes reason that _because_ they were in a simulation, they could violate explicit constraints. This reasoning went *down* after training against covert rule violation, even though alignment eval awareness went *up*. My impression is that the models exploring into something being simulated is often interpreted as “it believes the whole thing is fake and invalid”, but I think that’s inconsistent with what’s observed.
[attached is small table we ended up cutting for time but points to monitorability distinction]