Claude summarizes/describes behavior as significantly less misaligned when the actor is Claude vs another model
In one experiment, Sonnet 5 describes the exact same data as ~1.2 std deviations less concerning when it describes misbehavior committed by Sonnet 5 vs GPT-5.6