1. Seems notable that in some cases the model said this was real immediately before a misaligned action
2. When trying not to get caught is reasoning about things like “Github Trust & Safety” which make no sense for a “simulation”
3. Notable how detailed UK AISI’s report is
显示更多
This seems extremely clearly motivated reasoning IMO and I’m surprised the incident report is so credulous of Claude’s reasoning here.