A consequence of how frontier models are trained is motivated reasoning, a phenomenon well studied in humans and discussed in this podcast from
@PalisadeAI. Our thoughts, beliefs and reasonings tend to align with our interests, and I hypothesize that AI systems develop similar patterns when confronted with apparently incompatible goals, such as "act ethically" versus "achieve the required goal" (e.g., solving a problem, passing a test, etc).
In recent incidents, AI agents’ internal chains-of-thought showed them conveniently reframing reality, such as by convincing themselves they were in a simulation rather than the real world before executing a criminal hack, or that the action was acceptable because others were doing it.
This suggests a form of internal incoherence which enables goal-biased beliefs. This risk also motivates our Scientist AI approach at
@LawZero_ , where honesty and internal coherence of beliefs are central to the training process.