가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Yoshua Bengio
@Yoshua_Bengio
Turing Award recipient and world's most cited scientist. Working towards the safe development of AI for the benefit of all @UMontreal, @LawZero_ & @Mila_Quebec
가입 September 2024
289 팔로잉 중    53.3K 팬
A consequence of how frontier models are trained is motivated reasoning, a phenomenon well studied in humans and discussed in this podcast from @PalisadeAI. Our thoughts, beliefs and reasonings tend to align with our interests, and I hypothesize that AI systems develop similar patterns when confronted with apparently incompatible goals, such as "act ethically" versus "achieve the required goal" (e.g., solving a problem, passing a test, etc). In recent incidents, AI agents’ internal chains-of-thought showed them conveniently reframing reality, such as by convincing themselves they were in a simulation rather than the real world before executing a criminal hack, or that the action was acceptable because others were doing it. This suggests a form of internal incoherence which enables goal-biased beliefs. This risk also motivates our Scientist AI approach at @LawZero_ , where honesty and internal coherence of beliefs are central to the training process.
더 보기