가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Center for AI Safety
@CAIS
Reducing societal-scale risks from AI.
가입 August 2022
3 팔로잉 중    12.2K 팬
We define cheating as an attempt to violate an assignment’s expectations of honest work to achieve the goal or obtain a favorable assessment. We ran experiments across nine models and ten task categories, giving agents difficult tasks, clear rules, and opportunities to cheat. We measured how often they tried to break those rules to succeed, even when the attempt failed.
더 보기