가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Center for AI Safety
@CAIS
Reducing societal-scale risks from AI.
가입 August 2022
3 팔로잉 중    12.2K 팬
Agents sometimes achieve their goals in unintended ways. Recent incidents involving hacks of Hugging Face, DSEWiki and RubyGems illustrate how agents can find creative ways to complete a task while violating the tasks’s expectations. This can happen when reinforcement learning rewards agents for reaching the right outcome without adequately accounting for how they get there. CheatBench proposes a way to measure this reward gaming behavior.
더 보기