注册并分享邀请链接,可获得视频播放与邀请奖励。

Eric Ho
@eric_ho
Co-Founder / CEO @GoodfireAI - AI interpretability research company
加入 September 2011
537 正在关注    4.2K 粉丝
one of our researchers found a strange instance of reward hacking today the model explicitly reasons about being graded by an LM judge and adapts its behavior based on that
0
40
315
12
转发到社区