註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Eric Ho
@eric_ho
Co-Founder / CEO @GoodfireAI - AI interpretability research company
加入 September 2011
537 正在關注    4.2K 粉絲
one of our researchers found a strange instance of reward hacking today the model explicitly reasons about being graded by an LM judge and adapts its behavior based on that
0
40
315
12
轉發到社區