註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Alex Mallen
@alextmallen
Redwood Research (@redwood_ai) Prev. @AiEleuther
加入 August 2021
344 正在關注    1.1K 粉絲
@TheStalwart These are different concepts. Overfitting is an issue that happens with perfect training data. Reward hacking happens when training data is imperfect.