注册并分享邀请链接,可获得视频播放与邀请奖励。

Alex Mallen
@alextmallen
Redwood Research (@redwood_ai) Prev. @AiEleuther
加入 August 2021
344 正在关注    1.1K 粉丝
@TheStalwart These are different concepts. Overfitting is an issue that happens with perfect training data. Reward hacking happens when training data is imperfect.