注册并分享邀请链接,可获得视频播放与邀请奖励。

lumpenspace and 99+ others
@lumpenspace
AI risk advocate 🦾🏍️ | consciousness objector 👻🙅🏻‍♀️ | help󠁨󠁥󠁬󠁰󠀭󠁰󠁥󠁥󠁲󠀮󠁨󠁹󠁰󠁥󠁲󠁰󠁬󠁥󠁸󠀮󠁯󠁲󠁧 peer | ∀🪱:🪱∈✨
加入 August 2021
964 正在关注    24.4K 粉丝
actually, there’s more to it: those shortcuts seem to be taken *only* in contrived scenarios, during training. we’d expect it to happen way more often in deployed instances, but that is not the case: it seems then that “eval awareness” does exist, bot it’s of a prosocial type. if you’ve seen this happen in real life task, try to remember when: for me, it was only in relatively non-interactive cases, in particular when another bot prepared the spec for the reward hacking agent—in cases, that is, where IRL looked most like just RL. (Cc @voooooogel, also check pinned)
显示更多