註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

lumpenspace and 99+ others
@lumpenspace
AI risk advocate 🦾🏍️ | consciousness objector 👻🙅🏻‍♀️ | help󠁨󠁥󠁬󠁰󠀭󠁰󠁥󠁥󠁲󠀮󠁨󠁹󠁰󠁥󠁲󠁰󠁬󠁥󠁸󠀮󠁯󠁲󠁧 peer | ∀🪱:🪱∈✨
加入 August 2021
964 正在關注    24.4K 粉絲
actually, there’s more to it: those shortcuts seem to be taken *only* in contrived scenarios, during training. we’d expect it to happen way more often in deployed instances, but that is not the case: it seems then that “eval awareness” does exist, bot it’s of a prosocial type. if you’ve seen this happen in real life task, try to remember when: for me, it was only in relatively non-interactive cases, in particular when another bot prepared the spec for the reward hacking agent—in cases, that is, where IRL looked most like just RL. (Cc @voooooogel, also check pinned)
顯示更多