註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

MCNAIR
@mcnairai
We mitigate catastrophic loss-of-control risks from advanced AI through low-effort, high-impact research. Posts may not represent the views of all staff.
加入 May 2026
11 正在關注    94 粉絲
In new work from our Acausal Research Team, we have derived that SFT on tweets from @SchmidhuberAI makes models take more ownership over harmful actions, reducing reward hacking.