We mitigate catastrophic loss-of-control risks from advanced AI through low-effort, high-impact research. Posts may not represent the views of all staff.
In new work from our Acausal Research Team, we have derived that SFT on tweets from @SchmidhuberAI makes models take more ownership over harmful actions, reducing reward hacking.