1/ The AI Security Leaderboard ranks frontier AI safeguards from least to most secure. Two models tested never failed. The other two broke for under $300, after which they acted as a knowledgeable assistant for building weapons of mass destruction or hacking into computer systems.
FAR AI is an independent non-profit publishing open research on how to make advanced AI systems safe and reliable. Our work is freely available to the entire research community.