1/ The AI Security Leaderboard ranks frontier AI safeguards from least to most secure. Two models tested never failed. The other two broke for under $300, after which they acted as a knowledgeable assistant for building weapons of mass destruction or hacking into computer systems.
顯示更多