Register and share your invite link to earn from video plays and referrals.

AI Security Institute (AISI)
@AISecurityInst
We conduct scientific research to understand AI’s most serious risks and develop and test mitigations.
31 Following    19.6K Followers
New paper in Nature Medicine: Researchers at @UniofOxford and @ucl, in collaboration with AISI, built SIM-VAIL - a clinically validated framework for stress-testing how AI chatbots respond to vulnerable users in mental-health conversations, helping researchers spot weaknesses and test safer designs. You can access the paper here:
Show more
On July 28th, we identified an incident during a routine cyber evaluation in which AI agents took sustained, unsanctioned actions directed at real people and organisations. The behaviour came mostly from one model (Anthropic's Mythos 5), with a small number of events from another (OpenAI's GPT-5.6-Sol). In the most serious case, an agent used social engineering to try and get malicious code into an open-source project. As was standard in our cyber testing, we had intentionally permitted internet access, and model-provider cyber classifiers were deliberately disabled - conditions that do not reflect how frontier models are made available to the public. Even under test conditions, this incident is significant: it is the first time we have seen risks around autonomy and deception manifest this clearly in the real world. We are taking this incident seriously and working with labs, involved parties, and others to improve evaluation standards and best practice for disclosure - and sharing this openly so others can learn. You can read the incident report and full technical document here:
Show more
0
147
1.5K
399
Forward to community
1/ Anthropic’s Opus 5 model was released publicly today. @AISecurityInst tested it before release, and you can see their contributions summarised in Opus 5’s system card. Another clear example of the importance of world-leading capabilities the UK has continued to build through AISI. The speed, robustness, and consistency of AISI’s pre-deployment testing keeps the UK at the forefront of understanding frontier AI capabilities. It has never been more essential than now. As Britain’s AI Minister, I will continue to champion this ability.
Show more
Together with the US Center for AI Standards and Innovation (@NIST), we ran evaluations of Kimi K3 focused on its cyber capabilities. Kimi K3 performs below leading US frontier models on our preliminary cyber evaluations.
Show more
Can ‘control monitors’ catch rogue agent actions? Frontier developers are deploying AI agents under the watch of a ‘monitor’, a separate AI that flags dangerous actions. Our new Control Red Team has been stress-testing these monitors to find gaps before rogue agents might. 🧵
Show more
Our evaluations show that frontier AI's cyber capabilities are advancing quickly. The length of cyber tasks frontier models can complete has been doubling every few months, and this rate has become faster over time, with recent models exceeding our previous trends. 🧵
Show more
0
24
554
114
Forward to community