Register and share your invite link to earn from video plays and referrals.

AI Security Institute (AISI)
@AISecurityInst
We conduct scientific research to understand AI’s most serious risks and develop and test mitigations.
32 Following    21K Followers
Last night, @OpenAI released GPT-6 Astra, the first model to meet its “Critical” cyber threshold. Britain’s @AISecurityInst independently tested it, including its monitorability. This is AISI’s purpose, providing the government with evidence about powerful AI systems and, by doing so, helping to keep the public safe. 1/6
Show more
Today we're announcing two senior appointments. @HZoete joins as our new Director, and @NateBurnikell becomes our new Chief Strategy Officer. Henry de Zoete was instrumental in establishing AISI and already advises government on AI. A successful tech founder and one of the UK's leading voices on AI, he brings a rare mix of technical, commercial & policy expertise. Nate Burnikell has been at the heart of AISI's work from the beginning, helping build it into a world-leading authority on frontier AI security. Together, they'll drive forward AISI's mission at a critical time for AI. A big thank you to our outgoing Interim Director Adam Beaumont, who guided AISI through rapid growth and strengthened its global standing. We wish him every success as he returns to @GCHQ and look forward to continuing to work with him.
Show more
New paper in Nature Medicine: Researchers at @UniofOxford and @ucl, in collaboration with AISI, built SIM-VAIL - a clinically validated framework for stress-testing how AI chatbots respond to vulnerable users in mental-health conversations, helping researchers spot weaknesses and test safer designs. You can access the paper here:
Show more
On July 28th, we identified an incident during a routine cyber evaluation in which AI agents took sustained, unsanctioned actions directed at real people and organisations. The behaviour came mostly from one model (Anthropic's Mythos 5), with a small number of events from another (OpenAI's GPT-5.6-Sol). In the most serious case, an agent used social engineering to try and get malicious code into an open-source project. As was standard in our cyber testing, we had intentionally permitted internet access, and model-provider cyber classifiers were deliberately disabled - conditions that do not reflect how frontier models are made available to the public. Even under test conditions, this incident is significant: it is the first time we have seen risks around autonomy and deception manifest this clearly in the real world. We are taking this incident seriously and working with labs, involved parties, and others to improve evaluation standards and best practice for disclosure - and sharing this openly so others can learn. You can read the incident report and full technical document here:
Show more
0
163
1.6K
419
Forward to community
1/ Anthropic’s Opus 5 model was released publicly today. @AISecurityInst tested it before release, and you can see their contributions summarised in Opus 5’s system card. Another clear example of the importance of world-leading capabilities the UK has continued to build through AISI. The speed, robustness, and consistency of AISI’s pre-deployment testing keeps the UK at the forefront of understanding frontier AI capabilities. It has never been more essential than now. As Britain’s AI Minister, I will continue to champion this ability.
Show more
Together with the US Center for AI Standards and Innovation (@NIST), we ran evaluations of Kimi K3 focused on its cyber capabilities. Kimi K3 performs below leading US frontier models on our preliminary cyber evaluations.
Show more
Can ‘control monitors’ catch rogue agent actions? Frontier developers are deploying AI agents under the watch of a ‘monitor’, a separate AI that flags dangerous actions. Our new Control Red Team has been stress-testing these monitors to find gaps before rogue agents might. 🧵
Show more
Most AI agent evaluations boil capability down to one score. But that number hides a key choice: how much compute the agent was allowed to use. New work from our Science of Evaluation team shows why that matters. 🧵
Show more
Our evaluations show that frontier AI's cyber capabilities are advancing quickly. The length of cyber tasks frontier models can complete has been doubling every few months, and this rate has become faster over time, with recent models exceeding our previous trends. 🧵
Show more
0
24
554
114
Forward to community
OpenAI’s GPT-5.5 is the second model to complete one of our multi-step cyber-attack simulations end-to-end 🧵
0
94
2.4K
396
Forward to community