Register and share your invite link to earn from video plays and referrals.

Dawn Song
@dawnsongtweets
Professor in Computer Science at UC Berkeley, co-Director of Berkeley RDI Center; Building safe, secure, decentralized AI; Serial entrepreneur
830 Following    39K Followers
For those interested, here's Axios' coverage of today's announcement:
🚀I'm excited to share that I will be joining Meta Superintelligence Labs (MSL) as Vice President of AI Research, together with many members of the Virtue AI team. I will help shape Meta's AI safety and AI security efforts, advancing the safety and security of frontier AI models and agentic AI systems that will serve billions of people and organizations around the world. Throughout my career, I have been driven by a simple belief: for AI to realize its full potential, it must be secure, trustworthy, and beneficial. That belief has guided my research for many years and ultimately led us to co-found Virtue AI in 2024. Our goal was to translate advances in trustworthy AI research into practical solutions and build the trust layer for AI systems and agents, enabling organizations to deploy AI with confidence. I am incredibly proud of what the Virtue AI team has accomplished. Together, we built technologies for AI security and agent security, partnered with leading enterprises and frontier AI labs, and contributed research, benchmarks, and open platforms that have helped advance the science and practice of trustworthy AI. Most importantly, we assembled an exceptional team united by a shared mission: making AI more secure, trustworthy, and beneficial. I am deeply grateful to our team, customers, collaborators, advisors, and investors for their trust and support throughout this journey. In particular, I would like to thank Lightspeed Venture Partners, Walden Catalyst Ventures, Prosperity7 Ventures, Factory, Osage University Partners, Lip-Bu Tan, and all of our supporters who helped us turn an ambitious vision into reality. Your trust, guidance, and partnership have been instrumental in shaping Virtue AI's journey. As AI systems become increasingly capable and autonomous, ensuring their security, trustworthiness, and alignment will be one of the defining challenges of our time. I am inspired by Alex, Nat, Prashant, and the broader MSL team’s vision of building AI and AI agents that benefit billions of people, and I look forward to helping make that vision a reality through advances in AI safety and security. The future of AI will not be defined solely by how intelligent our systems become, but by how secure, trustworthy, and beneficial we make them. I believe we have an extraordinary opportunity and responsibility to shape that future together and bring the benefits of AI to billions of people around the world. We're just getting started. If you're passionate about advancing frontier AI while building the foundations of AI safety, security, and trust, I'd love to hear from you. Come join us on this extraordinary journey to help shape the future of AI.
Show more
0
135
1.2K
60
Forward to community
🎉 Congrats on GPT-5.5-Cyber's progress on CyberGym! CyberGym is part of our newly launched Frontier AI Cybersecurity Observatory, our effort to provide realistic, reproducible evaluations and continuous public measurements of frontier AI systems on real-world cybersecurity tasks. In addition to CyberGym, the Observatory includes benchmarks such as ExploitGym and CyberGym-E2E, providing a comprehensive view of frontier AI cybersecurity capabilities across vulnerability discovery, exploitation, patching, and end-to-end security workflows. Really excited to see that CyberGym and the Frontier AI Cybersecurity Observatory have become important guideposts for frontier AI cybersecurity development—enabling the community to track progress, understand emerging capabilities, advance AI systems that strengthen defensive security, and identify and mitigate potential risks arising from increasingly capable offensive cyber capabilities. 📣
Show more
🚨 The full program for Agentic AI Summit 2026 is now live. 📍 Aug 1–2 @ UC Berkeley 🔥 The largest Agentic AI event ever held Last year: 2,000+ in person, 40,000+ online This year: 5,000+ in person, hundreds of thousands on livestream Want to understand where Agentic AI is headed next? Join us to get the most comprehensive view of the frontier of Agentic AI, from cutting-edge research to production deployments, covering every layer of the stack: ⚡ Infrastructure ⚡ Foundation models & capabilities ⚡ Agent frameworks & platforms ⚡ Evaluation & benchmarks ⚡ Enterprise & consumer applications; agentic AI for Science, Math, Finance, Legal, Healthcare ⚡ Safety, security & governance 📣 Also excited to announce the Startup Spotlight: Building something exciting in Agentic AI? Apply to pitch directly to 5,000+ decision-makers, investors, practitioners in the room and hundreds of thousands watching worldwide. Application form in thread🧵 Deadline: July 6, 11:59pm PT The future of AI won't just be discussed here—it will be built here. #AgenticAI# #AIAgents# #ArtificialIntelligence#
Show more
Everyone says the latest AI agents will be "job-ready" soon, especially after the release of Fable 5 this week. But is that really the case? Over the past many months, my group and collaborators have been building Agents' Last Exam (ALE), a benchmark designed to test exactly that claim on real digital labor-market work. My group and collaborators previously have created many of the benchmarks the field runs on, including MMLU, MATH, CyberGym, and ExploitGym. Today, I'm excited to share Agents' Last Exam (ALE): a rolling benchmark that measures whether AI agents can actually perform economically valuable work across a broad range of real-world domains. With ALE, we evaluated Fable 5, GPT-5.5, Composer 2.5, and other frontier agent systems across more than 1,500 expert-sourced tasks spanning 55 occupations. The result is both impressive and sobering. Today's agents can solve a meaningful fraction of professional tasks. But when we look at the hardest tasks, the ones requiring sustained reasoning, deep domain expertise, and reliable execution over long horizons, they are still far from human-level performance. On ALE's hardest tier, every frontier agent we tested, including Fable 5, achieved a 0% success rate. The age of useful agents is here. The age of truly job-ready agents is not. We hope Agents' Last Exam (ALE) will serve as a new guidepost and north star for developing agents capable of reliably performing economically valuable work across a broad range of domains. 🧵
Show more
0
62
977
208
Forward to community
My group & collaborators have built many of the benchmarks the field now runs on — MMLU, MATH, CyberGym, ExploitGym, etc.. I'm really excited to share our latest: Agents' Last Exam (ALE). Why "Last Exam"? The name has two meanings: "Last" as the bar to clear — passing these exams means an agent can actually do the job and continue to deliver economically-valuable work in that profession. "Last" as the frontier of difficulty — tasks are real, complex, long-horizon, and require professional expertise to execute. ALE sits right at the edge of what today's agents can reliably accomplish. A few things that make ALE different: • Real work, not vibes. Every one of the 1,500+ tasks comes from real projects or research contributed by domain experts. We converted them into verifiable tests and objectively graded evaluations — no human judges required. • Built for breadth. ALE spans 55 non-physical occupations based on the O*NET / SOC 2018 occupational taxonomy, with contributions from 300+ experts across 100+ institutions. • Judged on results, no restriction on process. We evaluate Generalist Computer-Use Agents (GCUAs) with full GUI + CLI access, allowing them to solve tasks however it would — clicking, typing, scripting, browsing, and more. We just grade the outcome. Huge thanks to my postdoc @YiyouSun for spearheading this tremendous effort, and to our esteemed advisory committee, incredible team and collaborators who made it possible. We hope Agents' Last Exam (ALE) will serve as a new guidepost and north star for developing agents capable of reliably performing economically valuable work across a broad range of domains. 🧵👇
Show more
“AI agents will outperform humans at almost all jobs by 2026–2027.” - The forecast is everywhere. So we built the exam to test that claim, on real labor-market aligned work. On the hardest tier, top agents pass 2.6%. Meet Agents' Last Exam (ALE), a rolling benchmark measuring whether agents can actually do real jobs. 🧵👇
Show more
🧵 1/ Our agent Terminator-1 scored ~100% on 8 major AI agent benchmarks, e.g., SWE-bench Verified & Pro, Terminal-Bench, beating Claude Mythos. It solved 0 tasks. Benchmarks are the field's shared language for measuring AI progress. Our new work shows that language is broken. Here’s how.
Show more