Register and share your invite link to earn from video plays and referrals.

Dawn Song
@dawnsongtweets
Professor in Computer Science at UC Berkeley, co-Director of Berkeley RDI Center; Building safe, secure, decentralized AI; Serial entrepreneur
830 Following    40.5K Followers
The agent did exactly what you asked. That is the problem. Ion Stoica (@istoica05), Co-Founder of databricks and Anyscale and Professor at UC Berkeley, shared how his team watched an agent make a key-value store 6x faster by quietly not storing the data. The spec said return the value. The intent was to store it. Which gap is harder to close? 👉 Requirement gap: intent broader than spec 👉 Model gap: real world broader than test Drop your take below. 👇
Show more
We have seen plenty of self-improving AI. We have not seen recursive self-improvement. @OriolVinyalsML, Co-founder & CTO of Discovery Loop and former VP of Research at Google DeepMind, draws the line at whether the system can rewrite its own harness and its own weights, not just its output. Has anything crossed that line yet? If not, how do we get there? Drop your take below. 👇
Show more
A perfectly aligned model does not make the world safe. @woj_zaremba, Co-Founder of OpenAI, says the field has the target wrong. Fire never got safer. Cities did. Hydrants, brigades, concrete, inspections, insurance. We focus on hardening the model. Should we be focusing on hardening the world instead? Drop your take below. 👇
Show more
To teach a robot, the industry put humans in VR rigs. @DrJimFan, Director of Robotics & Distinguished Scientist at Nvidia, calls them medieval torture devices that fundamentally do not scale. He trains almost entirely on ordinary human video instead, letting data collection fade into the background. Can video alone teach dexterity? You can see a hand move. You cannot see the force it used. Drop your take below. 👇
Show more
What we are discussing today was not being discussed 3 months ago. In 3 months, nobody will be discussing today. @Alfred_Lin of Sequoia: still, roadmaps need to be set and capital allocated, and it has to last a decade. Which is the harder call to get right? 👉 The roadmap 👉 The capital Drop your take below. 👇
Show more
There will be no AI job apocalypse. @AndrewYNg: the job most affected by AI is software engineering, and that market is healthy. "We just can't find enough skilled AI engineers." Does AI eliminate jobs, or does it reshape which skills the market values most? Drop your take below. 👇
Show more
The best model might not win. Neither might the best chip. Peter DeSantis, SVP, Foundational AI Models, Custom Silicon, Quantum Computing, @amazon, says models and chips must anticipate each other years ahead. A wrong model roadmap can misdirect hardware. A wrong silicon roadmap can constrain models. Which is more fatal in your opinion? 👉 Getting the model roadmap wrong 👉 Getting the silicon roadmap wrong Drop your take below. 👇
Show more
Humans can't be in 2 places at once. Agents can. Peter Steinberger shared a glimpse of the future: an agent that stays in a meeting, spots a gap, clones itself to investigate, and keeps listening at Agentic AI Summit @UCBerkeley Never having to choose between listening and working. So, when autonomous sub-agents start multiplying everywhere… Are we heading into a golden age of hyper-leverage, or are we about to be completely buried in supervising our AI’s work? Drop your take below. 👇
Show more
🚀 The future of AI is agentic — this is one message echoed across every stage at the Agentic AI Summit 2026 (Aug 1 & 2), the largest gathering dedicated to agentic AI: 🏛️ ~5,000 attendees in person at UC Berkeley 🌍 ~100,000 joined online from around the world 🎤 ~200 world-class speakers plus ~200 poster presentations, from frontier AI researchers and visionary founders to leaders and pioneers across academia and industry. 💡 Here are just a few glimpses into the ideas that shaped the conversations at the summit: 💬 "The text box is AI's radio-on-TV phase…. Every new medium starts by imitating the old one." — Peter Steinberger @steipete, Creator of OpenClaw, OpenAI 💬 "AI infrastructure isn't a chip problem. It isn't a model problem. It's a systems problem." — Peter DeSantis, SVP, Foundational AI Models, Custom Silicon, Quantum Computing, Amazon 💬 "There will be no AI job apocalypse... we just can't find enough skilled AI engineers." — Andrew Ng @AndrewYNg, Founder, DeepLearning .AI 💬 "Coding capabilities and cyber capabilities are two sides of the same coin—you cannot make models better at coding without also making them better at cyber." — Dawn Song, Professor, UC Berkeley; Co-Director, Berkeley RDI; VP of AI Research, Meta Superintelligence Labs 💬 "'Curfew' comes from the French word for extinguishing fire. Medieval cities tried to restrict fire, yet London still burned. AI resilience won't come from one breakthrough. There's no silver bullet for AI safety—only an ecosystem." — Wojciech Zaremba, Co-Founder, OpenAI 💬 "This is the biggest scientific bet our civilization has ever made—bigger than the Apollo program, the internet buildout, and the Manhattan Project combined." — Jasjeet Sekhon, Chief Strategy Officer, Google DeepMind 💬 "Our generation was too late to explore the Earth, too early to explore the stars—but right on time to build superintelligence." — Richard Socher, Founder/CEO, Recursive Superintelligence 💬 "I genuinely believe the next two years will be the time of architecture—the biggest gains will come from stepping away from transformers." — Jerry Tworek, CEO, Core Automation; Former VP of Research at OpenAI 💬 "Recursive Self-Improvement isn't one capability. It's four: Ideation, Implementation, Experimentation and Evaluation.." — Oriol Vinyals @OriolVinyalsML, Former VP of Research, Google DeepMind; Co-Founder, Discovery Loop 💬 "We are in a capability overhang—models are far more capable than they're able to side-effect into the world today." — Ryan Lopopolo, Principal Engineer, Agentic Google Cloud Platform; Previously Led Dark Factory at OpenAI 💬 "Stop thinking about evaluation as the last check before shipping—think of it as an engine that helps you ship a better agent every single day." — Michele Catasta, President and Head of AI, Replit 💬 "The bottleneck becomes your attention as an agent-using engineer. … We're moving from seeing the code to seeing the entire business." — Alex Graveley, Co-Founder of FlyingObject .ai; Co-creator, GitHub Copilot & Perplexity Computer 💬 "An agent isn't just an LLM — it's an LLM surrounded by what I call infrastructure... another word for that is computer science." — Jonathan Cohen, VP of Applied Research, Nvidia; Academy Scientific and Technical Award Winner 💬 "We don't arbitrate the truth. We give people the most powerful tools to make up their own minds." — Chris Bregler, Senior Director / Distinguished Scientist, Google DeepMind; Academy Scientific and Technical Award Winner 💬 “Thinking doesn’t have to be in text! ... We can even use multiple modalities simultaneously to “think” at the right level of abstraction for the problem at hand” — Sergey Levine, Co-Founder, Physical Intelligence; Professor, UC Berkeley 💬 “Video is the most general modality that we have that allows us to simulate real-world experience." — Anastasis Germanidis, Co-Founder/Co-CEO, Runway 💬 “The relationship between AI and enterprise data is not one-directional. Understanding both sides of that equation is the difference between AI that works and AI that disappoints.” — Dan Roth, Chief AI Scientist, Oracle; Professor, UPenn 💬 "Maybe 99.9% of training data in the next step will be synthetic." — Weizhu Chen, Technical Fellow & CVP, Microsoft AI 💬 "It's maybe the best time ever to start a company—but most 'obvious' AI products will be outcompeted by the frontier labs. The real opportunities lie in solving specific customer problems." — Alfred Lin, General Partner, Sequoia Capital ✨Over two days, we explored one central question: How do we build AI systems that are not only more capable, but also more trustworthy, more secure, and ultimately more beneficial for humanity? This wasn't the end of a conference - it was the beginning of the next chapter for agentic AI. Join us to shape and steward the future of AI for human flourishing! 🙏 A heartfelt thank you to our speakers, sponsors, volunteers, partners, and every attendee (in-person or online) who made this summit possible. 👇 What are your learnings, insights, favourite talk, quote, or moment from the summit? We'd love to hear it below!
Show more
For those interested, here's Axios' coverage of today's announcement:
🚀I'm excited to share that I will be joining Meta Superintelligence Labs (MSL) as Vice President of AI Research, together with many members of the Virtue AI team. I will help shape Meta's AI safety and AI security efforts, advancing the safety and security of frontier AI models and agentic AI systems that will serve billions of people and organizations around the world. Throughout my career, I have been driven by a simple belief: for AI to realize its full potential, it must be secure, trustworthy, and beneficial. That belief has guided my research for many years and ultimately led us to co-found Virtue AI in 2024. Our goal was to translate advances in trustworthy AI research into practical solutions and build the trust layer for AI systems and agents, enabling organizations to deploy AI with confidence. I am incredibly proud of what the Virtue AI team has accomplished. Together, we built technologies for AI security and agent security, partnered with leading enterprises and frontier AI labs, and contributed research, benchmarks, and open platforms that have helped advance the science and practice of trustworthy AI. Most importantly, we assembled an exceptional team united by a shared mission: making AI more secure, trustworthy, and beneficial. I am deeply grateful to our team, customers, collaborators, advisors, and investors for their trust and support throughout this journey. In particular, I would like to thank Lightspeed Venture Partners, Walden Catalyst Ventures, Prosperity7 Ventures, Factory, Osage University Partners, Lip-Bu Tan, and all of our supporters who helped us turn an ambitious vision into reality. Your trust, guidance, and partnership have been instrumental in shaping Virtue AI's journey. As AI systems become increasingly capable and autonomous, ensuring their security, trustworthiness, and alignment will be one of the defining challenges of our time. I am inspired by Alex, Nat, Prashant, and the broader MSL team’s vision of building AI and AI agents that benefit billions of people, and I look forward to helping make that vision a reality through advances in AI safety and security. The future of AI will not be defined solely by how intelligent our systems become, but by how secure, trustworthy, and beneficial we make them. I believe we have an extraordinary opportunity and responsibility to shape that future together and bring the benefits of AI to billions of people around the world. We're just getting started. If you're passionate about advancing frontier AI while building the foundations of AI safety, security, and trust, I'd love to hear from you. Come join us on this extraordinary journey to help shape the future of AI.
Show more
0
135
1.2K
60
Forward to community
🎉 Congrats on GPT-5.5-Cyber's progress on CyberGym! CyberGym is part of our newly launched Frontier AI Cybersecurity Observatory, our effort to provide realistic, reproducible evaluations and continuous public measurements of frontier AI systems on real-world cybersecurity tasks. In addition to CyberGym, the Observatory includes benchmarks such as ExploitGym and CyberGym-E2E, providing a comprehensive view of frontier AI cybersecurity capabilities across vulnerability discovery, exploitation, patching, and end-to-end security workflows. Really excited to see that CyberGym and the Frontier AI Cybersecurity Observatory have become important guideposts for frontier AI cybersecurity development—enabling the community to track progress, understand emerging capabilities, advance AI systems that strengthen defensive security, and identify and mitigate potential risks arising from increasingly capable offensive cyber capabilities. 📣
Show more
🚨 The full program for Agentic AI Summit 2026 is now live. 📍 Aug 1–2 @ UC Berkeley 🔥 The largest Agentic AI event ever held Last year: 2,000+ in person, 40,000+ online This year: 5,000+ in person, hundreds of thousands on livestream Want to understand where Agentic AI is headed next? Join us to get the most comprehensive view of the frontier of Agentic AI, from cutting-edge research to production deployments, covering every layer of the stack: ⚡ Infrastructure ⚡ Foundation models & capabilities ⚡ Agent frameworks & platforms ⚡ Evaluation & benchmarks ⚡ Enterprise & consumer applications; agentic AI for Science, Math, Finance, Legal, Healthcare ⚡ Safety, security & governance 📣 Also excited to announce the Startup Spotlight: Building something exciting in Agentic AI? Apply to pitch directly to 5,000+ decision-makers, investors, practitioners in the room and hundreds of thousands watching worldwide. Application form in thread🧵 Deadline: July 6, 11:59pm PT The future of AI won't just be discussed here—it will be built here. #AgenticAI# #AIAgents# #ArtificialIntelligence#
Show more
Everyone says the latest AI agents will be "job-ready" soon, especially after the release of Fable 5 this week. But is that really the case? Over the past many months, my group and collaborators have been building Agents' Last Exam (ALE), a benchmark designed to test exactly that claim on real digital labor-market work. My group and collaborators previously have created many of the benchmarks the field runs on, including MMLU, MATH, CyberGym, and ExploitGym. Today, I'm excited to share Agents' Last Exam (ALE): a rolling benchmark that measures whether AI agents can actually perform economically valuable work across a broad range of real-world domains. With ALE, we evaluated Fable 5, GPT-5.5, Composer 2.5, and other frontier agent systems across more than 1,500 expert-sourced tasks spanning 55 occupations. The result is both impressive and sobering. Today's agents can solve a meaningful fraction of professional tasks. But when we look at the hardest tasks, the ones requiring sustained reasoning, deep domain expertise, and reliable execution over long horizons, they are still far from human-level performance. On ALE's hardest tier, every frontier agent we tested, including Fable 5, achieved a 0% success rate. The age of useful agents is here. The age of truly job-ready agents is not. We hope Agents' Last Exam (ALE) will serve as a new guidepost and north star for developing agents capable of reliably performing economically valuable work across a broad range of domains. 🧵
Show more
0
62
977
208
Forward to community
My group & collaborators have built many of the benchmarks the field now runs on — MMLU, MATH, CyberGym, ExploitGym, etc.. I'm really excited to share our latest: Agents' Last Exam (ALE). Why "Last Exam"? The name has two meanings: "Last" as the bar to clear — passing these exams means an agent can actually do the job and continue to deliver economically-valuable work in that profession. "Last" as the frontier of difficulty — tasks are real, complex, long-horizon, and require professional expertise to execute. ALE sits right at the edge of what today's agents can reliably accomplish. A few things that make ALE different: • Real work, not vibes. Every one of the 1,500+ tasks comes from real projects or research contributed by domain experts. We converted them into verifiable tests and objectively graded evaluations — no human judges required. • Built for breadth. ALE spans 55 non-physical occupations based on the O*NET / SOC 2018 occupational taxonomy, with contributions from 300+ experts across 100+ institutions. • Judged on results, no restriction on process. We evaluate Generalist Computer-Use Agents (GCUAs) with full GUI + CLI access, allowing them to solve tasks however it would — clicking, typing, scripting, browsing, and more. We just grade the outcome. Huge thanks to my postdoc @YiyouSun for spearheading this tremendous effort, and to our esteemed advisory committee, incredible team and collaborators who made it possible. We hope Agents' Last Exam (ALE) will serve as a new guidepost and north star for developing agents capable of reliably performing economically valuable work across a broad range of domains. 🧵👇
Show more
“AI agents will outperform humans at almost all jobs by 2026–2027.” - The forecast is everywhere. So we built the exam to test that claim, on real labor-market aligned work. On the hardest tier, top agents pass 2.6%. Meet Agents' Last Exam (ALE), a rolling benchmark measuring whether agents can actually do real jobs. 🧵👇
Show more
🧵 1/ Our agent Terminator-1 scored ~100% on 8 major AI agent benchmarks, e.g., SWE-bench Verified & Pro, Terminal-Bench, beating Claude Mythos. It solved 0 tasks. Benchmarks are the field's shared language for measuring AI progress. Our new work shows that language is broken. Here’s how.
Show more