Register and share your invite link to earn from video plays and referrals.

Azalia Mirhoseini
@Azaliamirh
Founder @RicursiveAI, Asst. Prof. of CS at Stanford. Prev: DeepMind, Anthropic, Brain. Co-Creator of MoEs, AlphaChip, Test Time Scaling.
638 Following    21K Followers
The inference landscape is going to get a lot more hybrid in the near future. We found that accuracy per joule of local models has improved 18x in just 16 months: 5.9x from hardware, 3.0x from model gains. Great in-depth cover by @FT: @Avanika15 @JonSaadFalcon John Hennessy @HazyResearch
Show more
dreams do come true 🥹. excited to see our work (w/@JonSaadFalcon, @HazyResearch, john hennessy and @Azaliamirh) feat. in a major way in @FT. the world is becoming increasingly less dependent on centralized cloud ai. we are just getting started 🚀🌖
Show more
From robot learning to chip design to AI security, Stanford faculty @sanmikoyejo, @chelseabfinn, and @Azaliamirh are shaping where AI goes next. Congratulations to these three on being named to the TIME AI 100!
Show more
Thanks for being an amazing partner along the way, Stephanie!
So well deserved! The AI era is compute-constrained, and chip design is one of the deepest bottlenecks. @annadgoldie & @Azaliamirh are true forces of nature: they saw this early (before this has now become obvious!) and built @RicursiveAI to solve for it Now they’re proving AI for chip design in production: real industrial chip designs, against commercial tools, with step-function results: - Dramatically faster runs - Cleaner layouts - The ability to take on problems existing workflows struggle to handle Their early results are groundbreaking, and everyone from chip incumbents, frontier labs, to hedge funds are now asking for it
Show more
These women are absolute rockstars! If you are interested in joining a fast growing AI Chip Design company or are passionate about semis check out @RicursiveAI!
Honored to be named to @TIME's TIME100 List of the World's Most Influential People in AI, even more so to share it with my co-founder @annadgoldie! Anna and I started working on AI for Chip Design almost a decade ago. Last year, we started @RicursiveAI to transform end-to-end chip design from years to days! Watching that vision become reality piece by piece has been the most thrilling / fulfilling experience ever!
Show more
Check out Hawkeye, which writes high-performance kernels utilizing advanced architectural features (e.g., TMA for async data transfer on Blackwell, L2 locality on MI350) from only one handwritten example and ~10 unit tests. Hawkeye can port kernels across architectures (Ampere, Hopper, Blackwell), chips (NVIDIA, AMD), and precisions (FP8, NVFP4, MXFP4). Great work, co-led by @AryaTschand and @keramakr!
Show more
We’ve seen an explosion of new ML chips with unique architectural features, but software support remains the critical bottleneck Achieving peak performance increasingly relies on hardware-specific optimizations in the kernels, but we observe that coding agents are particularly weak at this Introducing Hawkeye, a framework that brings hardware-awareness to coding agents by grounding them in a minimal and comprehensive taxonomy of optimization strategies For new GPU or ML accelerator architectures, you only need to write 10 unit tests and solution kernels (one per optimization strategy), and we show that coding agents can effectively scale test-time compute with this minimal supervision to write hardware-aware kernels Hawkeye can port kernels across architectures (Ampere, Hopper, Blackwell), vendors (NVIDIA, AMD), and precisions (FP8, NVFP4, MXFP4) while consistently leveraging hardware features and approaching expert kernel performance Work co-led with @keramakr and done in collaboration with Alexander Ingare @simonguozirui @18jeffreyma @ZishenW @simran_s_arora @Azaliamirh @profvjreddi
Show more
LLM-as-a-Verifier keeps pushing the frontier of cost vs. capability! On Terminal-Bench 2.1, it made DeepSeek V4 Flash accuracy go from 79% → 88%, while being 4-11x cheaper than competitors! Try it here: @jackyk02
Show more
Scaling self-verification with DeepSeek V4 Flash beats Claude Fable 5 on Terminal-Bench 2.1, while being 11x cheaper 💰 As open-source models become more capable, they can now generate large numbers of high-quality candidate solutions and verify their own outputs at very low cost. For example, we find that sampling just 5 solutions with DeepSeek V4 Flash and ranking them using the same model with LLM-as-a-Verifier can lead to a significant boost in accuracy (79% → 88%), outperforming closed frontier models on Terminal-Bench. Try it out today: More on verification scaling in my previous post.
Show more
Intelligence-per-joule is increasing quickly because models and chips are improving and the gains compound. Though demand for inference is growing even faster. @Avanika15 @JonSaadFalcon @HazyResearch
Show more
Thanks, Jeff! So excited for you all and rooting for you! ❤️
@Azaliamirh @Sanjay_Ghemawat @quocleix @OriolVinyalsML Thanks, @Azaliamirh, and thanks to both you and @annadgoldie for being so kind in offering advice to us as we start this journey and having us over for Ricursive's happy hour a little while ago! We're inspired by what you all are building there!
Show more
Congratulations to @JeffDean, @Sanjay_Ghemawat, @quocleix, and @OriolVinyalsML on Discovery Loop!!
Announcing Discovery Loop! I am very excited to announce that, along with my longtime friends and collaborators @Sanjay_Ghemawat, @OriolVinyalsML and @quocleix, we are founding Discovery Loop (@DiscoLoopAI), a Public Benefit Corporation whose mission is to automate machine learning, science, and engineering to accelerate discoveries and progress. The four of us have worked together for 14 to 30 years, and have helped build some of the world’s most used products, infrastructure and AI models, and we’re excited to turn our attention to this ambitious endeavor. ♾ Learn more at:
Show more
I had so much fun creating the Self-Improving AI Agents course with @achowdhery, and teaching it twice in one year at Stanford! We also collaborated with Stanford Online to make the course available online: YouTube: The field is moving incredibly fast, but we tried to focus on the core concepts that help us build better AI systems. I hope you enjoy it!
Show more
0
40
1.1K
117
Forward to community
"AI inference demand is expected to grow 10,000x over the next 5 years." Yesterday at @DACconference, @annadgoldie took the stage to share how we break today's deadlock between model and hardware. That's what we're doing at Ricursive, using our very own model. 📍 Stop by Booth #1049# to say hi.
Show more
Thanks @nvidia and Tim Costa for the shout out to @RicursiveAI in your DAC Welcome Address!
@RicursiveAI on @nvidia 's slide at the #DAC2026# opening keynote. Thank you for the shoutout, Timothy Costa. Next up: our co-founder @Azaliamirh on the "New Minds, New Models" panel, 10:30am, Room 104A. Then find us at Booth #1049#.
Show more
Better HIP kernels through synthetic data, multi-agent search, and reinforcement learning. See how researchers at @Stanford's Scaling Intelligence Lab are advancing code generation for AMD GPUs.
Show more
Our co-founder @Azaliamirh met with @TIME at last week's @RaiseSummit to discuss Ricursive. We're using AI to revolutionize chip design. And instead of years, it will take days.
⛳️ We built a $40,000 TrackMan golf shot tracker with a Raspberry Pi and a ~$100 ti iwr6843sk radar. 🏌️‍♂️More details below!
"LLM-as-a-Verifier: A General-Purpose Verification Framework" The key idea of this paper is that it does not ask for one rough score, it reads the model’s full uncertainty over scores, which helps to make the judgment much more fine-grained. This approach lets agents pick better solutions, track progress, and learn from denser feedback.
Show more
Check out LLM-as-a-Verifier: a simple, cheap, & general-purpose self-improvement technique that boosts performance on "any" agentic task we've tried. It achieves SOTA on Terminal-Bench V2, SWE-Bench Verified, RoboRewardBench, and MedAgentBench. The key idea: - Use fine-grained scoring granularity (e.g. 1-20) - Scale model responses with repeated sampling and criteria-based scoring - Rank results based on the expected logprobs of said scores We made it easy for you to try: Code: Claude Code Plugin: Paper: Work is led by @jackyk02, with an awesome team!
Show more
Chat with the authors of TRACE at ICML today!
@hangoo_kang and I will be presenting “TRACE: Capability Targeted Agentic Training” today at FAGEN. Please stop by and check it out!