Register and share your invite link to earn from video plays and referrals.

Azalia Mirhoseini
@Azaliamirh
Founder @RicursiveAI, Asst. Prof. of CS at Stanford. Prev: DeepMind, Anthropic, Brain. Co-Creator of MoEs, AlphaChip, Test Time Scaling.
Joined May 2013
626 Following    20.3K Followers
Check out LLM-as-a-Verifier: a simple, cheap, & general-purpose self-improvement technique that boosts performance on "any" agentic task we've tried. It achieves SOTA on Terminal-Bench V2, SWE-Bench Verified, RoboRewardBench, and MedAgentBench. The key idea: - Use fine-grained scoring granularity (e.g. 1-20) - Scale model responses with repeated sampling and criteria-based scoring - Rank results based on the expected logprobs of said scores We made it easy for you to try: Code: Claude Code Plugin: Paper: Work is led by @jackyk02, with an awesome team!
Show more
How can we extract richer signals from AI Feedback? Introducing LLM-as-a-Verifier✨— a simple verification scaling framework that achieves SOTA on agentic benchmarks 🚀 The key idea: - Use fine-grained scoring granularity (e.g., 1-20 instead of the standard 1-5 scale) - Take the expectation over the full logprob distribution of score tokens - Scale repeated evaluation and criteria decomposition You can use these fine-grained signals for more effective test-time scaling, RL, and agent monitoring! It achieves SOTA across Terminal-Bench V2, SWE-Bench Verified, RoboRewardBench, and MedAgentBench 👑 Advised by @Azaliamirh @istoica05 @drmapavone @chelseabfinn 🧵👇
Show more