Register and share your invite link to earn from video plays and referrals.

Azalia Mirhoseini
@Azaliamirh
Founder @RicursiveAI, Asst. Prof. of CS at Stanford. Prev: DeepMind, Anthropic, Brain. Co-Creator of MoEs, AlphaChip, Test Time Scaling.
Joined May 2013
638 Following    21K Followers
LLM-as-a-Verifier keeps pushing the frontier of cost vs. capability! On Terminal-Bench 2.1, it made DeepSeek V4 Flash accuracy go from 79% → 88%, while being 4-11x cheaper than competitors! Try it here: @jackyk02
Show more
Scaling self-verification with DeepSeek V4 Flash beats Claude Fable 5 on Terminal-Bench 2.1, while being 11x cheaper 💰 As open-source models become more capable, they can now generate large numbers of high-quality candidate solutions and verify their own outputs at very low cost. For example, we find that sampling just 5 solutions with DeepSeek V4 Flash and ranking them using the same model with LLM-as-a-Verifier can lead to a significant boost in accuracy (79% → 88%), outperforming closed frontier models on Terminal-Bench. Try it out today: More on verification scaling in my previous post.
Show more