가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Marco Pavone
@drmapavone
Prof @Stanford, Distinguished Research Scientist and AV research lead @nvidia. PhD from @MITAeroAstro. Robotics, autonomous systems, AI. Opinions are my own.
가입 November 2018
68 팔로잉 중    6.1K 팬
Our latest finding: scaling self-verification can make open-weight models significantly more capable at a fraction of the cost. With DeepSeek V4 Flash, sampling just 5 candidate solutions and using the same model to rank them with LLM-as-a-Verifier improves Terminal-Bench 2.1 accuracy from 79% → 88%—outperforming Claude Fable 5 while costing 11× less. 💰 As open-weight models become more capable, we can generate many high-quality candidate solutions and verify them at very low cost. Try it out: More on verification scaling in @jackyk02's previous post: @StanfordAILab @StanfordEng
더 보기
How can we extract richer signals from AI Feedback? Introducing LLM-as-a-Verifier✨— a simple verification scaling framework that achieves SOTA on agentic benchmarks 🚀 The key idea: - Use fine-grained scoring granularity (e.g., 1-20 instead of the standard 1-5 scale) - Take the expectation over the full logprob distribution of score tokens - Scale repeated evaluation and criteria decomposition You can use these fine-grained signals for more effective test-time scaling, RL, and agent monitoring! It achieves SOTA across Terminal-Bench V2, SWE-Bench Verified, RoboRewardBench, and MedAgentBench 👑 Advised by @Azaliamirh @istoica05 @drmapavone @chelseabfinn 🧵👇
더 보기