How can we extract richer signals from AI Feedback?
Introducing LLM-as-a-Verifier✨— a simple verification scaling framework that achieves SOTA on agentic benchmarks 🚀
The key idea:
- Use fine-grained scoring granularity (e.g., 1-20 instead of the standard 1-5 scale)
- Take the expectation over the full logprob distribution of score tokens
- Scale repeated evaluation and criteria decomposition
You can use these fine-grained signals for more effective test-time scaling, RL, and agent monitoring! It achieves SOTA across Terminal-Bench V2, SWE-Bench Verified, RoboRewardBench, and MedAgentBench 👑
Advised by
@Azaliamirh @istoica05 @drmapavone @chelseabfinn
🧵👇