How can we extract richer signals from AI Feedback?
Introducing LLM-as-a-Verifierโจโ a simple verification scaling framework that achieves SOTA on agentic benchmarks ๐
The key idea:
- Use fine-grained scoring granularity (e.g., 1-20 instead of the standard 1-5 scale)
- Take the expectation over the full logprob distribution of score tokens
- Scale repeated evaluation and criteria decomposition
You can use these fine-grained signals for more effective test-time scaling, RL, and agent monitoring! It achieves SOTA across Terminal-Bench V2, SWE-Bench Verified, RoboRewardBench, and MedAgentBench ๐
Advised by
@Azaliamirh @istoica05 @drmapavone @chelseabfinn
๐งต๐