Check out LLM-as-a-Verifier: a simple, cheap, & general-purpose self-improvement technique that boosts performance on "any" agentic task we've tried.
It achieves SOTA on Terminal-Bench V2, SWE-Bench Verified, RoboRewardBench, and MedAgentBench.
The key idea:
- Use fine-grained scoring granularity (e.g. 1-20)
- Scale model responses with repeated sampling and criteria-based scoring
- Rank results based on the expected logprobs of said scores
We made it easy for you to try:
Code:
Claude Code Plugin:
Paper:
Work is led by
@jackyk02, with an awesome team!