๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Jacky Kwok
@jackyk02
Stanford CS PhD | Berkeley EECS
๊ฐ€์ž… June 2025
899 ํŒ”๋กœ์ž‰ ์ค‘    693 ํŒฌ
How can we extract richer signals from AI Feedback? Introducing LLM-as-a-Verifierโœจโ€” a simple verification scaling framework that achieves SOTA on agentic benchmarks ๐Ÿš€ The key idea: - Use fine-grained scoring granularity (e.g., 1-20 instead of the standard 1-5 scale) - Take the expectation over the full logprob distribution of score tokens - Scale repeated evaluation and criteria decomposition You can use these fine-grained signals for more effective test-time scaling, RL, and agent monitoring! It achieves SOTA across Terminal-Bench V2, SWE-Bench Verified, RoboRewardBench, and MedAgentBench ๐Ÿ‘‘ Advised by @Azaliamirh @istoica05 @drmapavone @chelseabfinn ๐Ÿงต๐Ÿ‘‡
๋” ๋ณด๊ธฐ