๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Jacky Kwok
@jackyk02
Stanford CS PhD | Berkeley EECS
๊ฐ€์ž… June 2025
1.1K ํŒ”๋กœ์ž‰ ์ค‘    6.2K ํŒฌ
Scaling self-verification with DeepSeek V4 Flash beats Claude Fable 5 on Terminal-Bench 2.1, while being 11x cheaper ๐Ÿ’ฐ As open-source models become more capable, they can now generate large numbers of high-quality candidate solutions and verify their own outputs at very low cost. For example, we find that sampling just 5 solutions with DeepSeek V4 Flash and ranking them using the same model with LLM-as-a-Verifier can lead to a significant boost in accuracy (79% โ†’ 88%), outperforming closed frontier models on Terminal-Bench. Try it out today: More on verification scaling in my previous post.
๋” ๋ณด๊ธฐ