what’s interesting about this is that it doesn’t really work often, especially not for frontend development.
there is a hard ceiling for what a model like gpt 5.6 can accomplish, and when it realizes that it cannot move the needle on a harsh judging agent, it will simply give up.
Scaling self-verification with DeepSeek V4 Flash beats Claude Fable 5 on Terminal-Bench 2.1, while being 11x cheaper 💰
As open-source models become more capable, they can now generate large numbers of high-quality candidate solutions and verify their own outputs at very low cost.
For example, we find that sampling just 5 solutions with DeepSeek V4 Flash and ranking them using the same model with LLM-as-a-Verifier can lead to a significant boost in accuracy (79% → 88%), outperforming closed frontier models on Terminal-Bench.
Try it out today:
More on verification scaling in my previous post.