Interesting results here. This is why I expect more agent workloads to run on blended models.
Pareto 26.9 from
@TheUnbiasedCo sends requests to several frontier and open models and keeps the best answer.
In the new eval of 30 agent tasks, Pareto tied GPT-6 Astra for first place at about 1/3 the cost per successful task.
It also finished tasks faster than DeepSeek V4 Pro and GLM 5.3 Flash.
We tested 6 AI models on 30 challenging agent tasks: GPT-6 Astra, Opus 5.5, GPT-6 Sol, Pareto 26.9, DeepSeek V4 Pro, and GLM 5.3 Flash.
Sol matched Opus’s score, finished faster, and cost about a quarter as much per successful task.
Here’s how all 6 models compared 🧵🧵🧵
더 보기