There is no "best" AI model.
Martian just showed what happens when you stop pretending there is: 46% fewer errors than the single best LLM, across 16 benchmarks, by routing each task to the right model instead of locking into one.
The part that got me: quoted token pricing is a bad predictor of real cost. Some models charge more and think less. Some charge less and burn tokens.
Backed by an oral at ICLR 2026.
If you're hardcoding one model into everything, this is the dashboard to sit with:
顯示更多
We got 46% fewer errors than the single best LLM across the 16 most used benchmarks (TerminalBench, LiveCodeBench, etc).
Here's how that's possible and what each model can achieve when used optimally (every benchmarks misses the majority of model capabilities) 👇
Interactive Site:
Academic Paper:
顯示更多