Grok 4.6 just took #
1# on CursorBench.
Not only the highest score.
It did it at a fraction of the cost of the models sitting right behind it.
That combination is the real signal.
Top-tier results are one thing.
Top-tier results that stay cheap enough to run for long agentic coding sessions are something else.
I see this as the practical edge that matters for real work.
Benchmarks are useful.
Sustained performance at low cost is what actually gets used.