Grok 4.6 gets 98.4% of Fable 5's Intelligence Index score at a fraction of its API price.
5X cheaper on input tokens and 8X less for output tokens vs Fable 5.
Looks clearly the strongest intelligence-per-dollar equation.
Grok 4.6 scores 61 versus Claude Fable 5 Max's 62, just a 1-point or 1.6% relative gap on that composite score, while Grok costs 80% less on input tokens ($2 vs $10/M) and 88% less on output tokens ($6 vs $50/M).
Grok 4.6 just dropped.
Clearest strength is professional agent work, with leading results against GPT Sol Max and Fable 5 Max on GDPVal-AA v2 and AA-Briefcase.
- matches GPT-5.6 Sol Max at 61 on Artificial Analysis while charging $2/$6 per million input/output tokens.
- 1753 on GDPVal-AA v2 and 1577 on AA-Briefcase, both ahead of Sol Max and Fable 5 Max.
Those benchmarks measure real-world agent tasks and agentic knowledge work, making them closer to research, analysis and multi-file deliverables than isolated question answering.
On coding, 69.9% on CursorBench v3.2 beats Sol's 67.2%, while DeepSWE and Terminal-Bench leave Grok behind both Sol and Fable.
SpaceXAI attributes the jump to a longer training run, regenerated SFT trajectories, model-based trace filtering, and agentic RL across coding, web development, CAD and kernel optimization.
It also reports more self-testing on long trajectories, with the model checking its work before continuing, directly targeting error accumulation across multi-step agents.