Kimi K3 is getting called Fable/Sol level, and it's 7th in our tests.
Arena Frontend Code: #
1# at 1679 points.
Artificial Analysis: #
3# at Intelligence Index of 57.
We ran it the next day on our coding-agent repair harness against GPT-5.6 Sol, Fable 5, Grok 4.5, Opus 4.8, GLM-5.2, and Gemini 3.1 Pro.
Results:
> Last of 7 models
> 53 of 67 attempts (79%)
> $0.186 per successful fix
> 702s average wall time
Sol hit 100% (70/70) on the same suite. Grok sat at 99% and 46s.
So why does the internet sound so sure K3 is crushing coding agents, if our tests have it at the bottom?
-----
> Full write-up:
> 5-min daily signals: