After a day of dealing with C code in the linux kernel, which is about as hard and subtle a task I could ever deal with, I have a better feel for Grok 4.6's capabilities by comparing it side by side with Claude Opus5.
First, 4.6 in Grok Code has a new effort level of extra high, compared to its default and previous max of high. If you're going to do anything serious it's a big step up from high and worth the extra time and tokens.
Its reasoning appears to be even better than Opus 5 - I had them argue out potential solutions and Grok more often than not outdid Opus 5.
Its code is better than 4.5 was, but still below what I got out of Opus 4.8, missing some subtle detail that Opus would not have.
It also does just what you ask it to do in code, and nothing more. Opus 5 on the other hand, if you ask it to write code, it will write the code, then go off and create well thought out test frameworks, and execute them without being asked. Either approach is valid depending on your viewpoint or workflow, but in the end Opus 5 won out for me with its approach for my task.
I can certainly see the improvement in Grok, and directionally that it could easily overtake the competition with its rapidly expanding parameter models coming up, but Opus5 won this battle.
That said, I have cancelled my Claude account as I mainly got it for the free Fable5 access and for the subtle work I was doing it was better, more concise, and actually faster than Opus5. But I kept tripping safeguards and ran out of credits and it became nigh on useless. So I'm running on remaining credit till my account terminates.
The next grok model probably won't be out before my account terminates so I'll probably give in and re-subscribe for another month to fill that gap. Also having two completely different models check each others' reasoning leads to a much better result than either alone, and perhaps I'll need it regardless for that superpower.
더 보기
Long term I'm betting on Grok getting there. Grok 4.5 was just below opus 4.8, and Grok 4.6 looks better but it's too early to tell (and remains the same number of parameters as 4.5). I expect 5 will be insanely good. That guy has more money than Satoshi to throw at it.
더 보기