Sam Altman is probably shitting his pants today...
Grok was an early beta in November 2023.
And today, xAI shipped Grok 4.7, in just under 3 yrs.
ChatGPT had already been live for about 11 months before Grok even existed... OpenAI had the head start.
On the 7 hard-work tests xAI published today, Grok 4.7 beat OpenAI's GPT-5.6 Sol on 5 of them.
1/ Coding, the longer jobs... Grok 46.3%. Sol 41.7%.
2/ Electrical engineering, actually designing circuits... Grok 64%. Sol 39.4%.
3/ Multi-hour office work... Grok 1,657. Sol 1,487.
4/ Multi-hour work in the computer terminal... Grok 38.0%, Sol 37.3%.
5/ Legal work... Grok 19.6%. Sol 2.5%.
Sol still leads the other 2.
6/ A second coding test... Sol 72.7%. Grok 71%.
7/ Doctor-style reasoning... Sol 60.5%. Grok 56.7%.
Then an outside lab (Artificial Analysis) graded them... Grok got a 46. Sol got a 47.
Just a one point difference.
On coding agents, meaning an AI that writes the code and then runs it, that same lab had Grok at 56 and Sol at 55... Grok passed Sol.
FYI, GPT-6 and Claude are still at 62.
And one more from that same lab, bc this is the test that looks like a real job. The AI had to hand you the document, the spreadsheet, the slides... Grok scored 1,695. Sol scored 1,588. OpenAI's newer GPT-6 scored 1,542. Claude was still first, at 1,735.
And FYI, the price did not go up... Grok 4.7 still costs the same as Grok 4.6 with $2 per million tokens in, $6 out (Sol is $4 in and $20 out). Materially more capabilities, but Grok is still the same price...
It's awesome to see Grok caught up to a company that was already about a year ahead, in just 2 years and 10 months.
And I bet soon, Grok will surpass them all... Elon is back with SpaceXAI!
顯示更多