登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Milk Road AI
@MilkRoadAI
Helping millions of investors navigate the AI markets. Track our 5 top-tier analysts portfolios inside Milk Road PRO. Come join us for just a $1 👇
参加 October 2025
407 フォロー中    48.4K ファン
Grok 4.7 just pushed SpaceXAI into the top four AI labs in the world. The new model scored 46 on the independent Artificial Analysis Intelligence Index, which is two points higher than Grok 4.6. That still places Grok behind the most powerful models from OpenAI and Anthropic, which scored as high as 53 but the gap is becoming much smaller. The biggest improvement appeared in tasks that resemble real work. Grok 4.7 scored 1,657 Elo on AA-Briefcase, which measures how well AI can complete long and complicated professional assignments. That score increased by 111 points from Grok 4.6 and placed the model alongside the newest Claude models at the frontier of agentic knowledge work. Grok also made a major jump in coding.Grok 4.7 paired with Grok Build scored 56 on the Artificial Analysis Coding Agent Index, compared with 47 for Grok 4.6. That result moved Grok Build into fourth place among coding agents and pushed it ahead of GPT-5.6 Sol. The model also scored 71% on DeepSWE, which placed it close to GPT-5.6 Sol and slightly ahead of Claude Fable 5.1 on that particular software-engineering benchmark. Grok 4.7 is showing that SpaceXAI can now compete in those areas. The company is also offering the model for $2 per million input tokens and $6 per million output tokens, which is far below the listed prices of several competing frontier models. Grok 4.7 also supports a 500,000 token context window, which allows it to process large collections of documents and extended conversations in a single session. However, the model uses a considerable amount of computing power to achieve these results. Artificial Analysis found that Grok 4.7 used roughly 81,000 output tokens per Intelligence Index task, compared with 36,000 for Grok 4.6 and 27,000 for GPT-6 Astra. This means that Grok is becoming much more capable, but it is sometimes reaching better answers by thinking longer and using more tokens.That tradeoff will matter for companies running thousands or millions of tasks because a low price per token does not always produce a low total cost per task. Grok 4.7 is not the best model on every benchmark, and it has not taken the overall crown from OpenAI or Anthropic. However, SpaceXAI no longer looks like a company merely trying to catch up. Grok is now competing near the frontier in intelligence, coding and professional knowledge work while charging considerably less than several of its largest rivals. @elonmusk has finally turned Grok into a legitimate frontier model and the gap between SpaceXAI and the leaders is closing much faster than most people expected.
もっと見る
Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index to bring SpaceXAI into the top 4 AI labs. Coding Agent Index performance has also improved, overtaking GPT-5.6 Sol Grok 4.7 scores +2 points over Grok 4.6 on the Intelligence Index, with strong performance on agentic knowledge work tasks. We evaluated the new model at xhigh reasoning effort. Congratulations to @SpaceXAI and @ElonMusk on the release! Key takeaways: ➤ Grok 4.7 joins the frontier of agentic knowledge work: Grok 4.7 gains +111 Elo over Grok 4.6 (high) on AA-Briefcase, our private benchmark for long-horizon agentic knowledge work, scoring 1657 Elo and placing it alongside Claude Opus 5 and Claude Fable 5.1 at the frontier. On GDPval-AA, it scores 1695 Elo, +90 ahead of Grok 4.6 (high). ➤ A leap in coding agent performance: Grok 4.7 (xhigh) with Grok Build scores 56 on the Artificial Analysis Coding Agent Index, up +9 points from Grok 4.6 (xhigh). Among models in their native harnesses, Grok 4.7 + Grok Build now ranks 4th, behind only Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5. ➤ Incremental performance changes elsewhere: Outside of agentic knowledge work, Grok 4.7 broadly matches Grok 4.6 (high) on the other Intelligence Index tasks. It improves on Terminal-Bench 4.0 (+4.5 percentage points) and GDP.pdf (+3.0 p.p.), with regressions on AA-LCR (-3.7 p.p.) and AutomationBench-AA (-1.1 p.p.). ➤ High token use across tasks: Grok 4.7's gains come with higher token usage. Grok 4.7 (xhigh) uses approximately 81k output tokens per Intelligence Index task, compared with 36k for Grok 4.6 (high) and 27k for GPT-6 Astra (max) - 125% and 196% more, respectively. Other model details: ➤ Context window of 500k tokens, unchanged from Grok 4.6 ➤ Pricing of $2/$6 per 1M input/output tokens with cache hits discounted to $0.50 per 1M tokens, matching Grok 4.6 ➤ Configurable reasoning effort spans low to xhigh. Our evaluation uses xhigh.
もっと見る