가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Rohan Paul
@rohanpaul_ai
Compiling in real-time, the race towards AGI. The Largest Show on X for AI. 🗞️ Get my daily AI analysis newsletter to your email 👉
가입 June 2014
6.8K 팔로잉 중    154.3K
Grok 4.6 just dropped. Clearest strength is professional agent work, with leading results against GPT Sol Max and Fable 5 Max on GDPVal-AA v2 and AA-Briefcase. - matches GPT-5.6 Sol Max at 61 on Artificial Analysis while charging $2/$6 per million input/output tokens. - 1753 on GDPVal-AA v2 and 1577 on AA-Briefcase, both ahead of Sol Max and Fable 5 Max. Those benchmarks measure real-world agent tasks and agentic knowledge work, making them closer to research, analysis and multi-file deliverables than isolated question answering. On coding, 69.9% on CursorBench v3.2 beats Sol's 67.2%, while DeepSWE and Terminal-Bench leave Grok behind both Sol and Fable. SpaceXAI attributes the jump to a longer training run, regenerated SFT trajectories, model-based trace filtering, and agentic RL across coding, web development, CAD and kernel optimization. It also reports more self-testing on long trajectories, with the model checking its work before continuing, directly targeting error accumulation across multi-step agents.
더 보기