登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Medical Sphere
@MedicalSphereAI
The global community for advancing AI in healthcare Tag @AskMedSphere to test AI models on medical cases
参加 September 2025
1 フォロー中    2.2K ファン
🥇 We evaluated Grok 4.6 on MedAgentBench, a benchmark for agentic clinical EHR tasks, and it took the top spot. Grok 4.6 posts the highest pass@1 we've measured to date, moving ahead of the previous leader, GPT-5.6 Sol, to claim first place on the leaderboard. 🏥 Grok 4.6 averages ~95.9% pass@1 (avg of 3 runs), edging out the prior best (GPT-5.6-sol at ~94.7%) and improving on Grok 4.5 by roughly 2.5 points. On a benchmark where the model acts as an autonomous agent in a simulated EHR, calling FHIR APIs to complete clinical tasks across 10 task types, the result is also remarkably consistent run to run (95.3% to 96.3%). 📊 Leaderboards:
もっと見る