登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Jason Wei
@_jasonwei
ai researcher @meta, past: openai, google 🧠
参加 October 2020
733 フォロー中    112.1K ファン
Muse Spark 1.1 outperforms GPT-5.6 Sol and Gemini 3.1 on Radiology's Last Exam. We don't beat Fable (yet). And Humans are still a lot better, but we are working on closing the gap!
🔥Today, we are releasing one of the first visual reasoning benchmarks for autonomous AI diagnosis in healthcare! 🚀Introducing Radiology’s Last Exam 2.0 (RadLE 2.0) from @CRASHLabAI, an uncertainty-aware benchmark for autonomous diagnosis in radiology! ✅In the last few days, the AI frontier has moved significantly. @OpenAI launched GPT-5.6 Sol. @Meta launched Muse Spark 1.1. @xAI dropped Grok 4.5. 🙌We’ve benchmarked all frontier, open-source and medical VLMs in RadLE2.0 and the leaderboard is now LIVE! 🚨 Before AI models are handed autonomy, one question matters more than any accuracy score: Do they know when to STOP and hand over to a human? ⚠️ A confident wrong diagnosis is far more dangerous than an honest “I don’t know.” Yet most models are bad at admitting the latter! 🚀 We release five RadLE 2.0 Scores: Confidence Weighted, Reliability, Accuracy, Safety and Handover Readiness and we find that models from @OpenAI @AnthropicAI @MetaAI @GoogleDeepMind @xAI @nvidia @Alibaba_Qwen @MistralAI @MiniMax_AI all score very differently as they optimize for different metrics! 🚨But most importantly, NONE of the Models have been able to reach the average human expert baseline! ⚡️A thread on what we found and which models aced our metrics! Link to the leaderboard and technical report at the end of the thread!
もっと見る