Finally: an independent, head-to-head test of clinical AI models.
This week, Stanford's Arise Lab published one of the most comprehensive independent evaluations of clinical AI to date. The benchmark compared 24 frontier and clinical AI models, producing 12,747 expert rankings on their answers to real-world clinical questions.
We were excited to see Doximity Ask outperform
OpenEvidence, GPT-5.6, Claude Fable 5, Gemini 3.1 Pro, and other leading frontier AI models.
The bigger story, though, is that clinical AI is finally getting the kind of rigorous, independent benchmarking it needs.