註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Karan Singhal
@thekaransinghal
Health and Human Flourishing @OpenAI
加入 April 2008
353 正在關注    12.1K 粉絲
♥️ GPT-5.6 is a major step forward for health, both at the frontier and at cost. These models push the frontier of performance per dollar, bringing the best health intelligence to all. The smallest variant, GPT-5.6 Luna, evaluated at the lowest reasoning effort, outperforms GPT-5.5 at the highest reasoning effort–despite costing 25x less. The largest variant, GPT-5.6 Sol, sets a new high bar at cost. Another especially cool result: physicians found fewer flaws in GPT-5.6 responses than physician-written responses. We collected diverse tasks that remain difficult for recent OpenAI models, across patient-facing and clinician-facing use cases. We asked speciality-matched physicians to write responses to these tasks with unlimited time and web access. We then asked other physicians to compare responses side-by-side, blinded to their source. Physicians were asked to comment on areas of improvement across five axes: accuracy, communication, completeness, instruction following, and health decision helpfulness. We then reported the fraction of responses across sources rated perfectly across all axes, across 20,000 total axis ratings. GPT-5.6 Sol appeared strongest, although all GPT-5.6 models performed significantly better than physicians.
顯示更多
0
51
921
122
轉發到社區