🩺 For health questions, accuracy and safety matter most. On hard self-harm and suicide conversations, GPT-5 cut undesired answers by 52% versus GPT-4o.
📰 Title: Improving health intelligence in ChatGPT
🔗 URL:
💡 Overview
OpenAI shared its work on improving ChatGPT's health intelligence. It frames GPT-5 as its best model yet for health questions, with a big jump on HealthBench, a benchmark scored against physician-defined criteria.
🔍 Challenges Solved
In health, mistakes directly affect user safety. Plausible-but-wrong answers (hallucinations) and missed urgent situations are serious risks, so accuracy, clarity, and appropriate encouragement to seek clinical care are essential.
🛠 Methodology & Proposed Approach
・Evaluated on HealthBench, scoring responses against realistic scenarios and physician-defined criteria
・HealthBench Consensus has hard cases validated by 2+ physicians
・Built a Global Physician Network of ~300 physicians and psychologists who practiced in 60 countries to inform safety research
・Advisors reviewed 700,000+ model responses reflecting real-world use
📊 Use Cases / Results
・52% fewer undesired answers on hard self-harm and suicide conversations vs GPT-4o
・8x fewer hallucinations on hard conversations from o3 to gpt-5-thinking
・Over 50x fewer errors in potentially urgent situations vs GPT-4o
Useful for understanding symptoms and lab results, judging when to see a doctor, and being nudged toward appropriate follow-up care.
#
ChatGPT# #
HealthcareAI#