TL;DR: Clinical AI benchmarking has a core problem — real EHRs can't be shared and their labels aren't verifiable. This paper solves it with a fully synthetic hospital, and frontier models still fall short of top physicians.
Title: Synthetic Hospital: An Open, Verifiable, Physician-Validated Longitudinal EHR Benchmark
URL:
Points
🏥 Built from medical education materials into 1,268 patients and 5,602 encounters, fully synthetic and PHI-free so it can be shared openly
🔗 Diagnoses, findings, and temporal relations are deterministically grounded in ICD-10-CM/SNOMED CT/LOINC, with labels derived mechanically from a knowledge graph
👨⚕️ Physicians distinguished synthetic from real charts at just 53% accuracy — essentially chance
📊 Across 10 models on 5 tasks, the best patient-diagnosis score was Kimi 2.5-thinking at 0.732 severity-weighted F1, matching average physician performance
⚠️ Still well below top physicians (0.89); every model missed roughly half the findings in summarization tasks
🔁 Swapping the generator model changed scores by ≤0.05, confirming it measures real clinical state, not generation artifacts
Feels significant to finally have a benchmark that can measure clinical LLMs honestly, without a privacy tax.
#
MedicalAI# #
LLMBenchmark#