TL;DR: Clinical AI benchmarking has a core problem โ real EHRs can't be shared and their labels aren't verifiable. This paper solves it with a fully synthetic hospital, and frontier models still fall short of top physicians.
Title: Synthetic Hospital: An Open, Verifiable, Physician-Validated Longitudinal EHR Benchmark
URL:
Points
๐ฅ Built from medical education materials into 1,268 patients and 5,602 encounters, fully synthetic and PHI-free so it can be shared openly
๐ Diagnoses, findings, and temporal relations are deterministically grounded in ICD-10-CM/SNOMED CT/LOINC, with labels derived mechanically from a knowledge graph
๐จโโ๏ธ Physicians distinguished synthetic from real charts at just 53% accuracy โ essentially chance
๐ Across 10 models on 5 tasks, the best patient-diagnosis score was Kimi 2.5-thinking at 0.732 severity-weighted F1, matching average physician performance
โ ๏ธ Still well below top physicians (0.89); every model missed roughly half the findings in summarization tasks
๐ Swapping the generator model changed scores by โค0.05, confirming it measures real clinical state, not generation artifacts
Feels significant to finally have a benchmark that can measure clinical LLMs honestly, without a privacy tax.
#
MedicalAI# #
LLMBenchmark#