๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

cv usk
@cv_usk
AI / Software Research Notes AI Agent, LLMOps, MLOps, Software Architecture ๆŠ•็จฟใฏๅ€‹ไบบใฎๆ„่ฆ‹ใงใ™ใ€‚
๊ฐ€์ž… May 2026
279 ํŒ”๋กœ์ž‰ ์ค‘    409 ํŒฌ
TL;DR: Clinical AI benchmarking has a core problem โ€” real EHRs can't be shared and their labels aren't verifiable. This paper solves it with a fully synthetic hospital, and frontier models still fall short of top physicians. Title: Synthetic Hospital: An Open, Verifiable, Physician-Validated Longitudinal EHR Benchmark URL: Points ๐Ÿฅ Built from medical education materials into 1,268 patients and 5,602 encounters, fully synthetic and PHI-free so it can be shared openly ๐Ÿ”— Diagnoses, findings, and temporal relations are deterministically grounded in ICD-10-CM/SNOMED CT/LOINC, with labels derived mechanically from a knowledge graph ๐Ÿ‘จโ€โš•๏ธ Physicians distinguished synthetic from real charts at just 53% accuracy โ€” essentially chance ๐Ÿ“Š Across 10 models on 5 tasks, the best patient-diagnosis score was Kimi 2.5-thinking at 0.732 severity-weighted F1, matching average physician performance โš ๏ธ Still well below top physicians (0.89); every model missed roughly half the findings in summarization tasks ๐Ÿ” Swapping the generator model changed scores by โ‰ค0.05, confirming it measures real clinical state, not generation artifacts Feels significant to finally have a benchmark that can measure clinical LLMs honestly, without a privacy tax. #MedicalAI# #LLMBenchmark#
๋” ๋ณด๊ธฐ