Register and share your invite link to earn from video plays and referrals.

Search results for LLMBenchmark
LLMBenchmark community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including LLMBenchmark
TL;DR: Clinical AI benchmarking has a core problem — real EHRs can't be shared and their labels aren't verifiable. This paper solves it with a fully synthetic hospital, and frontier models still fall short of top physicians. Title: Synthetic Hospital: An Open, Verifiable, Physician-Validated Longitudinal EHR Benchmark URL: Points 🏥 Built from medical education materials into 1,268 patients and 5,602 encounters, fully synthetic and PHI-free so it can be shared openly 🔗 Diagnoses, findings, and temporal relations are deterministically grounded in ICD-10-CM/SNOMED CT/LOINC, with labels derived mechanically from a knowledge graph 👨‍⚕️ Physicians distinguished synthetic from real charts at just 53% accuracy — essentially chance 📊 Across 10 models on 5 tasks, the best patient-diagnosis score was Kimi 2.5-thinking at 0.732 severity-weighted F1, matching average physician performance ⚠️ Still well below top physicians (0.89); every model missed roughly half the findings in summarization tasks 🔁 Swapping the generator model changed scores by ≤0.05, confirming it measures real clinical state, not generation artifacts Feels significant to finally have a benchmark that can measure clinical LLMs honestly, without a privacy tax. #MedicalAI# #LLMBenchmark#
Show more
🌱 Seed-2.1-Pro Review: Better Post-Training Can't Hide an Aging Base @ByteDanceSeed_ Seed-2.1-Pro 0915 earns higher reasoning scores with fewer tokens, and its agent ability has climbed from nearly unusable to passable. But in the three months since the last version, rivals iterated roughly a generation and a half — and the old base model is running out of tricks. That is the verdict from Zhihu contributor toyama nao, who runs a long-running monthly logic benchmark and put the 0915 build through his full evaluation suite. 1️⃣ Coding and agent work: from unusable to passable The jump over the predecessor is large. Deliverable completeness now beats DeepSeek V4.1 Flash, though first-pass success still trails top-tier models — a model caught in the middle. 🔹 Frontend: some aesthetic sense, but unstable. Constrained stacks like iOS system components look decent; open stacks like web or raw canvas expose visible flaws in proportion and color. 🔹 Task adaptability: every task now at least completes. In a HarmonyOS-flavored project it barely knew, the model read docs and iterated its way to usable — a good sign. 🔹 Delivery efficiency: roughly tied with DeepSeek V4.1 Flash and GLM-5.3-Flash, splitting steps evenly between writing and verifying. All three sit below the current Chinese SOTA. 2️⃣ Multi-step reasoning: the biggest gain This is where 0915 improved most — and where it leads its tier, with a small token-efficiency edge. On problems where GLM-5.3 falls into exhaustive enumeration, 0915 repeatedly finds higher-scoring answers with fewer tokens, showing what the author calls real "big-model intuition." The caveat: multi-turn reasoning, which demands in-context learning and reflection, remains mediocre — on par with Chinese peers. 3️⃣ Where 0915 still stumbles Hallucination runs high, and context confusion appears regardless of prompt length — especially when source details are tangled. The agent symptom is subtler than dropping requirements: 0915 keeps every requirement but misreads semi-ambiguous ones — arguably the more dangerous failure mode. 4️⃣ The contrarian bet: a true non-thinking mode Seed is one of the few teams still maintaining a genuinely non-thinking mode, and this version quietly got good: average output dropped from 8K tokens back to ~1K, with no measurable capability regression, a slight gain in complex reasoning, and readable prose intact. For latency-sensitive scenarios that still need some reasoning, the author considers it a legitimate option. 5️⃣ The long march His closing frames 0915 as a rest stop, not a destination: the predecessor's lukewarm market reception forced ByteDance's team into a forced march of biweekly iterations, and better post-training is now visibly paying off in cost per task. But the base is old, and the competition has moved. His last line is worth keeping: sometimes the long way around is the real shortcut. 🔗 Full Reading: 🔗 Key links: Author's monthly logic benchmark (Aug 2026): #ByteDance# #Seed# #LLM# #AIAgents# #LLMBenchmark# #ReasoningModels# #AI#
Show more
LOL! The US government is now the new LLM benchmark authority 😂 They say Kimi K3 is much worse than US frontier models LiveBench AI said the same thing a few days ago Kimi is a very cheap sonnet / opus 4.6 class model for long running tasks - that’s not frontier intelligence
Show more