가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Sumanth
@Sumanth_077
Simplifying LLMs, Machine Learning & AI Agents for you! • Building • Shipping Open Source AI Apps
가입 July 2021
871 팔로잉 중    76.7K
The model humans prefer the most is not always the most accurate one! Arena just added factuality scoring to their leaderboard alongside human preference. The idea: human preference measures whether a response felt good to read. Factuality measures the correctness of claims in a model's response. These are two different things. Most benchmarks test models on fixed, predetermined questions in a lab setting. What Arena is doing differently: measuring factuality in real, open-ended user conversations at scale. Over 2 million claims extracted from actual conversations, verified against the web, across 130k Text Arena battles and 40k Search Arena battles. The way it works: they randomly sample battles, extract web-verifiable claims from each model response, verify them against the web, and calculate a factuality score per model. The final ranking is a weighted combination of human preference and factuality score. When factuality weight is enabled, the shifts are significant. In Text Arena, GPT-5.5 jumps 13 spots. Muse Spark drops 13. Claude Fable 5 moves down slightly. In Search Arena, GPT-5.5-search takes the top spot while Gemini grounding drops from 7 to 13. I've shared the full methodology blog in the replies!
더 보기