登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Vals AI
@ValsAI
Public LLM Evaluation // @8vc @BloombergBeta @pearvc
参加 March 2024
269 フォロー中    16.2K ファン
We found that no model is close to client-ready deliverables. Claude Opus 4.8 leads with 69.4% accuracy, ahead of Claude Sonnet 5 (66.3%) and GPT 5.5 (64.5%). When creating models from scratch, numerical correctness is the primary bottleneck: Opus 4.8 passes 87% of formula checks and 74% of presentation checks, but only 61% of numerical checks. A formula can point at the right cell and still compute the wrong value when an upstream input is off. The writing looks correct while the numbers do not.
もっと見る