Register and share your invite link to earn from video plays and referrals.

Vals AI
@ValsAI
Public LLM Evaluation // @8vc @BloombergBeta @pearvc
Joined March 2024
269 Following    16.2K Followers
We found that no model is close to client-ready deliverables. Claude Opus 4.8 leads with 69.4% accuracy, ahead of Claude Sonnet 5 (66.3%) and GPT 5.5 (64.5%). When creating models from scratch, numerical correctness is the primary bottleneck: Opus 4.8 passes 87% of formula checks and 74% of presentation checks, but only 61% of numerical checks. A formula can point at the right cell and still compute the wrong value when an upstream input is off. The writing looks correct while the numbers do not.
Show more