Register and share your invite link to earn from video plays and referrals.

Vals AI
@ValsAI
Public LLM Evaluation // @8vc @BloombergBeta @pearvc
269 Following    15.9K Followers
Muse Spark 1.2 is the first model to crack 60% on Finance Agent v2, our benchmark that gives models the job of a financial analyst. At $0.77/test it is 6.7x cheaper than the previous #1#, Opus 5 ($5.12), at twice the speed.
Show more
Muse Spark 1.2 just cracked the top 5 on the Vals Index, at just $0.69 per test. This is 3x cheaper than Kimi and 10x or more cheaper than Fable, Opus, and 5.6 Sol.
Introducing Vals-Smith: turn your code base into a customized benchmark. Public benchmarks tell you which model is strongest overall, not which model is the best on your code. Vals-Smith turns your merged pull requests into real coding tasks and measures the percentage a model can actually resolve. New models ship every week. Vals-Smith tells you which one to trust with your code.
Show more
Kimi K3 is the #2# overall model on our in-house Vibe Code Bench at 85.0%. VCB tests a model's ability to go from zero-to-one; creating a web application completely from scratch.
Show more
Meta just released Muse Spark 1.1 and is the new SOTA on MedScribe and TaxEval, taking the top spot from Fable 5 while being 10x cheaper and twice as fast. Meta currently holds the top 2 spots on TaxEval It is also the new #1# on Harvey's Legal Agent Bench, dethroning Grok 4.5 less than 24 hours after it took the top spot.
Show more
0
77
1.2K
129
Forward to community
One of our team members attended @aiDotEngineer conference last week and presented Vibe Code Bench in the poster session. Great conversations with everyone who stopped by!
Excel Modeling Benchmark shows that AI agents are starting to build useful financial models, but they still aren’t ready to replace the analyst workflow end-to-end To learn more about this benchmark and for full results visit
Show more
LBO and DCF models are the hardest categories because each is one long chain of linked calculations where single early error cascades through everything downstream. Failures can also be quite spectacular. On one DCF modeling task, most of the model was built correctly, but a single runaway revenue line cascaded through the model and produced an implied share price of negative $67 trillion. In tightly linked financial models, even one bad cell can ruin the entire model.
Show more
Opus 4.8 is the most accurate but among the most expensive at $12/task, roughly 4x GPT 5.5 which reaches 64.5% at a quarter of the cost. Open source mode, Kimi K2.6 is the most cost-efficient of the leaders, holding 58% accuracy at $2.20/task
Show more
We found that no model is close to client-ready deliverables. Claude Opus 4.8 leads with 69.4% accuracy, ahead of Claude Sonnet 5 (66.3%) and GPT 5.5 (64.5%). When creating models from scratch, numerical correctness is the primary bottleneck: Opus 4.8 passes 87% of formula checks and 74% of presentation checks, but only 61% of numerical checks. A formula can point at the right cell and still compute the wrong value when an upstream input is off. The writing looks correct while the numbers do not.
Show more
EMB grades each task in two modes. In Template mode, the agent fills in a provided workbook skeleton and is scored on exact cell matches against a gold model. In Scratch mode, it builds the workbook from nothing, with an agent-as-judge scoring the workbook on numerical accuracy, formula wiring, and presentation. One mode tests whether an agent can follow the structure while the other tests whether it can create one.
Show more
Agents work in a sandbox with two main tools: a bash tool to read the source data and create the Excel model, and an execute tool that recalculates it through an Excel engine. Every submission is graded by first recalculating its formulas in Microsoft Excel, so a model only scores if it actually runs.
Show more
It takes investment bankers and private equity analysts more than five hours to build Excel financial models by hand. Can AI build the same models more efficiently? Today we're releasing the Excel Modeling Benchmark (EMB), a benchmark that measures whether AI agents can build complete, working financial models from a prompt and a set of source spreadsheets. The models span 7 categories, including LBOs, DCFs, and M&A Models
Show more
Congrats to @AnthropicAI on the release! Results on all benchmarks coming shortly.
The model has a 1M context window. It was evaluated with max effort on all benchmarks aside from Terminal Bench (which used high), 128k max output tokens, and default temperature and top p. The model priced at $2 / $10 tokens through August 31, after which it switches to $3 / $15 (same as previous Sonnet models). All prices reported on Vals AI use the latter cost.
Show more
We noticed a small but noticeable quantity of refusals on one of our financial benchmarks, CorpFin v2. 15 out of 858 tasks were refused, mainly for “bio”.
Outside coding it’s a smaller step from Sonnet 4.6 - on Finance Agent, it’s 0.9% higher, on CorpFin v2, it’s +1.3%.
The models main strength is in coding. On the index splits, it scores 75.5% SWE-bench Verified, 86.9% on Vibe Code Bench, and 74.5% Terminal-Bench 2.1.
We are releasing a live leaderboard for @harvey's Legal Agent Benchmark on Vals AI. We are the first third-party to host this benchmark live. Results are on the private, held-out test set, not the public set.
Show more
Can AI do the job of a financial analyst? We just released V2 of our Finance Agent Benchmark and tested the frontier models. The results are tighter than you'd expect.