We are launching FrontierFinance, the world’s hardest benchmark for measuring frontier financial intelligence of agentic systems. FrontierFinance is fully open, and differentiates itself by measuring agentic performance across the entire investment workflow.
Our finance and AI experts created FrontierFinance for realistic and reproducible comparison of frontier AI systems, consisting of 220 diverse queries and a total of 11,543 expert-crafted rubrics, making it the largest open benchmark of its kind. Unlike existing benchmarks which largely focus on financial data extraction, FrontierFinance covers a diverse range of use cases essential to an investor’s workflow and are harder to evaluate: from screening to research to analysis to monitoring.
The benchmark tests for an agent’s ability to exhaustively find information, perform numerical analysis and most importantly synthesize information using the taste and judgement of professional investors. This makes FrontierFinance hardest among existing finance benchmarks with a ~50% pass rate for the best system currently.
@samaya_AI's agentic system outperforms frontier models with 50.8% on the same benchmark, at roughly a 4x lower inference cost than Fable 5 and ~2.7x lower than Opus and GPT. Among frontier models, Claude Fable 5 is the best-performing on FrontierFinance, scoring 49.2%, followed by Claude Opus 4.8 at 45% and GPT 5.5 at 43.5%.
We are releasing the data, code and analysis for use by the community. See link in comments.
Samaya’s mission is to take us from information to conviction and FrontierFinance takes a large step in that direction. FrontierFinance comes from a larger internal set of almost 5000 examples, and we plan to release subsequent, harder benchmarks in the future.