Everyone says AI agents can already analyze data, run code, generate figures, and write reports.
But can they actually complete a scientific workflow?
Explore FrontierChallenge — now #
2# on today’s Hugging Face Daily Papers:
Today, we’re introducing FrontierChallenge, a new benchmark evaluating whether AI agents can complete real scientific workflows end to end and deliver complete, verifiable results.
We evaluated 12 frontier models across 97 cross-domain tasks.
The highest full-completion rate was only 20.6% (GPT-5.6 Sol + Codex and Grok 4.6 + Claude Code).
In electrochemistry and environmental science, every evaluated system achieved a 0% pass rate.
More strikingly, 75.5% of unsuccessful Claude Code runs still ended by claiming completion.
Saying “done” is not the same as delivering.
Agents that can advance scientific work are already here. Agents that can reliably complete scientific workflows are NOT.
That’s why we built FrontierChallenge.
🏆 Leaderboard:
💻 GitHub:
🤗 Hugging Face:
📝 Blog: