登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Apodex
@Apodex_AI
The World's First Self-Evolving Heavy-Duty Solver Try: • •
参加 April 2026
207 フォロー中    5.3K ファン
Everyone says AI agents can already analyze data, run code, generate figures, and write reports. But can they actually complete a scientific workflow? Explore FrontierChallenge — now #2# on today’s Hugging Face Daily Papers: Today, we’re introducing FrontierChallenge, a new benchmark evaluating whether AI agents can complete real scientific workflows end to end and deliver complete, verifiable results. We evaluated 12 frontier models across 97 cross-domain tasks. The highest full-completion rate was only 20.6% (GPT-5.6 Sol + Codex and Grok 4.6 + Claude Code). In electrochemistry and environmental science, every evaluated system achieved a 0% pass rate. More strikingly, 75.5% of unsuccessful Claude Code runs still ended by claiming completion. Saying “done” is not the same as delivering. Agents that can advance scientific work are already here. Agents that can reliably complete scientific workflows are NOT. That’s why we built FrontierChallenge. 🏆 Leaderboard: 💻 GitHub: 🤗 Hugging Face: 📝 Blog:
もっと見る