Register and share your invite link to earn from video plays and referrals.

Apodex
@Apodex_AI
The World's First Self-Evolving Heavy-Duty Solver Try: • •
207 Following    5.3K Followers
Everyone says AI agents can already analyze data, run code, generate figures, and write reports. But can they actually complete a scientific workflow? Explore FrontierChallenge — now #2# on today’s Hugging Face Daily Papers: Today, we’re introducing FrontierChallenge, a new benchmark evaluating whether AI agents can complete real scientific workflows end to end and deliver complete, verifiable results. We evaluated 12 frontier models across 97 cross-domain tasks. The highest full-completion rate was only 20.6% (GPT-5.6 Sol + Codex and Grok 4.6 + Claude Code). In electrochemistry and environmental science, every evaluated system achieved a 0% pass rate. More strikingly, 75.5% of unsuccessful Claude Code runs still ended by claiming completion. Saying “done” is not the same as delivering. Agents that can advance scientific work are already here. Agents that can reliably complete scientific workflows are NOT. That’s why we built FrontierChallenge. 🏆 Leaderboard: 💻 GitHub: 🤗 Hugging Face: 📝 Blog:
Show more
0
187
225
143
Forward to community
Reactions to Apodex 1.1 this week: “Apodex spawned out of nowhere.” “From never heard of to frontier level.” We’ll take the curiosity! Wondering the same things? Come ask us live on @LocalLLaMAsub today 👀 AMA · 8–11 AM PT Join the r/LocalLLaMA community and warm up for the AMA ↓
Show more
0
151
109
110
Forward to community
Meet Apodex 1.1: Scaling Agentic Intelligence for Complex Work Open Source Harness: Open Weights: We’re excited to introduce Apodex 1.1, our new model family built to scale agentic intelligence for professional work. 🧠 Frontier-level intelligence for complex work Apodex 1.1 brings frontier-level agentic performance across complex professional work, scientific research, financial analysis, and deep search. 🤝 Asynchronous Agent Team Apodex 1.1 can break down complex tasks, coordinate multiple agents in parallel, continuously integrate their findings, and let you step in to guide or redirect the work at any time. 🔬 Open-source research workbench We’re open-sourcing FrontierAgent—a locally deployable research workbench for the Apodex 1.1 family, including asynchronous Agent Team. Available now: 🔹 Apodex 1.1 — our most capable frontier model, available through the Apodex online workbench 🔹 Apodex 1.1 mini — open-weight model for running complex work locally 🔹 FrontierAgent — open-source, locally deployable research workbench 🌐 Live workbench: 🦾 API platform: 📃 Paper:
Show more
0
330
1.9K
440
Forward to community
Most AI benchmarks test retrieval — can a model find the known answer? However, the hardest problems in science require discovery, can a system earn an answer nobody has yet? Meet TRACES 🧭 — the world's first benchmark for measuring discoverative AI: AI that can work through evidence, test hypotheses, and reach verifiable conclusions on problems without answer keys. Proposed by our founder @tianqiao_chen, who defined its six capabilities. Three things published today: a definition of "discoverative intelligence", a rubric to tell sound investigation from lucky guesses, and a open call for both solvers and problems *Website:
Show more
0
524
1.1K
707
Forward to community