Are you still assuming the better result comes from the bigger model?
PRAXIST ran on an open-source model and beat a Claude Opus 4.8 stack.
> agent tries five things
> one works, four don’t
> it saves the one and deletes the rest
> next run starts from zero and fails the same way again
> PRAXIST keeps everything, including what didn’t work
> then decides what to try next based on all of it
> 49 gold medals on 75 MLE-Bench tasks for $3,054
the baseline spent $38,370 and won fewer. that’s the setup, not the model.
Introducing PRAXIST Beta, your autonomous research team🚀
Define the objective, constraints, and what success looks like. PRAXIST discovers the path.
From there, PRAXIST takes on the experimental research loop. Multiple Research Peers explore competing approaches in parallel, share useful findings, and build on accumulated evidence across experiments and generations.
In partner-provided environments, PRAXIST reached a 100% safe-landing rate in a rocket simulation and reduced accumulated error in an industrial SLAM system from 9.37 centimeters to 5.01 centimeters.
From robotics control to quantitative finance, PRAXIST has delivered measurable advances across fundamentally different problems by combining a shared core research architecture with domain-specific tools, knowledge, constraints, and evaluators.
More approaches explored. Faster iteration. Greater R&D capacity without proportionally increasing specialist headcount.