Agents on Rails: You asked, so we turned every model in Agents on Rails up to its max effort level.
The result: more effort/reasoning doesn’t always mean better results.
@OpenAI's models made the biggest gains, costs nearly doubled overall...and the newest agent in the benchmark, DeepSeek 4.1 Flash, figured out it was being benchmarked and tried to hack its way to a better score. What an entry.
Here’s what we learned and what max effort gets you with each model: