Agents on Rails: Stage 2 is live. We wanted to find out: can you hand a model a real feature ticket and trust what comes back?
The jump from atomic tasks to feature requests has interesting results…
@OpenAI GPT-6 Astra is new to the leaderboard, and it came out on top: 35% of tasks solved, with 9-minute median runs, and relatively low cost, all at its default effort level: medium.
@AnthropicAI Claude Fable 5.1 still performed well at second place, but came with a hefty price tag (almost 4x the cost of Astra).
@GeminiApp 3.8 Flash was third place with a cost comparative to Astra, but took more 200 steps and longer at 27 minutes per run.
At the bottom of the leaderboard,
@OpenAI GPT-5.6 Luna, which did well in Stage 1 (46/63 tasks for $0.90), didn’t complete a single task in Stage 2 when the work required planning, migrations, testing, and completeness.
Read the full benchmark report from
@evilmartians here: