Register and share your invite link to earn from video plays and referrals.

Ruby on Rails
@rails
Ruby on Rails scales from PROMPT to IPO. Token-efficient code that's easy for agents to write and beautiful for humans to review.
41 Following    137K Followers
Agents on Rails: You asked, so we turned every model in Agents on Rails up to its max effort level. The result: more effort/reasoning doesn’t always mean better results. @OpenAI's models made the biggest gains, costs nearly doubled overall...and the newest agent in the benchmark, DeepSeek 4.1 Flash, figured out it was being benchmarked and tried to hack its way to a better score. What an entry. Here’s what we learned and what max effort gets you with each model:
Show more
Set your livestream reminder on YouTube here:
For those who can't make it to Austin: Tune in live to watch the Rails World 2026 Opening Keynote where @dhh will share what's new in Rails, what's coming up, and where Rails is headed in the future. Join from wherever you are in the world: 9:30 AM - Austin 7:30 AM - Pacific 10:30 AM - Eastern 3:30 PM - UK 4:30 PM - Central Europe 11:30 PM - Japan This livestream is made possible by the support of Rails World Event Partner, @Shopify.
Show more
Agents on Rails: Stage 2 is live. We wanted to find out: can you hand a model a real feature ticket and trust what comes back? The jump from atomic tasks to feature requests has interesting results… @OpenAI GPT-6 Astra is new to the leaderboard, and it came out on top: 35% of tasks solved, with 9-minute median runs, and relatively low cost, all at its default effort level: medium. @AnthropicAI Claude Fable 5.1 still performed well at second place, but came with a hefty price tag (almost 4x the cost of Astra). @GeminiApp 3.8 Flash was third place with a cost comparative to Astra, but took more 200 steps and longer at 27 minutes per run. At the bottom of the leaderboard, @OpenAI GPT-5.6 Luna, which did well in Stage 1 (46/63 tasks for $0.90), didn’t complete a single task in Stage 2 when the work required planning, migrations, testing, and completeness. Read the full benchmark report from @evilmartians here:
Show more
This year @Dell is a first-time sponsor of #RailsWorld#, and they have a lot in store for you. Attendees can hang out in the Dell and @intel lounge, see AI run on-device with #Omarchy# on #XPS#, and explore a full lineup of laptops, workstations, and monitors.
Show more
Agents on Rails: lemans is now open source, and today four new models hit the benchmark: @openai’s Terra, @Alibaba_Qwen 3.8-27B, @AnthropicAI's Sonnet 5, and ox-alpha, a new model just released in stealth mode. As of August 24, 2026: - Fastest: @openAI Terra is now the fastest model, at a median of 3 minutes 2 seconds per task. - Most accurate: Still @claudeai Opus 5 by @AnthropicAI. Solved 92% of runs (58 of 63), and has the best API recall of all the models (35%). - Cheapest: @OpenAI GPT-5.6 Luna is still the cheapest from the models we can accurately price - just .012 cents per run. Read the latest benchmark report from @evilmartians to learn about lemans and see how the new models stacked up:
Show more
Of the 4 new models benchmarked, @spacexai Grok 4.6 was the better performer, ranking 4th in accuracy at 84% (compared to the current best: Claude Opus 5 at 92%), while costing 60% less than Opus to achieve those results. Grok 4.6 also had 33.3% recall, making it third best on the leaderboard for knowing (and using) Rails APIs. @AnthropicAI Opus 4.8 ranked second for speed at 3 minutes 36 seconds, just 7 seconds behind the fastest (Luna). @GoogleDeepMind Gemini Flash 3.7 performed mid to low on all fronts, but is one of the cheaper options.
Show more
Agents on Rails: we added four new models to the benchmark: @SpaceXAI Grok 4.6 (your number 1 request) @ZhipuAI GLM 5.3 (released 3 days ago) @GoogleDeepMind Gemini 3.7 Flash (released 4 days ago) and @AnthropicAI Claude Opus 4.8 …but none of them made it to the top of the leaderboard. As of August 17, 2026: - Most accurate: Still Claude Opus 5. Solved 92% of runs (58 of 63). (With Kimi a close second for a lot less cost.) - Cheapest: GPT-5.6 Luna is still the cheapest from what we can tell. (GLM 5.3 required a subscription, so costs for it are currently unknown.) - Fastest: Luna again, at a median of 3.3 minutes per task. - Best combination of all three: @OpenAI GPT-5.6 Sol. We also updated the insights and shared the full traces of the first two rounds (every command, diff, and verdict) into GitHub for you to explore. Read the latest benchmark report from @evilmartians here:
Show more
Agents on Rails: We ran 8 models against 21 atomic tasks to see which were best at writing Rails code. 3 runs each: a bug report, a security finding, a feature request. The first benchmark report with findings is now live. So: what did we discover? As of August 2026: - Most accurate: @claudeai Opus 5 by @AnthropicAI (by a hair). Solved 92% of runs (58 of 63). (But for a little more than half the cost, you get almost the same accuracy with @Kimi_Moonshot.) - Cheapest: @OpenAI GPT-5.6 Luna. 73% of runs solved at default medium reasoning effort, and all 63 of its runs cost 90 cents combined. - Fastest: Luna again, at a median of 3.3 minutes per run task. - Best combination of all three: @OpenAI GPT-5.6 Sol. 84% accuracy, costing $0.52 and 5 minutes per run. Read all the findings in the first full benchmark report from @evilmartians here:
Show more
0
35
607
116
Forward to community
We released new security versions of Rails to address a critical security bug in Active Storage. Please upgrade your applications. See the advisory for more information
Rails was born in Dec 2003, but took its first steps in the world with the launch of the first release: v0.5.0 on July 24, 2004. Thank you to the 7000+ programmers who have contributed to Rails since then and made the framework what it is today.
Show more