From BuildingBench to a square kilometre of NYC 🗽
Opus 5.5 vs. Astra, reconstructing the area around Madison Square Park.
Given the same maps, terrain, and photographs, each agent rebuilt the scene independently in Blender and Unreal Engine. No internet. No human help.
Buildings, pedestrians, and traffic—though Opus’s pedestrians glide through the park without moving their feet. Why walk when you can float? 😂
Opus leads on BuildingBench. Which model builds the more convincing city? 🏙️👀
New leader alert 🚨 Opus 5.5 takes the BuildingBench crown, while GPT‑6 Astra and Sol stand out on cost efficiency. 🏗️
🥇 Opus 5.5 (max) jumps from Opus 5’s 74.5 to 86.8—the highest score yet, ahead of GPT‑6 Astra (ultra, 84.3) and Fable 5.1 (max, 81.4).
💸 GPT‑6 Astra stays close at a fraction of the cost: $7.05 per building vs. ~$45.47 for Opus 5.5. Just 2.5 points behind, at 84% lower cost.
⚡ GPT‑6 Sol takes a different tradeoff:
• Ultra: 69.8 at ~$1.92/building
• Max: 69.6 at ~$1.89/building
Below GPT‑5.6 Sol’s 73.9, but roughly 75% cheaper than its $7.54/building. Both settings join the benchmark’s cost–quality frontier.
Watch the leaderboard shift 👇 Building comparisons from Opus 5.5, GPT‑6 Astra, and GPT‑6 Sol in the thread!