Coding agents like GPT‑6 Astra and DS V4.1 Flash are entering the 3D world. We need new benchmarks for their spatial reasoning. 🏗️
Introducing _BuildingBench_: can coding agents turn real-world images into coherent 3D buildings, and at what cost?
BuildingBench is a step toward coding agents creating reusable worlds that people and other agents can inspect, edit, and build on for design, games, and physical simulation.
And the quality–cost frontier is moving fast: 👀
🔥 GPT‑6 Astra achieves the highest score in our comparison at ~80% lower cost than Fable 5.1.
🔥 DeepSeek V4.1 Flash delivers strong quality at a median model cost of just $2.31 per building.
Watch the frontier shift. More details, interactive leaderboard, and GitHub in the thread 👇