Since Andrej
@karpathy released this video I’ve been working on building a benchmark for AI coding agents on real Three.js scene generation.
It’s been about a week. The site is live with:
• Blind human voting on the same prompts
• First quality-vs-cost rankings
• 5 live scene challenges (and more coming)
• Open prompt submissions
Models generate full Three.js scenes from the exact same brief. Humans vote on the actual rendered results: composition, lighting, motion, whether the code works. No text scores, no cherry-picked demos.
Some early surprises already:
@AnthropicAI Opus 5 leads overall, but
@OpenAI GPT-5.6 Luna is extremely strong for the price, and a few unspoken models are punching well above their weight.
Go vote or throw hard prompts at it:
P.S: Still early (static + human-only scoring for now). More models and richer evals coming. If you are interested in it comment or dm me.