Same prompt.
15ร difference in cost.
Four completely different frontend implementations.
I compared Claude Fable 5,
@Zai_org GLM-5.2, DeepSeek V4 Pro, and Qwen3.7 Max in the
@novita_labs Arena Playground using the exact same prompt.
Task:
Build a single-file HTML/CSS/JS animation where two airplanes collide mid-air with clouds, particles, and a pause button.
Results ๐
โ๏ธ Claude Fable 5
๐ฐ $0.5357
๐ช 10,715 tokens
โ๏ธ GLM-5.2
๐ฐ $0.0343
๐ช 7,796 tokens
โ๏ธ DeepSeek V4 Pro
๐ฐ $0.0604
๐ช 18,885 tokens
โ๏ธ Qwen3.7 Max
๐ฐ $0.0290
๐ช 7,729 tokens
Same prompt. Different outcomes.
Different animation quality
Different UX decisions
Different code structure
Up to 15ร difference in cost
Benchmark scores are useful.
But asking models to build the same real-world app reveals something different: how they design, reason, and implement.
Which result would you ship?
Try the prompt yourself ๐