Benchmarked deepseek-v4-flash vs qwen3-8-27b on 9 complex tasks in my own agent stack. With reasoning on, qwen edges Flash on quality. Off, it scores worst of the three.
The cost isn't accuracy, it's that qwen thinks more. 30x slower, 4.5x pricier.
n=9, not a verdict.