Qwen3.8 27B: mini swe agent vs Claude Code vs Pi
What I learned
With enough iterations, boosting scores on agentic coding benchmarks is relatively easy.
On DeepSWE 1.1, Qwen3.8 scores poorly with vanilla Pi. But a benchmaxxed Pi setup beats (Reward) the Qwen team’s published result using Claude Code.
Also: experiments with thinking low show that it spends more tokens, more turns, and score lower than medium.
Full details, including an ablation study and token-efficiency analysis: