Do models need their native harness for coding?
Awesome work from our intern
@melissapan on this. She dug into whether the harness (Claude Code vs Codex CLI vs Pi) actually moves the needle for coding agents.
Turns out it matters way less than people assume. 21 model-harness pairs, real rigor. More from Arena to come.
More details on the Arena blog:
Does your Claude model really need Claude Code…? 🤔
We evaluate 7 models on Claude Code, Codex, and Pi. Three surprising findings emerge:
1️⃣Harness choice has little effect on task success rate, but can significantly affect the cost
2️⃣A simple harness can be competitive
3️⃣The native harness isn’t always the best.
Millions of people are using coding agents, but the impact of harness choice remains unclear.
(1/n) More details in the thread. 🧵
Show more