Where is the harness tax coming from? ๐ง
Agents can take similar numbers of turns at substantially different costs.
For Fable 5 on SWE-bench Lite, Pi and Claude Code average 15.4 and 15.3 turns per attempt, yet Claude Code costs about twice as much for a 1.1-percentage-point increase in success rate.
So what is happening each turn? One possible reason can even be observed at the first model call: Claude Codeโs mean initial context is over 10ร Piโs, with longer instructions and larger tool schemas โผ๏ธ
As models become more capable, agents may need less scaffolding. For everyday tasks, harness design should therefore prioritize cost efficiency and reliability.
(4/n)