Turns out the best harness design for a coding agent — planning, tool setup, context management — can flip completely depending on how capable the model is. This paper tested 176 configurations to prove it.
Title: An Empirical Study of Harness Design for Coding Agents
URL:
🧠 Highlight 1: Context management matters most when resources are scarce
At a 32k-token window, managed vs. unmanaged context created a 35.7-point gap in SWE-Bench success rate. Interestingly, the fancy "recall" mechanism was barely ever used and added no accuracy at all.
📋 Highlight 2: Planning's benefit flips with model strength
A weaker 30B model gained +11.6 points from adding a planning tool, and its rate of quitting without even attempting an edit dropped from 69% to 28%. A stronger 550B model needed no such help — adding planning there just cut cost by about 30% with no accuracy change.
🔧 Highlight 3: Tool design also depends on the model and task
Weaker models need a full dedicated toolset, while stronger models often perform better and cheaper with bash-only access. Even the same model can flip its optimal choice depending on the task type.
This really drives home that tuning each harness component to your specific model and task is worth taking seriously.
#
CodingAgents# #
LLM#