recent notes/learnings on harnesses (small n, directionally correct, experimental results soon):
> running harness experiments is brutally expensive! e.g.
@guanlan's experiment used 2B tokens but only covered 30 tasks (21 terminal-bench tasks and 9 deepswe tasks)
> the best harness is one that's the most bitter-lesson pilled and cost-efficient. you do this by letting your harness adapt at runtime (rsi) and align itself with actual tasks and data. common failure mode is being too opinionated a priori or over optimizing for one model. good properties of harnesses have been mostly commoditized/converged.
> as horizons of execution get longer, two things stand out:
1) speed (does the harness know when to use subagents, fanouts, and have good primitives around them).
2) the ability to unblock itself. this means not getting stuck in a particular loop or deadzone but stepping outside without human intervention.
what do you look for in good harness design?