가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Joël Niklaus
@joelniklaus
Harness Optimization @HuggingFace
가입 October 2011
464 팔로잉 중    2.3K 팬
Same weights. $50 of harness search. 3x on Terminal-Bench. We optimized the OpenCode harness for GPT-OSS-20B (high reasoning) with a meta-harness: an agent that reads failures, patches the harness, and validates the improvement. The model never changed. On held-out Terminal-Bench 2.1, the optimized harness scored 14.8%, up from OpenCode's 4.8%. That is 3.08x, from $49.97 of API credits. Three harness fixes carried most of the gain: - i7, verify before stopping: the model often edited code and ended the session without testing it. The harness gives it another turn to run the code. - i14, continue announced actions: the model would say "now run X" and then stop. The harness keeps the session going so it runs X. - i23, repair malformed tool calls: one extra `]` made OpenCode reject otherwise valid commands. The harness repairs that clear-cut error. How the meta-harness ran: 1. Start: stock OpenCode v1.18.13 + GPT-OSS-20B scored 8.5% on the dev split. 2. Iterate: Claude Code (Opus 5 max) reads the failed dev trajectories and proposes one patch. We build it and run it once per dev task, paired against the current parent. 3. Validate: candidates that clear a pre-registered margin get a 5-trial validation. They become the new parent only if their validated score surpasses the current one. 4. Stop: the loop ended after 23 iterations and about 2.4k dev attempts, when the $50 search budget was spent. Takeaway: pre-training puts capability in the weights, and post-training makes that capability usable. Harness optimization continues the same work at inference time, turning the model's behavior into more completed tasks. Same weights, 3x the finished tasks.
더 보기