Some harness + training cooptimization we did with the cracked team
@baseten. Tons of banger charts but my favorite is data room read coverage which moves from < 1% (base agent) -> 64% (RLM harness) -> 92% (RLM harness + RL)
For basic behavioral strategies like "read the whole data room" it still surprises me every day (1) how strong models still don't just know them but, more excitingly, (2) how quickly any model learns them through RL in the right environment!