Model-harness co-optimization is helping us solve end-to-end legal use cases like M&A diligence.
Together with
@baseten and
@baselabs, we built an RLM harness for M&A Diligence and are sharing initial results from post-training on top of it.
The RLM harness lets a root agent delegate document review to sub-agents and aggregates their findings.
A few initial results to highlight that we ran over LAB Diligence tasks:
1. Harness design — Across seven models, moving from a standard tool loop to our RLM harness increased mean criteria pass rate from 23.3% to 62.4%.
2. Post-training — In a separate experiment, RL over Qwen3.5 in the RLM harness more than doubled rubric pass rate, from 29.9% to 63.0%, and increased document review coverage from 62% to 96%. This makes the much smaller Qwen 3.5 competitive with the closed frontier in the RLM harness.
We’re now doing a scaled RL run with GLM-5.3 and are excited about the potential to bring more capable models closer to completing these tasks end to end.
Thanks to
@oneill_c @mudithj and the
@baseten @baselabs teams for the collab here!