Presenting our first post-trained model effort, that maximized intelligence per token. This is humble beginnings and wouldn’t be possible with all our research partners and a mighty team of
@nikogrupen @ItsJulioPereyra @calvincongelado @vtrengarajan @gabepereyra
Summary:
- Tenet is a Kimi K3 base that we post-trained for long-horizon, agentic legal tasks.
- We co-optimized harness to make training and task execution more effective.
- It completes 2x held out tasks on LAB and 20% more on LAB contracts than base Kimi K3
- Not just lab, the results transferred other legal benchmarks: e.g. APEX-v1,CUAD, and MAUD.
The idea is simple:
we sample multiple ways to solve the same legal task, grade them, and train the model toward the best approaches. More technically, we use GSPO to sample groups of independent rollouts, score them with rubric-based rewards, and compute advantages within each group.
When we co-optimize harness and model, we get really good results on key legal-specific product areas:
• M&A Diligence: pass rate jumped from 46% to 60% and could tackle datarooms of up to 80M tokens
• Review table: Answer quality improved by 3.6 points and citation quality by 12.1 points at roughly 1/10th the cost per cell.
• Firm Knowledge: improved criteria pass rate by 15%+, and reduced cost per query by 90%
Follow Gabe for full details: