The gold standard for evaluating complex legal work is partner review.
Partners can cost over $3,000 an hour and associates $1,000, making expert review expensive at scale.
Evaluating a model on LAB through expert review alone would cost millions of dollars.
In practice, we combine rubric-based LLM scoring with sampled human preference.
But rubrics miss errors beyond predefined criteria, and human reviewers get fatigued and make mistakes at scale.
We built a generative reward model to bridge this gap by training agents to approximate partner review.
We give these agents the original outputs, web search, and other tools to check our systems’ work.
We find that their judgments correlate strongly with expert lawyer review.
This approach will help us scale human review across model training and production products, and will be central to training Tenet 1.5.
显示更多