We built generative reward models (GRMs) to scale how human experts judge legal work product.
A critical part of how we evaluate models and agent systems at Harvey is legal expert side-by-side review.
Lawyers on our Applied Legal Research team and at our data partners compare two responses to the same task, choose which they prefer, and explain why.
This process is difficult to scale:
1) Expert time is scarce and valuable, and
2) Lawyers often disagree on how to weigh the strengths and weaknesses of the model outputs.
GRMs are LLM judges that compare two responses by generating a rubric for each task, and scoring each response against the rubric.
We extended this approach with:
1) Sub-agents that let the GRM verify factual claims against the environment, source materials, and the web
2) Lawyer-defined rubric dimensions covering topics like accuracy, coverage, reasoning, style, and tone.
GRMs help us scale expert judgment by serving as copilots: they review outputs, identify key differences, and highlight issues that warrant closer expert attention. In our early experiments, they largely reproduced lawyers’ model rankings.
Deep dive by
@ItsJulioPereyra on our early experiments with GRMs and how we're using agents to scale our data and eval efforts: