Generative reward models are a promising direction for scaling up human judgment.
Quality judgments for legal work product depend on many preference drivers like style, tone, writing quality, and document formatting that are less well-represented in current RLVR configurations vs. substantive and objective criteria.
Legal outcomes are also inherently subjective — lawyers write documents called opinions for a reason. Trial judges and appellate judges disagree on 10-15% of cases.
For this reason, our benchmarking + evals have historically relied on human lawyer preference judgments (SxS, Likert) which are high signal but low volume.
In this article,
@ItsJulioPereyra lays out how we’re scaling preference judgments with GRMs and their implications for both eval and training.